Skip to main content
Glama

Examplary MCP Server

A Model Context Protocol (MCP) server for Examplary, providing AI assistants with access to the Examplary exam management API.

Features

  • 60+ API Endpoints - Full access to Examplary's exam management capabilities

  • Auto-generated Tools - All API operations automatically available as MCP tools via OpenAPI specification

  • Secure Authentication - User-specific API key authentication

  • Cross-platform - Works on macOS, Linux, and Windows via UV runtime

  • One-click Installation - Install as a .mcpb package in Claude Desktop

Related MCP server: MCPaeroedu

What is Examplary?

Examplary is an AI-powered exam management platform that helps educators:

  • Create and manage exams with AI assistance

  • Generate questions from source materials

  • Grade student responses automatically

  • Organize content in collaborative workspaces

  • Share resources via publisher libraries

Installation

Option 1: Claude Desktop (.mcpb package)

  1. Download the latest examplary-mcp.mcpb from the releases page

  2. Open Claude Desktop

  3. Go to Settings → Extensions

  4. Click "Install Extension" and select the downloaded .mcpb file

  5. When prompted, enter your Examplary API key

Option 2: Manual Installation

  1. Clone this repository:

    git clone https://github.com/examplary-ai/mcp.git
    cd mcp
  2. Install UV (if not already installed):

    curl -LsSf https://astral.sh/uv/install.sh | sh
  3. Run the server:

    export EXAMPLARY_API_KEY="your-api-key-here"
    uv run src/server.py

Getting Your API Key

  1. Log in to Examplary

  2. Navigate to Account → Developer

  3. Click "Generate New API Key"

  4. Copy the API key and save it securely

Note: Keep your API key secret. Never commit it to version control or share it publicly.

Usage

Once installed in Claude Desktop, you can ask Claude to interact with Examplary:

  • "Create a new exam about Python programming"

  • "List all my exams"

  • "Generate questions from this document"

  • "Get the results for exam ID 12345"

  • "Create a new organization workspace"

The MCP server provides access to all Examplary API endpoints, including:

Exam Management

  • Create, read, update, and delete exams

  • Duplicate exams

  • Export to PDF or Word

  • AI-powered question generation

Question Bank

  • Store and organize reusable questions

  • Public and private question types

  • Bulk import from various formats

Student Sessions

  • Create grading sessions

  • Scan documents for answers

  • AI-powered grading and feedback

  • Accept or reject grading suggestions

Organizations & Collaboration

  • Manage workspaces

  • Invite team members

  • Organize exams in folders

Publisher Library

  • Browse public content

  • Share exam templates

  • Discover featured resources

Development

Project Structure

examplary-mcp/
├── manifest.json       # MCPB manifest configuration
├── pyproject.toml      # Python dependencies
├── src/
│   ├── __init__.py
│   └── server.py       # Main MCP server
├── .mcpbignore        # Files excluded from .mcpb package
└── .github/
    └── workflows/
        └── release.yml # Automated release workflow

Local Development

  1. Install dependencies:

    uv pip install -e .
  2. Set your API key:

    export EXAMPLARY_API_KEY="your-api-key"
  3. Run the server:

    uv run src/server.py

Building the .mcpb Package

To build the .mcpb package manually:

zip -r examplary-mcp.mcpb \
  manifest.json \
  pyproject.toml \
  src/ \
  icon.png \
  -x "*.pyc" -x "__pycache__/*" -x ".git/*"

Testing

The server uses stdio transport for communication with Claude Desktop. To test locally:

  1. Run the server with your API key set

  2. The server will start and listen for MCP protocol messages on stdin

  3. Send MCP requests via stdin and receive responses on stdout

API Rate Limits

Examplary enforces rate limits on certain operations:

  • Exam generation: 3 requests per 60 seconds

  • Question operations: 10 requests per 60 seconds

  • Other endpoints: Standard rate limiting applies

Security

  • API keys are stored securely by Claude Desktop

  • All communication uses HTTPS

  • API keys can be revoked at any time from the Examplary dashboard

  • Never share your API key or commit it to version control

Support

License

MIT License - see LICENSE file for details

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository

  2. Create your feature branch (git checkout -b feature/amazing-feature)

  3. Commit your changes (git commit -m 'Add some amazing feature')

  4. Push to the branch (git push origin feature/amazing-feature)

  5. Open a Pull Request

Credits

Built with:

Available Tools

115 tools
attributes_createC

Create a new custom attribute in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
attributeTypeNo
optionsNo
showInStudentDashboardNo
showAsStudentPreferenceNo
positionNo
appearanceColorNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates a mutation operation ('create') but provides no details on side effects, authentication requirements, or constraints. With no annotations, the description carries the full burden but fails to disclose behavioral traits beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence, which is appropriate for a simple tool. However, it could be expanded to add value without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 7 parameters, no output schema, and no annotations, the description is grossly incomplete. It does not address the purpose of parameters, return behavior, or any edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 7 parameters with 0% description coverage. The description adds no meaning to any parameter, leaving the agent to infer usage from names alone. This is insufficient for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('create') and the resource ('custom attribute'), and specifies the scope ('in the current workspace'). It effectively differentiates from sibling tools like 'attributes_update' and 'attributes_delete'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as when to create versus update or delete attributes. No prerequisites or context are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attributes_deleteA

Delete a custom attribute from the current workspace. Default attributes cannot be deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that default attributes cannot be deleted, which provides some behavioral context. However, it does not mention permissions, side effects, or irreversible nature, leaving gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, using two short sentences to convey the essential action and a key constraint. No unnecessary words, front-loaded with the verb.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool with one parameter and no output schema, the description covers the core action and a constraint. However, it omits details like return value confirmation or parameter format, which would be expected for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the 'id' parameter at all. It fails to compensate for the lack of schema documentation, providing no added meaning about what the ID represents or how to obtain it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (delete) and the resource (custom attribute), and distinguishes from sibling tools like attributes_create by focusing on deletion. It also specifies that default attributes cannot be deleted, adding precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives. It implies usage for custom attributes but lacks guidance on when not to use or compare to other attribute tools. The constraint about default attributes is helpful but not sufficient for clear differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attributes_listA

Get all custom attributes configured for the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool performs a read operation ('Get all custom attributes'), but does not mention permissions, rate limits, or response details. For a simple list tool, this is adequate but not exceptional.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, brief sentence that is front-loaded with the key action and resource. Every word serves a purpose with no superfluous text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, no output schema, and performs a straightforward list operation, the description provides all necessary information for an AI agent to understand and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters, so schema coverage is 100%. The description does not add parameter semantics because none are needed. Baseline for 0 parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves all custom attributes for the workspace, using a specific verb ('Get') and resource ('custom attributes'). It distinguishes from sibling tools like attributes_create, attributes_delete, which perform different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context that this tool lists all attributes, but does not explicitly state when to use it versus alternatives. However, the context is sufficient and the sibling names make differentiation obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attributes_reorderC

Update the position of multiple attributes at once.

ParametersJSON Schema
NameRequiredDescriptionDefault
orderNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior but only states 'update the position', implying mutation. It omits details like side effects, required permissions, position indexing, or handling of invalid IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence. While efficient, it sacrifices informativeness; a slightly longer description could include parameter hints without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the description lacks essential context such as position indexing (0-based or 1-based), whether updates are incremental or absolute, and what happens on error. The missing output schema further limits completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should elaborate on the 'order' parameter. However, it adds no meaning beyond the schema structure, leaving the agent to deduce that each item requires 'id' and 'position'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies that the tool updates positions for multiple attributes at once, distinguishing it from sibling tools like attributes_update which likely handles single attribute updates. However, it could be more precise by mentioning 'reorder' explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like attributes_update. No mention of prerequisites or context for reordering attributes, leaving the agent to infer usage without support.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attributes_updateC

Update an existing attribute in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo
attributeTypeNo
optionsNo
showInStudentDashboardNo
showAsStudentPreferenceNo
positionNo
appearanceColorNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It only states 'Update', which implies mutation, but does not mention side effects, permissions, merge vs overwrite, or any constraints. The agent is left unaware of important behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that states the purpose concisely and is front-loaded. However, it is too brief to be fully informative; some key details about parameters or usage could be added without losing conciseness. It is adequate but not excellent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (8 parameters, no output schema, no annotations, many siblings), the description is severely incomplete. It does not cover return values, error handling, prerequisites, or the effects of updates. The agent cannot fully understand how to use this tool correctly from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning no parameter descriptions exist, and the tool description adds no explanations for any of the 8 parameters. The agent must rely on parameter names and types only, which may be insufficient for proper usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Update' and the resource 'existing attribute' and scopes to the current workspace. It distinguishes from siblings such as attributes_create (create), attributes_delete (delete), attributes_list (list), and attributes_reorder (reorder).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites, context, or exclusions. Given siblings like attributes_create and attributes_delete, explicit usage criteria would help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deleteRubricsidC

Delete a rubric from the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only states 'Delete' without elaborating on destructiveness, irreversibility, or side effects on associated resources. This is insufficient for a delete operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but it is too terse and omits critical context. While not verbose, it sacrifices necessary detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter documentation, the description is barely adequate. It misses details on return values, error conditions, and implications of deletion.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no meaning to the 'id' parameter. It does not clarify that 'id' refers to the rubric ID or provide any format or usage constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'delete', the resource 'rubric', and the scope 'from the current workspace'. This distinguishes it from sibling tools like getRubrics (retrieve) and patchRubricsid (update).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool or when not to. It lacks prerequisites (e.g., rubric must exist, permissions needed) and does not mention alternatives or consequences of deletion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

embedSessions_createD

Create a new embed session. This allows you to embed the exam generation flow into your own application.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
statusYes
embedUrlYes
flowYes
actorYes
enabledResponseModesYes
createdAtYes
expiresAtYes
createdByYes
presetsYes
outputsYes
themeYes
metadataYes

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It merely says 'Create', which implies mutation, but provides no details on side effects, permissions, lifecycle, or what the created entity represents. The description is virtually silent on behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but wasteful: the first sentence is a tautology ('Create a new embed session'), and the second sentence is inaccurate. It lacks a front-loaded summary of key details like supported flows or required parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multiple anyOf flows, many parameters), the description is completely inadequate. It does not mention the variety of flows, the required fields in body, or what the output (session token/URL) contains. The presence of an output schema does not excuse the lack of high-level context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no parameter information beyond the name. It does not mention that a body is required, nor the structure of the body (actor, theme, flow, presets, etc.). The description fails to compensate for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Create a new embed session' (verb+resource) but then incorrectly narrows the purpose to 'embed the exam generation flow', ignoring other supported flows like edit-rubric, generate-question, mark-answer, and edit-exam. This misleads the agent about the tool's actual scope and does not distinguish it from sibling tools (embedSessions_get, embedSessions_revoke).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor does it describe the different flows or prerequisites. The agent is left without context for appropriate use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

embedSessions_getB

Retrieve an embed session by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
statusYes
embedUrlYes
flowYes
actorYes
enabledResponseModesYes
createdAtYes
expiresAtYes
createdByYes
presetsYes
outputsYes
themeYes
metadataYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits such as error handling (e.g., if session not found), authentication needs, or side effects. It only states the basic operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no unnecessary words. However, it is so brief that it sacrifices completeness for conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema, the description does not need to explain return values. However, it lacks context about error scenarios, authorization, or typical usage, making it minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% parameter description coverage, and the description adds no extra meaning to the 'id' parameter beyond 'by its ID'. It fails to explain what the ID represents or any constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Retrieve an embed session by its ID', specifying the action (retrieve), resource (embed session), and the key parameter (ID). It distinguishes from sibling tools like embedSessions_create and embedSessions_revoke.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The context implies retrieval, but there is no mention of exclusions, prerequisites, or comparison to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

embedSessions_revokeB

Revoke access to an embed session by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It states 'Revoke access' implying destruction but offers no details on reversibility, permissions, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 8-word sentence is concise and front-loaded, though could include more detail without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a revoke action with no annotations, output schema, or param details, the description is insufficient. It lacks behavioral context like permanence, permissions, or error states.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds 'by its ID' to the single 'id' parameter, but schema coverage is 0% and the parameter name is self-explanatory. Baseline score of 3 as the description does not add significant value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Revoke access') and resource ('embed session'), clearly distinguishing it from sibling tools like embedSessions_create and embedSessions_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor any prerequisites or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_cancelGenerationC

Cancel an ongoing job to generate new questions for the exam using AI.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries full burden but only mentions cancellation. It does not disclose side effects (e.g., partial results, resumability) or required state (e.g., job must be ongoing).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence of 13 words, concise and to the point, though it could benefit from structure like listing parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single parameter and no output schema, the description lacks completeness: it does not explain the parameter, prerequisites, or result of cancellation, leaving critical gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter 'id' with no description, and the tool description does not clarify what the parameter represents (e.g., exam ID, job ID), leaving ambiguity for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Cancel' and the resource 'ongoing job to generate new questions for the exam using AI', differentiating it from other exam tools like 'exams_startGeneration' and 'exams_questions_generate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, such as when to cancel vs. wait for completion, or any conditions like the job must be in progress.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_createC

Create a new exam within your workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
sourceMaterialIdsNo
contextNo
studentLevelNo
subjectNo
languageNo
taxonomyIdNo
intendedDurationNo
allowedGenerationQuestionTypesNo
metadataNo
questionsNo
permissionsNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only states 'create', implying mutation, but does not mention side effects, authorization needs, rate limits, or what happens to existing data. The one-sentence description fails to provide adequate transparency for a creation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (one sentence), but it is under-specified for the tool's complexity (12 parameters). Conciseness should not come at the expense of clarity; here, too much information is omitted, making the description ineffective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (12 parameters, no annotations, no output schema), the description is severely incomplete. It fails to cover return values, parameter details, or any contextual information needed for correct invocation. The description adds minimal value beyond the tool name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 12 parameters with 0% description coverage at the top level. The description does not elaborate on any parameter (e.g., name, sourceMaterialIds, subject). It adds no meaning beyond the bare verb, leaving the agent without guidance on how to populate the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (create) and resource (exam) with scope (within workspace). However, it does not differentiate from sibling tools like exams_duplicate or exams_import, which also create exams but from different sources. The verb 'create' is unambiguous but lacks context for selection among alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus other exams_* tools. There is no mention of prerequisites, limitations, or alternative tools for similar tasks. The description offers no contextual help for choosing this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_deleteC

Delete the exam.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must convey behavioral traits. It only says 'Delete the exam' without disclosing consequences like whether it is irreversible, cascading effects on related entities (e.g., questions, sessions), or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, which is efficient but too minimal. It is front-loaded but lacks any structure or additional context that would help the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deletion tool with no output schema and simple parameters, the description should at least mention that the operation is permanent or any side effects. It is insufficient for an agent to understand the full context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'id' has no description in the schema (0% coverage). The description does not explain what the id represents or what format it should be in, leaving the agent without necessary guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (delete) and the resource (exam). It distinguishes from siblings like exams_cancelGeneration which cancel generation instead of deleting the exam. However, it is very brief and lacks any additional context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention when deletion is appropriate or any prerequisites, such as the exam not being in use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_duplicateC

Duplicate the exam's questions and settings to a new exam.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo
includeSessionsNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It states duplication of questions and settings but does not disclose important behaviors such as whether sessions are duplicated (even though the parameter includeSessions exists), whether the original is modified, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, achieving conciseness, but it is too short given the tool has three parameters. It front-loads the purpose but omits essential parameter info, making it insufficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, no annotations, and parameters with zero description coverage. The description only covers basic purpose, leaving the agent with no information on parameter semantics, return values, or error conditions, rendering it incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no explanation for any of the three parameters (id, name, includeSessions). The agent must guess their meaning from the schema alone, which lacks descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Duplicate' and resource 'exam', specifying it duplicates questions and settings to a new exam, distinguishing it from sibling tools like exams_create which create from scratch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool over alternatives, nor any prerequisites or conditions for use. The minimal description does not help an agent decide when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_export_qti21ZipB

Export an exam as a QTI 2.1 ZIP package.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYesURL to download the exported QTI ZIP package.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided; the description only states the action and output format, omitting behavioral traits such as whether the operation is read-only, required permissions, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no redundancy, but adding structured information (e.g., what the ZIP contains) could improve clarity without adding much length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter) and the presence of an output schema, the description provides basic completeness, but lacks mention of prerequisites or how to obtain the exam ID.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'id' is not described; with 0% schema description coverage, the agent must infer its meaning from context (exam ID), which is minimal and not explicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action 'Export an exam' and the output format 'QTI 2.1 ZIP package', distinguishing it from the sibling tool exams_export_qti3Zip by version.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (to get a QTI 2.1 ZIP), but does not provide explicit guidance on prerequisites, when not to use it, or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_export_qti3ZipB

Export an exam as a QTI 3 ZIP package.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYesURL to download the exported QTI ZIP package.

TDQS

B3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only states the purpose without disclosing behavioral traits such as destructive potential, read-only nature, processing time, or whether the output is a file or URL. The description fails to compensate for the missing annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that is front-loaded with the core purpose. Every word carries meaning without redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of a sibling tool for QTI 2.1, the description should differentiate or mention the specific QTI version. Also, an output schema exists but is not visible; the description should hint at the output type (e.g., file download). This leaves the agent with incomplete information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one required parameter 'id' with no schema description coverage (0%). The description implies 'id' is the exam identifier, adding minimal context beyond the schema. For a single parameter, this is adequate but not informative about format or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool exports an exam as a QTI 3 ZIP package, specifying the verb (export), resource (exam), and format (QTI 3 ZIP). This differentiates it from the sibling tool 'exams_export_qti21Zip' by version number.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'exams_export_qti21Zip' or other export options. The description does not mention prerequisites, required permissions, or context for usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_getA

Get a single exam by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description only says 'Get' with no mention of side effects, permissions, or error behavior (e.g., not found case). Minimal but sufficient for a simple read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no unnecessary words, perfectly concise and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-ID tool with no output schema, the description is largely complete. It could mention what fields are returned, but the core intent is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter 'id' with no additional description; schema coverage is 0%, but the parameter is self-explanatory. The description adds no extra meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get a single exam by its ID' clearly states the action (get) and resource (exam), and distinguishes from sibling tools like exams_list (multiple) and exams_create/delete/update (mutations).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a single exam ID is known, but lacks explicit guidance on when not to use it (e.g., for multiple exams use exams_list) or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_getContextSuggestionsA

Get AI-generated suggestions for context/instructions to include in the exam generation prompt, based on the source materials linked to the exam.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It indicates the tool is AI-generated and based on source materials, but does not disclose behavioral traits like whether it is read-only, requires authentication, or has any side effects. While not misleading, it lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the action and resource. Every word adds value, with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (one parameter, no output schema), the description provides a clear purpose and context. It explains the input (exam id) and output (suggestions) adequately. While more detail on return format would be helpful, the description is largely complete for its role.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description connects the 'id' parameter to the exam by stating 'based on the source materials linked to the exam', adding meaning beyond the schema's bare 'string' type. However, the parameter is not explicitly described, and schema coverage is 0%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get', the resource 'AI-generated suggestions for context/instructions', and the context 'based on the source materials linked to the exam'. It distinguishes this tool from siblings like exams_questions_generate which generate questions, not context suggestions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you need suggestions for context/instructions for the exam generation prompt, but it does not explicitly state when to use it vs alternatives or provide any exclusions. Usage context is clear but guidance is minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_importC

Create an exam by importing one or more source materials (e.g. an existing exam).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
sourceMaterialIdsNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only says 'Create an exam by importing', omitting details on side effects, permissions, or whether source materials are copied or linked. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise (one sentence), but at the cost of omitting essential details. Conciseness without completeness reduces utility for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of an import operation, the description lacks information about return values, error handling, or what happens post-import. Incomplete for a creation tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 0% description coverage; the description only hints at sourceMaterialIds with the example, but does not explain the 'name' parameter or that parameters are optional. Insufficient to guide correct parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action (create) and resource (exam) with method (importing source materials). Provides example in parentheses, distinguishing it from exams_create or exams_duplicate. However, 'source materials' could be more precisely defined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like exams_create, exams_duplicate, or exams_questions_import. Lacks any context about prerequisites or preferred use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_listB

Get a list of all exams within your workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
folderNoOptional filter by folder ID, e.g. `folder_1234`.
roleNoOptional filter by permission role.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It implies a read operation but does not disclose potential side effects, pagination behavior, or scope limitations beyond 'within your workspace'. This minimal disclosure results in a score of 2.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that conveys the core functionality without any extraneous words. It is well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (two optional parameters, no output schema), the description is adequate but lacks details about the return structure (e.g., fields, pagination). It provides enough context for basic use but is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters (folder and role). The tool description adds no additional meaning beyond what the schema already provides, yielding a baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Get a list' and the resource 'all exams within your workspace'. It is a specific verb+resource, but it does not explicitly differentiate from sibling tools like exams_get or exams_sessions_list, so it scores 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention scenarios such as filtering by folder or role, nor does it indicate when to use other list tools like exams_sessions_list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_printB

Export an exam to a PDF file or Word document.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must reveal behavioral traits. It only states 'Export' without clarifying side effects, permissions, or non-destructive nature. It does not confirm whether the tool alters the exam or merely generates a file, leaving behavior partially ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at one sentence (9 words), but it sacrifices necessary detail for brevity. While it front-loads the core action, it omits critical information about parameters and output, making it borderline under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple interface (one parameter, no output schema), the description should be sufficient, but it lacks details on output format selection or result behavior (e.g., file download, URL). The absence of any behavioral or output information makes it incomplete for an agent to reliably use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the single parameter 'id'. The description does not explain the parameter's meaning (e.g., exam ID) or expected format, relying entirely on the parameter name. The brief description fails to compensate for the lack of schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Export an exam') and the output formats ('PDF file or Word document'). It distinguishes itself from sibling tools like exams_export_qti21Zip and exams_export_qti3Zip which export to QTI formats, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for PDF/Word export but provides no explicit guidance on when to use this tool versus alternatives (e.g., QTI export). No prerequisites or when-not conditions are mentioned, leaving the agent to infer context from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_questions_generateC

Generate a new question for the exam using an AI prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
messageNo
questionTypeNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description fails to disclose behavioral traits such as whether the tool is synchronous, returns a result, or has any side effects. It only states it generates a question, leaving agents unaware of potential wait times or authorization needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, making it concise. However, it lacks structure and does not front-load key details like required parameters or output expectations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 3 parameters with no descriptions, no output schema, and no annotations, the description is grossly incomplete. It fails to explain what 'id' represents or what the tool returns, making it difficult for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain parameters but does not. 'id' is required but unclear (exam ID or question ID?), 'message' and 'questionType' are not explained at all, leaving the agent to guess their meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a new question for an exam using an AI prompt. It distinguishes from the sibling 'exams_questions_regenerate' with the word 'new'. However, it lacks specificity about the exam context or how the AI prompt is used.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'exams_questions_regenerate' or 'exams_startGeneration'. There is no mention of prerequisites or context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_questions_getD

Retrieve a question by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesExam ID, e.g. `exam_1234`.
questionIdYesQuestion ID, e.g. `question_1234`, or `first` to get the first question.
formatNoThe format to return the question in. `examplary_json` returns the question in the Examplary JSON format, while `qti_v3p0` and `qti_v2p1` return the question in QTI format version 3 and 2.1 respectively.examplary_json

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It fails to mention that the tool is read-only, what happens if the question is not found, or any side effects. The minimal description does not compensate for missing annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but lacks structure and important details. It is too sparse to be genuinely helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and the complexity of having two required IDs plus an optional format, the description is incomplete. It does not hint at return values, pagination, or error conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for each parameter, but the description adds no additional meaning. It even misleads by implying a single ID when two are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says 'Retrieve a question by its ID' but the input schema requires both an exam ID and a question ID, which is misleading. The verb 'Retrieve' is appropriate but the target is imprecisely specified.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus siblings like exams_questions_generate or exams_questions_import. No exclusions or alternatives mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_questions_importB

Import questions into an exam from a file. Support Word, PDF, Moodle quiz XML and plain text files.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNoFile name
typeNoMIME type of the file, e.g. application/pdf
urlNoURL of the file, required if no content is specified.
contentNoContent of the file as a string, as an alternative to URL. Only supported for exports from other systems (e.g. Moodle XML).
keepExistingSettingsNoWhether to keep existing exam settings when importing questions via AI.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose behavioral details such as whether importing overwrites or appends questions, side effects on existing exam data, required permissions, or error handling. For a mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary purpose, and wastes no words. It could be slightly more structured (e.g., separate behavior from file types) but is efficient for its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters, no output schema, and no annotations, the description should cover constraints, behavior (replace vs. merge), and potential errors. It does not mention that 'keepExistingSettings' is supported or what happens if neither 'url' nor 'content' is provided. The description is incomplete for reliable agent invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 83% of parameters with descriptions. The description adds value by clarifying supported file types and that 'content' is for exports from other systems. However, it does not explain the relationship between 'url' and 'content' or the meaning of 'keepExistingSettings' in detail. Overall, it provides marginal additional meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the exact action ('Import questions into an exam'), the target entity ('exam'), and supported file types ('Word, PDF, Moodle quiz XML and plain text files'). This clearly distinguishes from sibling tools like exams_import (imports whole exams) and exams_questions_generate (generates questions).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when needing to import pre-existing questions from files, but it does not explicitly state when to use versus alternatives (e.g., exams_import, exams_questions_generate) or provide any when-not-to-use guidance. The agent is left to infer context from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_questions_regenerateC

Update a question using an AI prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
questionIdYes
messageNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully convey behavioral traits. It indicates a write operation but does not disclose any side effects, permission requirements, rate limits, or the impact on the question. The behavior remains mostly opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it lacks essential details for an AI agent to correctly invoke the tool. It is under-specified rather than efficiently informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, no output schema, and no annotations, the description is woefully incomplete. An agent cannot determine the proper usage, required inputs, or expected outcomes from this description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain parameters. It only hints that 'message' is the AI prompt but does not explain 'id' and 'questionId'. The optionality and constraints of 'message' are not clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Update') and the resource ('a question') and the method ('using an AI prompt'), making the purpose understandable. However, it does not explicitly differentiate from siblings like exams_questions_generate, though the context suggests distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, conditions, or scenarios where this tool is preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_acceptAllSuggestionsC

Accept all AI-generated grading suggestions for the exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a state modification (accepting suggestions) but does not disclose whether the operation is destructive, reversible, or what the effects are on the suggestions. This lack of transparency leaves the agent uninformed about potential side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence without any redundancy. It is front-loaded with the key action. However, it could be slightly improved by including brief parameter explanations without becoming verbose, so it loses a point for being too sparse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema, no annotations, and zero parameter coverage, the description is severely incomplete. It fails to provide essential context such as parameter details, behavioral traits, return values, or usage constraints. An agent would struggle to correctly invoke this tool based solely on this description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate by explaining the parameters. However, it does not mention 'id' or 'sessionId' at all. The parameter names are somewhat self-explanatory but insufficient; the description should clarify which identifier corresponds to what (e.g., the exam session or suggestion). The agent has no way to know how to correctly populate these fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Accept all AI-generated grading suggestions') and the specific context ('for the exam session'). It effectively distinguishes from the sibling tool 'exams_sessions_acceptSuggestion' by using 'all', indicating a batch operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor does it mention any prerequisites or conditions. It only states what the tool does without any context for appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_acceptSuggestionC

Accept a single AI-generated grading suggestion for the exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes
questionIdYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description only states 'Accept' which implies mutation, but provides no details on effects, permanence, or required state. With no annotations, the description fails to disclose behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words, but is too sparse for the required information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the three parameters, no output schema, and many sibling tools, the description fails to explain the relationships between parameters or provide enough context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has three required parameters (id, sessionId, questionId) with no descriptions. Schema coverage is 0%, and the description adds no explanation of these fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Accept') and the resource ('a single AI-generated grading suggestion for the exam session'). It distinguishes from the sibling tool acceptAllSuggestions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like acceptAllSuggestions. No context about prerequisites or conditions is mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_deleteC

Remove an exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description only states 'Remove an exam session' without disclosing permanence, cascading effects, or error conditions. For a deletion tool, this is insufficient behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but not informative. It front-loads the action but lacks structure or additional details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and two undocumented parameters, the description fails to provide adequate context for correct invocation. The agent is left with ambiguity about required inputs and effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters (id, sessionId) are untyped strings with no descriptions in schema or description. The agent cannot discern which is which (e.g., id likely refers to exam id, sessionId to the session). Schema coverage is 0%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Remove an exam session' uses a specific verb and resource, clearly distinguishing it from siblings like exams_sessions_list or exams_sessions_get. However, it could be more precise about what 'remove' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, prerequisites, or conditions (e.g., session must exist and not be in progress). The description lacks context for appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_editAnswerC

Edit a single answer in an exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes
questionIdYes
answerNo
autoGradeNoWhether to automatically grade the answer after editing

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description is solely responsible for behavioral transparency. It merely says 'Edit' without disclosing side effects, whether it triggers auto-grading (though an 'autoGrade' parameter exists), or any state changes beyond the answer field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (one sentence) but overly terse. It front-loads the purpose but omits critical details, making it less useful than a slightly longer, more informative description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex input schema (nested object, mandatory fields), no output schema, and no annotations, the description is insufficient. It does not explain return values, prerequisites, or behavior such as idempotency or update timestamps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only autoGrade has a description). The tool description adds no meaning to the parameters id, sessionId, questionId, or the nested answer object, leaving their semantics unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Edit') and the resource ('a single answer in an exam session'), making the purpose straightforward. However, it does not differentiate from sibling tools that might also modify session data, such as exams_sessions_acceptSuggestion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidelines are provided on when to use this tool versus alternatives like exams_sessions_saveFeedback or exams_sessions_acceptSuggestion. The description lacks context for appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_generateOverallFeedbackC

Generate feedback for the test as a whole in an exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but only says 'generate feedback'. It fails to disclose side effects (e.g., whether it overwrites existing feedback), authentication requirements, rate limits, or if it is a read-only or mutation operation. The term 'generate' implies mutation, but details are absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but it is under-specified. While concise, it omits critical details that would fit in 2-3 sentences. Brevity here sacrifices usefulness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of generating overall feedback, the absence of output schema, annotations, and detailed description leaves the agent with insufficient context to understand prerequisites, return format, or expected behavior. The description is incomplete for effective tool selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with no parameter descriptions in the schema. The description does not explain what 'id' and 'sessionId' represent (likely exam or session identifiers), leaving the agent to guess. The description adds no semantic value beyond the parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates overall feedback for an exam session, specifying both the action (generate) and resource (feedback for test as a whole). However, it does not differentiate from sibling tools like exams_sessions_saveOverallFeedback, which could be confused as a similar but distinct operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., exams_sessions_saveOverallFeedback or per-question feedback generation). Prerequisites or context (e.g., requires answers to be present, should be called after scoring) are missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_getC

Get a single exam session by its ID, representing a set of answers from a student.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It only states it gets a session, without mentioning any side effects, permissions, or data volume. This is insufficient for a mission-critical tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 14-word sentence with the verb upfront. No wasted words, perfectly front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple retrieval tool, the description omits return format, relationship to other objects, and how to interpret the session. Given the lack of output schema and annotations, the description should provide more context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has two required parameters (id, sessionId) with 0% description coverage, but the description does not explain the difference between them or their roles. It simply says 'by its ID' without clarifying which parameter is which, adding no value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves a single exam session by ID and defines its content as student answers. It distinguishes itself from sibling like exams_sessions_list by emphasizing 'single', though it could explicitly contrast with list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus siblings like exams_sessions_list or exams_get. The description implies usage when you have a session ID, but does not state prerequisites or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_importC

Manually create a session, possibly from uploaded data.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
studentNameNo
sourceNomanual
answersNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It states 'manually create' and 'possibly from uploaded data', but does not clarify whether this is a destructive operation, required permissions, or data validation rules. This is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but it is vague and does not front-load critical information. It is under-informative rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (4 params, nested answers array) and lack of annotations or output schema, the description is far from complete. It omits essential details about session creation, data handling, and return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%; the description does not explain any of the four parameters (id, studentName, source, answers). The complex answers array and enum source are left undefined, forcing the agent to rely solely on the schema without additional context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it creates a session manually, possibly from uploaded data. However, the name 'import' suggests batch input, and the description lacks specificity about the tool's exact role among sibling session tools. It is not a tautology but remains vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool over alternatives like exams_sessions_scan or exams_sessions_editAnswer. There is no mention of prerequisites or exclusions, leaving the agent without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_listA

Get a list of all student sessions for an exam, representing a set of answers from a student.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears the full burden of disclosing behavioral traits. It only states the basic operation (listing sessions) without mentioning read-only nature, permission requirements, pagination, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 14 words, front-loading the core purpose. Every word is necessary, and there is no redundant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no output schema, the description adequately conveys that the result is a list of student sessions as sets of answers. However, it omits details about response format, filtering, ordering, or pagination, which would be helpful for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single required parameter 'id' with 0% description coverage. The description adds meaning by implying 'id' refers to the exam ('for an exam'), but does not explicitly state that it is the exam ID or provide further semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Get' and identifies the resource as 'list of all student sessions for an exam'. It clearly differentiates from sibling tools like exams_sessions_get (single session) and exams_sessions_delete (delete) through the 'list' action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (listing all sessions for an exam) but does not provide explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives like exams_sessions_get for retrieving a single session.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_saveFeedbackB

Set teacher feedback and/or a grade for a specific question in an exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes
questionNoQuestion ID to provide feedback for
pointsAwardedNoPoints to award for the answer
feedbackNoOptional teacher feedback for the answer

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must disclose behavioral traits. It indicates mutation ('Set') but does not mention permissions, idempotency, side effects, or return values. The tool may overwrite existing feedback, but this is not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no wasted words. However, it lacks structural elements like a use case or example. It is appropriately sized but could benefit from a second sentence for context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description should explain what the tool returns (e.g., success confirmation or updated object). It does not. It also omits prerequisites (e.g., session and question must exist). The information provided is insufficient for an agent to fully understand tool behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description maps high-level concepts ('feedback' and 'grade') to the parameters 'feedback' and 'pointsAwarded', but the schema already describes these with comments. With 60% schema coverage, the description adds minimal new meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Set') and the resource ('teacher feedback and/or a grade for a specific question in an exam session'). This distinguishes it from sibling tools like 'exams_sessions_saveOverallFeedback' which targets the entire session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for providing feedback on individual questions but does not explicitly mention when to use this tool versus alternatives (e.g., 'saveOverallFeedback' or 'editAnswer'). No exclusions or when-not-to-use guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_saveOverallFeedbackC

Set teacher feedback for the test as a whole in an exam session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes
feedbackNoTeacher feedback for the test as a whole

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility. It implies a write operation but does not disclose whether feedback is overwritten, merged, or cleared (despite schema allowing null). Missing details on idempotency, permissions, or side effects like triggering notifications or status changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no waste. However, it is too brief given the missing contextual information (e.g., parameter roles, when to use). Conciseness does not compensate for lack of completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and only 3 parameters with minimal schema coverage. The description fails to explain return values, error states, or prerequisites (e.g., session must exist, feedback must not be generated already). Incomplete for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only 33% of parameters have descriptions in the schema (feedback). The description adds no new meaning beyond the schema: id and sessionId are unexplained, and the feedback parameter description in the schema already states 'Teacher feedback for the test as a whole'. The null option is not clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Set teacher feedback for the test as a whole') and the resource ('exam session'). It distinguishes from sibling tools like exams_sessions_generateOverallFeedback (generation vs manual setting) and exams_sessions_saveFeedback (per-question vs overall).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., generateOverallFeedback, saveFeedback). No mention of prerequisites, such as whether feedback must be generated first or whether the session must be in a specific state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_scanC

Scan student answers from a document.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must convey behavior. It fails to disclose whether the operation is idempotent, asynchronous, or modifies state. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but it is overly terse for the tool's complexity, sacrificing necessary detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and minimal parameter info, the description should provide context on return values, side effects, and prerequisites. It provides none.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has one required 'id' parameter with 0% description coverage. The description does not explain what 'id' refers to (e.g., exam session ID, document ID), leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Scan student answers from a document.' which indicates the action and resource, but the verb 'scan' is ambiguous (optical vs. digital) and the description does not distinguish this from sibling tool 'exams_sessions_scanDocument'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. With a similarly named sibling, explicit usage context is needed but absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_sessions_scanDocumentC

Start a background task to extract student answers from an uploaded PDF document.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
documentUrlNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description mentions 'background task' implying asynchrony, but does not explain how to monitor progress or retrieve results. No annotations are provided, so the description should disclose behavioral traits like mutability, side effects, or required permissions. It does not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence without waste, but it is too concise and lacks essential details, making it less helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and no parameter explanations, the description is incomplete. The agent cannot fully understand how to use the tool or what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fails to explain the meaning or purpose of the 'id' and 'documentUrl' parameters. The agent has to infer from context, which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (start background task) and the resource (extract student answers from PDF). However, it does not differentiate from the sibling tool exams_sessions_scan, which may perform a similar function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. No context about prerequisites or integration with other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_startGenerationC

Initiate a job to generate new questions for the exam using AI, based on the source materials and metadata specified.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should disclose behavioral traits. It mentions 'initiate a job' implying asynchronicity, but lacks details on side effects, permissions, or how to track progress.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence that directly states the action and key inputs. No unnecessary words, though it could be slightly more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (job initiation, AI generation) and lack of output schema, the description should explain return value, job tracking, or next steps. It does not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'id' is not described at all, despite 0% schema description coverage. The description fails to explain what 'id' refers to or how it relates to generation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it initiates a job to generate exam questions using AI, based on specified source materials and metadata. However, it does not distinguish itself from similar siblings like 'exams_questions_generate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no prerequisites or context provided. The description is purely declarative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exams_updateC

Update an existing exam.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must fully disclose behavioral traits. It only states 'Update an existing exam,' implying a mutation, but fails to indicate if the update is partial or full, what happens on error, or whether it requires specific permissions. This is insufficient for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but lacks substance. While it avoids verbosity, it does not add value beyond the tool's name. It is minimal but not tautological, earning a mid-range score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and minimal input description, the description fails to provide complete context for a mutation tool. It does not address return values, error scenarios, or the scope of updates, making it inadequate for correct usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage for its single parameter 'id', and the description does not mention or explain this parameter. The description 'Update an existing exam' does not clarify that the exam ID is required or how to supply updated fields, leaving the agent without essential usage details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Update' and the resource 'exam', making the tool's purpose unambiguous. However, it does not differentiate from sibling tools beyond the name, which is acceptable since 'exams_update' is the only update tool for exams. The lack of detail on what fields can be updated slightly reduces clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like exams_create or exams_delete. There is no mention of prerequisites (e.g., exam must exist) or context such as typical workflows, leaving the agent without sufficient decision-making information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

folders_createC

Create a new folder in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only indicates a write operation without disclosing effects, permissions needed, or any side effects. It fails to add behavioral context beyond the verb 'create'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but it lacks structure such as separating purpose from usage. It is appropriately short for a simple tool but could benefit from more detail without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and minimal annotations, the description should compensate but only states the basic action. It omits important details like behavior when 'name' is omitted, potential errors, or return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the tool description does not explain the 'name' parameter beyond what the schema provides (a string with length constraints). No additional meaning or constraints are given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (create) and resource (folder) with scope ('in the current workspace'). It distinguishes from sibling tools like folders_delete and folders_list which handle different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as when to create vs. update a folder. No prerequisites or conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

folders_deleteA

Delete a folder from the current workspace. If there are any exams in the folder, they will be moved out of the folder.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait: if there are exams in the folder, they will be moved out. However, it does not mention irreversibility, required permissions, or other potential side effects. With no annotations, more detail would improve transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences front-load the primary action ('Delete a folder'), with no wasted words. Every sentence adds value, especially the behavior of moving exams.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main effect and a notable side effect. However, it lacks information about return values, error cases, or prerequisites. Given the tool's simplicity, it is minimally complete but has gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter 'id' with no description (0% coverage). The description does not explain the parameter beyond implying it's the folder ID. It provides no additional meaning or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Delete a folder' and specifies the resource is from the current workspace. It distinguishes from sibling tools like folders_create, folders_list, and folders_update by being the delete operation, and adds a unique detail about moving exams out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description is clear about when to use it (to delete a folder). There are no explicit alternatives or when-not-to-use scenarios, but the context of sibling tools makes it obvious. For a simple delete, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

folders_listA

Get a list of folders for exam organisation that exist within the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates a read-only, non-destructive operation ('Get a list'). With no annotations, it adds sufficient behavioral context. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with verb and resource, no redundant words. Highly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description fully covers what the tool does: listing folders for exam organisation in the current workspace.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, so schema coverage is 100%. Description adds no parameter info, but baseline for 0 parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and resource 'folders', specifies scope 'within the current workspace' and purpose 'for exam organisation', distinguishing it from sibling tools like folders_create or folders_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like folders_list, or when not to use it. However, for a simple list operation, the usage context is implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

folders_updateC

Update an existing folder in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the burden of behavioral disclosure. It only says 'Update', which implies mutation, but no details on side effects, reversibility, or required authorization.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no fluff, but it is overly minimal. It could include parameter details without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description fails to provide sufficient context. It omits return value, prerequisites, and any behavioral traits beyond the basic action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the 'id' (required) or 'name' (optional with constraints) parameters. The parameters are completely undocumented beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Update', the resource 'folder', and importantly specifies 'existing in the current workspace', distinguishing it from sibling tools like folders_create, folders_delete, and folders_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor any preconditions like folder ownership or permissions. The description is purely declarative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

getRubricsB

Get a list of rubrics that exist within the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only says 'Get a list of rubrics that exist within the current workspace.' It fails to disclose any behavioral traits such as pagination, ordering, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no wasted words, perfectly concise for a 0-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is too brief; it lacks output format details, pagination information, and does not set expectations for the returned list. Given no output schema and no annotations, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters (0), and schema description coverage is 100%. Baseline of 4 applies as the description need not add parameter info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'rubrics', and distinguishes from sibling tools like postRubrics (create) and deleteRubricsid (delete).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., attributes_list for other lists). No exclusions or context provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_addMemberA

Add a member to a group. Requires group manager/owner role or org admin/owner.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesGroup ID
bodyNo

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It only mentions the action and role requirement, omitting details such as idempotency, behavior if the member already exists, or any side effects. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that conveys purpose and permission requirements without any extraneous words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two complex request body alternatives (by actor or email) and no output schema. The description does not explain these alternatives nor any return values, making it incomplete for effective invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has two parameters: 'id' (described as 'Group ID') and 'body' (undescribed, with two alternative formats). Schema description coverage is 50%, and the description adds no extra parameter details. The agent must rely solely on the schema, which lacks documentation for the 'body' parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Add a member to a group.' It distinguishes from sibling tools like 'groups_removeMember' and 'groups_updateMember' by focusing on the addition action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies the required roles: 'Requires group manager/owner role or org admin/owner.' This provides clear context on when the tool can be used, though it does not explicitly mention when not to use it or list alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_createA

Create a new group. Requires admin or owner org role.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It mentions the role requirement but fails to address potential conflicts (e.g., duplicate group names) or the result of creation (e.g., no response details).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, using two short sentences with no superfluous words. It is front-loaded with the main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description is somewhat complete but lacks details on parameter semantics and what the tool returns after creation, which would be expected for a creation action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate by explaining the 'name' parameter. It does not reference the parameter at all, leaving its meaning implied by the parameter name alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Create a new group', providing a specific verb and resource. It implicitly distinguishes from siblings like groups_delete and groups_list by focusing on creation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies a prerequisite role ('Requires admin or owner org role'), guiding when to use the tool. However, it does not explicitly mention when not to use it or compare to alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_deleteA

Delete a group and revoke all associated permissions. Requires group manager/owner role or org admin/owner.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesGroup ID

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action (delete) and the side effect (revoke all associated permissions), which is clear and transparent for a deletion tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: first states action and side effect, second states required role. Front-loaded, no wasted words, efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool with one parameter and no output schema, the description provides all necessary context: what it does, side effect, and authorization requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100% with a single parameter 'id' described as 'Group ID'. The description does not add any additional meaning beyond what the schema provides, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Delete a group and revoke all associated permissions,' using a specific verb and resource. It distinguishes itself from sibling tools like groups_create, groups_update, and groups_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies the required roles: 'Requires group manager/owner role or org admin/owner.' This provides clear context for who can use the tool, but does not explicitly mention when not to use it or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_listA

List all groups in the organization. When restricted group visibility is enabled, only returns groups the user belongs to.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description discloses an important behavioral nuance: that when restricted group visibility is enabled, the tool only returns groups the user belongs to. This adds transparency beyond the function name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two short sentences with no excess verbiage. The purpose is front-loaded, and every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential purpose and a key condition. It is complete for a simple list tool with no parameters, though it could optionally mention output format or ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so the baseline is 4. The description adds context by explaining what the tool returns and the visibility condition, which goes beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'list all groups' and the resource 'groups in the organization'. It distinguishes from sibling tools like groups_listMembers which list members of a group, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides conditional guidance on when results may be limited (restricted visibility). While it doesn't explicitly exclude alternatives, the context implies it is the primary list tool for groups, and the guidance is helpful for understanding output variations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_listMembersA

List all members of a group with resolved user info. Respects restricted group visibility.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesGroup ID

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so description carries full burden. It reveals that user info is resolved and visibility restrictions are respected, which adds behavioral context beyond the schema. However, it omits details on read-only nature, pagination, or performance traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences, front-loaded with the primary purpose, and every word adds value. Zero wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool with one parameter and no output schema, the description covers core functionality and visibility considerations. However, details about return format, pagination, or member ordering are missing, which would aid completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the description does not add extra meaning to the single parameter 'id' beyond what the schema provides ('Group ID'). The baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly uses the verb 'list' and the resource 'members of a group' with additional detail 'with resolved user info' and 'respects restricted group visibility'. It distinguishes itself from sibling tools like groups_addMember, groups_removeMember, and groups_list by specifying membership retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for listing group members but does not explicitly state when to use this tool over alternatives like groups_list (which lists groups) or groups_addMember. No context on prerequisites or when not to use is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_removeMemberA

Remove a member from the group. Requires group manager/owner role or org admin/owner. Cannot remove the owner.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesGroup ID
userIdYesUser ID of the member

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It adds role requirements and the owner restriction, but does not mention side effects, reversibility, or error conditions. This is minimal for a destructive action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, straightforward, and contains no unnecessary words. It front-loads the action and then adds role and constraint details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with no output schema, the description covers purpose, requirements, and a constraint. However, it lacks details on outcome, error handling, or idempotency, which would be helpful for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for both parameters ('Group ID', 'User ID of the member'). The description does not add any extra meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Remove a member from the group', with the verb 'remove' and the resource 'member from group'. It distinguishes from sibling tools like groups_addMember and groups_updateMember.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies required roles ('group manager/owner role or org admin/owner') and a key constraint ('Cannot remove the owner'). While it does not explicitly compare to alternatives, the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_updateA

Rename a group. Requires group manager/owner role or org admin/owner.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesGroup ID
nameNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries full burden. It discloses role requirements, but omits other behavioral traits such as potential side effects (e.g., breaking references), whether the operation is reversible, or what the response looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences, no fluff. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite being a simple tool, it lacks critical context: no output schema, no mention of response, and ambiguity about the optional name parameter (if omitted, does it rename?).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% (id has a brief description, name has none). The description does not add meaning to the name parameter, such as its optionality or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Rename a group'. This specific verb-resource pair distinguishes it from siblings like groups_updateMember (which updates membership) and groups_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context by specifying required roles ('group manager/owner role or org admin/owner'), but does not explicitly state when not to use this tool or compare it to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

groups_updateMemberA

Change a member's role in the group. Requires group manager/owner role or org admin/owner.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesGroup ID
userIdYesUser ID of the member
roleNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It reveals the permission requirement, which is a key behavioral trait. However, it does not disclose potential side effects, idempotency, or error states, which would be beneficial for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise—two sentences, front-loaded, with no extraneous information. Every sentence serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple mutation with three parameters and no output schema, the description adequately covers purpose and prerequisites. However, it could mention the expected response (e.g., success/failure) for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, meaning two parameters (id, userId) have descriptions, and role has an enum. The description adds no extra information beyond the schema. Given moderate coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Change a member's role in the group', which is a specific verb (change) and resource (member's role in the group). This distinguishes it from sibling tools like groups_addMember, groups_removeMember, and groups_listMembers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It specifies the required roles: 'Requires group manager/owner role or org admin/owner.' This provides clear context on who can use the tool, though it does not explicitly state when not to use it (e.g., for adding/removing members).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobs_cancelC

Cancel a background job.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose any behavioral traits beyond the basic action. With no annotations, it fails to explain side effects, reversibility, or what happens to the job after cancellation. The agent gets minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (one sentence) but lacks necessary information. Conciseness should not sacrifice completeness; here it is under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a single parameter, the description should provide comprehensive context (e.g., success/error responses, idempotency, concurrency). It fails to do so, leaving the agent with inadequate information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'id' is not described beyond the schema. With 0% schema coverage, the description should clarify what 'id' refers to (e.g., job ID, task ID) but does not, adding no value over the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb 'cancel' and the resource 'background job', which clearly indicates the tool's function. However, it lacks detail to differentiate from other cancellation tools like exams_cancelGeneration, though the context of 'background jobs' makes it distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as jobs_get or other cancel tools. There is no mention of prerequisites, conditions, or context for cancellation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobs_getC

Poll the status of a background job.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral details, but it only says 'poll the status'. It does not explain what the returned status looks like, whether the tool is idempotent, error conditions, or any rate limiting. The agent lacks sufficient information about the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but too minimal. It lacks essential information, making it under-specified rather than efficiently concise. A well-structured description would front-load the purpose and then add necessary details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no parameter descriptions, and no annotations, the description should be more complete. It provides only the basic purpose, omitting details about return values, possible errors, and usage context. For a polling tool, this is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter 'id' with no description (0% coverage). The description does not add any meaning to this parameter; it simply repeats the parameter name. The agent cannot infer what 'id' refers to or its format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'poll' and clearly identifies the resource as 'the status of a background job'. It distinguishes from sibling tools like jobs_cancel, which is about canceling jobs, thus making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., when to poll a job rather than cancel it). There is no mention of prerequisites (e.g., job must exist) or context for polling frequency.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

media_uploadC

Returns a signed URL for uploading a file, and a public URL for accessing it afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
filenameNoOriginal attachment filenameimage.jpg
typeNoCategory of file, used for groupingavatar
contentTypeNoMIME type of the attachmentimage/jpeg

Output Schema

ParametersJSON Schema
NameRequiredDescription
uploadUrlNo
publicUrlNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the output (signed and public URLs) but lacks behavioral details such as authentication requirements, URL expiration, or that the actual file upload is separate. With no annotations, the description carries the full burden, which is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that communicates the core function without extraneous information. Every word earns its place, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite good schema coverage and an output schema (not described), the description omits critical context: it does not explain that the signed URL is used for subsequent uploads, mention URL expiration, or provide usage context among many sibling tools. The tool's simplicity does not excuse the lack of behavioral and practical details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with defaults and descriptions for all three parameters (filename, type, contentType). The description adds no additional meaning beyond the schema, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a signed URL for uploading and a public URL for access, specifying the verb 'returns' and resource 'URLs for file'. The name 'media_upload' reinforces the purpose. However, it does not explicitly differentiate from potential sibling upload tools, though siblings are unrelated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, context, or exclusion cases, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

me_getB

Get the current user's account details.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and description only states a read operation. Does not disclose authentication requirements, rate limits, or any side effects. Minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that is clear and to the point. No superfluous words. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description lacks details about returned data or limitations. For a tool with many siblings, more context on scope and output would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has zero parameters with 100% coverage, so baseline is 4. Description adds no additional parameter info, but none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'Get' and the resource 'current user's account details', making it specific and distinct from siblings like me_update or users_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., me_update for modifications or users_get for other users). Does not mention prerequisites or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

me_updateC

Update the current user's account details.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe user's name
preferencesNoA map of user preferences

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only states 'Update', implying a write operation, but omits details like required permissions, immediacy of effects, or any side effects. This is insufficient for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence is concise and front-loaded with the action and resource. While minimal, it avoids unnecessary words, earning a high conciseness score. Some structure improvement could include separating purpose and behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks information about return values (whether the updated account is returned) and any behavioral context. For a mutation tool with no output schema, this omission makes the description incomplete for an agent to anticipate outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the description adds no new meaning beyond the schema's description of name and preferences. The tool is simple, and the description does not elaborate on parameter constraints or usage, resulting in a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies the verb 'Update' and the resource 'current user's account details', clearly distinguishing it from reading tools like me_get. However, it could be more specific about the scope of updates beyond name and preferences, though the schema fills this gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description does not indicate when to use this tool versus alternatives such as me_get for reading or users_update for updating other users, leaving the agent without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_customDomain_getA

Get custom domain configuration

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral disclosure. While 'Get' implies read-only, it does not mention authentication needs, rate limits, or what happens if configuration is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise at four words, front-loaded with no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with no parameters and no output schema, the description adequately conveys the purpose. It could mention return value nature, but not required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, so schema coverage is 100% trivially. Baseline 4 is appropriate as the description adds no parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get custom domain configuration' clearly specifies the verb and resource, distinguishing it from sibling tools like org_customDomain_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_customDomain_updateC

Update custom domain configuration

ParametersJSON Schema
NameRequiredDescriptionDefault
customDomainNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. However, it only states 'Update' without explaining mutation implications, permissions, reversibility, or side effects. The description fails to compensate for the lack of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at just four words, but it lacks structure and depth. While it is not verbose, the extreme brevity sacrifices clarity and completeness. A bit more detail would be warranted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and a single undocumented parameter, the description is incomplete. It does not explain return values, side effects, or success criteria, leaving the agent with insufficient information to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter with 0% description coverage, and the description does not add any meaning beyond 'custom domain configuration'. It does not explain what 'customDomain' values are acceptable or how to use the null type. The description should clarify the parameter's role.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update custom domain configuration' clearly states the action (update) and the resource (custom domain configuration), and it distinguishes itself from the sibling tool 'org_customDomain_get' which retrieves the configuration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. There is no mention of prerequisites, conditions, or when not to use it, leaving the agent without context for appropriate invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_deleteB

Delete organization

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only says 'Delete', implying a destructive action, but does not explain consequences (e.g., irreversible, cascading deletes, required permissions). This is insufficient for a delete operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at two words. It is front-loaded and to the point, though it may be under-specified for a delete action. It earns its place with no extra content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool performs a notable delete operation with no parameters or output schema, the description is incomplete. It fails to clarify which organization is deleted, return behavior, or side effects. More detail is essential for safe usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, and schema description coverage is 100% (none to describe). Per guidelines, 0 parameters yields a baseline of 4. The description adds no parameter info, but none is needed as there are no parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete organization' clearly states the verb and resource, matching the tool name. It distinguishes from sibling tools as no other tool performs org deletion, but lacks specificity about which organization is targeted (e.g., current org vs. by ID), which is ambiguous given zero parameters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, or any prerequisites like admin rights. The description does not mention context or exclusions, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_getC

Get organization

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. Description implies read-only operation, but no confirmation of safety, auth requirements, or side effects. Does not disclose what data is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise at two words, but lacks critical details. Conciseness should not come at the expense of necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no parameters, the description fails to explain what the tool returns or how to interpret the result. Insufficient for a standalone tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has no parameters with 100% coverage, so description adds no value beyond schema. Baseline of 3 is appropriate as schema already fully documents parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states verb 'Get' and resource 'organization', which is clear but ambiguous given sibling tools like 'orgs_list' and 'org_customDomain_get'. Does not specify which organization (e.g., current context or by ID).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as 'orgs_list' or 'org_customDomain_get'. No context about prerequisites or scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

orgs_createC

Create a new workspace to collaborate with others on exams.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description must fully disclose behavioral traits. It only states the action 'Create' without mentioning side effects, permissions, return values, or idempotency. This leaves significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (one sentence) and front-loaded, but it omits valuable details that could make it more helpful without adding much length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and annotations, the description is insufficient. It does not explain what the tool returns or how to handle the response, which is critical for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% coverage, and the description does not explain the 'name' parameter (e.g., format, uniqueness, default behavior). It should compensate for the lack of schema descriptions but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Create', the resource 'new workspace', and the purpose 'to collaborate with others on exams'. It is easily distinguishable from sibling tools like orgs_list, org_get, org_update, and org_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, conflicts, or preferred use cases, leaving the agent without decision support.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

orgs_listB

List all organizations

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states 'List all organizations' but omits details like read-only nature, possible pagination, or rate limits. The behavioral insight is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no fluff. It is perfectly concise and front-loaded, earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of an output schema and zero parameters, the description adequately conveys the tool's purpose. However, it could be enhanced by explicitly stating the scope (e.g., 'all organizations' with no filtering) to fully contextualize its behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters (100% coverage), so the baseline is 3. The description adds no parameter-level information because none exist, but it does not detract. No extra semantics needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List all organizations' uses a specific verb 'list' and a clear resource 'organizations'. It distinguishes itself from siblings like org_get (single organization) and orgs_create (create), making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as org_get or orgs_create. The description implies usage for retrieving all organizations, but does not explicitly state when not to use it or offer comparisons to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_updateC

Update organization

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
logoImageUrlNo
menuColorNo
accentColorNo
sharingBlockPublicNo
settingsNo

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden of behavioral disclosure. 'Update organization' implies mutation, but no details about permissions, reversibility, side effects, or error conditions are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise but at the cost of under-specification. Two words do not provide sufficient value; the description should be longer to be useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested objects, no output schema, no annotations), the description is severely incomplete. It fails to explain behavior, return values, or usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description adds no information about parameters. The agent must guess the meaning of fields like menuColor, accentColor, settings, etc., without any textual hints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Update organization', which clearly indicates the verb (update) and resource (organization). However, it does not distinguish from sibling tools like 'org_delete' or 'orgs_create', nor does it specify which aspects of the organization can be updated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like 'org_get' or 'orgs_list'. No prerequisites, conditions, or exclusions are mentioned, leaving the agent without context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patchRubricsidC

Update an existing rubric in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo
scoringNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It states 'update' but does not disclose that it is a partial update (only provided fields are modified), what happens to omitted fields, required permissions, or side effects. This is insufficient for a mutation endpoint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at one sentence. It could be improved by adding a few words about partial update or constraints without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested scoring object), lack of annotations, and no output schema, the description is too minimal. It does not explain that only provided fields are updated, that the id must reference an existing rubric, or any error conditions. The agent lacks critical context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has detailed descriptions for nested fields (e.g., rubricType, criteria), but top-level parameters id and name lack descriptions. The description does not add any meaning beyond the schema, missing the opportunity to clarify the purpose of id or the partial update nature of the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update an existing rubric in the current workspace' clearly states the verb (update) and resource (rubric), distinguishing it from create (postRubrics), read (getRubrics), and delete (deleteRubricsid) siblings. However, it does not specify that it is a partial update (PATCH), which would strengthen clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. It implies usage for updating, but lacks when-not-to-use, prerequisites (e.g., must have an existing rubric ID), or comparisons with other tools. Context is solely implied by the verb 'Update'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

permissions_assignB

Create a new permission for a specific resource, or update the role of an existing actor.

ParametersJSON Schema
NameRequiredDescriptionDefault
resourceYesThe ID of the resource to assign the permission on, e.g. `exam_1234`.
bodyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
actorYes
roleYes
createdAtYes
updatedAtYes
createdByYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but only states mutation (create/update). It does not disclose side effects like invitation sending, idempotency, or required authentication.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence clearly conveys the action without unnecessary words. Well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and many siblings, the description is too brief. It omits return value, required permissions, and error conditions, leaving gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% and the description adds no parameter-specific details. The tool's description does not compensate for undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates or updates a permission for a specific resource, using verbs 'create' and 'update' with resource 'permission'. It distinguishes from sibling tools like permissions_list and permissions_delete by focusing on assignment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as permissions_getActorRole or permissions_autocomplete. No context about prerequisites or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

permissions_autocompleteA

Search for users and groups to share with. Returns matches by name or email. Respects restricted group visibility.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesSearch query (min 3 characters)
typeNoType of entity to search

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the burden. It discloses that results respect restricted group visibility, but lacks details on pagination, rate limits, or empty result handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with front-loaded purpose, no redundant or extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple autocomplete tool with no output schema, the description covers search purpose, matching criteria, and a visibility constraint. Minor gaps in result format but acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters with 100% coverage. The description adds value by clarifying that 'q' matches by name or email, which is not in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it searches for users and groups for sharing, specifying matching by name or email and respecting restricted group visibility. This distinguishes it from sibling permission tools like permissions_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for autocomplete search when sharing, but provides no explicit guidance on when to use this vs. other permissions tools, nor exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

permissions_deleteC

Remove a permission for a specific actor on a specific resource.

ParametersJSON Schema
NameRequiredDescriptionDefault
resourceYes
actorYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only says 'Remove', lacking details on irreversibility, required permissions, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (one sentence) but borderline under-specified. It could be improved by adding brief parameter hints without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of schema descriptions and output schema, the description is incomplete. It does not provide enough information for an agent to correctly invoke the tool (e.g., parameter formats, return type).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description should clarify parameter meanings. It mentions 'specific actor' and 'specific resource' but does not explain the pattern constraints (e.g., user_, group_, org_) or what resource refers to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the action (remove), the target (permission), and the key entities (actor and resource), distinguishing it from sibling tools like permissions_assign, permissions_list, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or alternative guidance is provided. The agent has to infer from the name alone that this tool is for deletion, with no prerequisites or exclusions mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

permissions_getActorRoleA

Get the role of a specific actor for a specific resource.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorYesThe ID of the actor to check role for, e.g. `user_abcd`.
resourceYesThe ID of the resource to check role for, e.g. `exam_1234`.

Output Schema

ParametersJSON Schema
NameRequiredDescription
actorYes
roleYes
createdAtYes
updatedAtYes
createdByYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears full burden for behavioral disclosure. It does not mention permission requirements, error cases, or what happens if the actor/resource is invalid. Only the basic action is stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single clear sentence with no unnecessary words. It is front-loaded and directly explains the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal but sufficient for a simple getter tool given that an output schema exists (which can explain return values). However, it does not add context about the role system or how this fits with other permission tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with clear parameter descriptions. The tool description adds no additional meaning beyond what the schema already provides, so baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('role of a specific actor for a specific resource'), clearly distinguishing it from sibling tools like permissions_assign or permissions_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. The name implies its purpose, but there is no mention of scenarios or exclusion of other tools like permissions_list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

permissions_listB

Get a list of users, groups and orgs that have access to a specific resource.

ParametersJSON Schema
NameRequiredDescriptionDefault
resourceYesThe ID of the resource to list permissions for, e.g. `exam_1234`.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. Description only says 'Get a list', implying read-only, but does not disclose any other behavioral traits such as required permissions, error handling, or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded, no filler words. Could be slightly more structured but is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description is too minimal. Lacks context on behavior for missing resources, pagination, or result format. With no annotations, more detail is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has one parameter with 100% coverage in description. Description adds no extra meaning beyond the schema example. Baseline 3 for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the verb 'Get' and the resource 'a list of users, groups and orgs that have access to a specific resource'. Distinguishes from sibling tools like permissions_assign and permissions_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage for listing permissions but provides no explicit when-to-use or when-not-to-use guidance, nor mentions alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postRubricsC

Add a rubric to the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
scoringNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must disclose behavioral traits. It only states 'Add', implying a write operation, but fails to mention idempotency, potential side effects (e.g., overwriting if name exists?), required permissions, rate limits, or the response format. The lack of detail leaves critical behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but inadequate. It front-loads no useful information beyond the basic action. Each word must earn its place, but here the brevity sacrifices essential context. An appropriate length would include at least a sentence about parameters and behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the nested 'scoring' parameter with multiple rubric types, the description provides no context on how to construct a rubric. The schema has some internal descriptions, but the tool description offers no high-level guidance. Sibling tools like getRubrics imply this is a create operation, but without further detail, the description is far from complete for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 2 complex parameters (name, scoring) with 0% schema description coverage (despite some nested descriptions in the schema itself). The tool description does not explain these parameters at all—e.g., what 'name' represents, or how to structure the 'scoring' object for different rubric types. This forces the agent to rely entirely on the schema, which is insufficient for constructing a valid request.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Add a rubric') and scope ('to the current workspace'). It is specific enough to distinguish from sibling tools like getRubrics, patchRubricsid, and deleteRubricsid, which handle other CRUD operations. However, it could be improved by clarifying the purpose of a rubric in the workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, such as postRubricsGenerate (which may create a rubric from a prompt) or patchRubricsid (for updates). There is no explanation of prerequisites, typical use cases, or indications that the tool is for creating new rubrics rather than editing existing ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postRubricsGenerateC

Start a background job to generate a rubric using AI.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageNo
sourceMaterialIdsNo
languageNo
contextNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that it is a background job, implying asynchronous execution, but does not elaborate on the job lifecycle, how to check status, or what the output contains. With no annotations, this minimal information leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but lacks necessary detail. It is not a model of efficient communication because it sacrifices information for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has four parameters and returns no output schema, yet the description does not explain what the background job returns (e.g., a job ID) or how to monitor its progress. This is insufficient for a complex background job tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not explain any of the four parameters (message, sourceMaterialIds, language, context). Since schema description coverage is 0%, the description must compensate but fails entirely to add meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it starts a background job to generate a rubric using AI. It specifies the verb 'start' and the resource 'background job for rubric generation', which distinguishes it from synchronous creation tools. However, it does not explicitly differentiate it from sibling tools like postRubrics or exams_questions_generate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention when a background job is preferred over synchronous generation, nor does it reference related tools such as jobs_get or jobs_cancel for tracking the job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_cancelGenerateTopicsB

Cancel an in-progress mastery topic generation job.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description implies it only applies to in-progress jobs, but no details on side effects, idempotency, permissions, or error states. Without annotations, more behavioral context would be helpful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence is concise and front-loaded. However, could include more detail without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a cancellation tool, description omits important details like idempotency, success/error responses, and what happens to related resources. With no output schema, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet description does not explain what the 'id' parameter represents (e.g., job ID, session ID). No added value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action ('cancel') and the resource ('in-progress mastery topic generation job'). It effectively distinguishes from sibling tools like practiceSpaces_generateTopics and practiceSpaces_getProgress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as exams_cancelGeneration or jobs_cancel. Lacks context on prerequisites or conditions for cancellation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_createC

Create a new practice space within your workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
sourceMaterialIdsNo
masteryTopicsNo
settingsNo
folderNo
permissionsNo
studentNotesNo
publishedNo
metadataNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should disclose behavioral traits. It only says 'create', implying a write operation, but no details on side effects, permissions, or return values.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it lacks necessary detail to be useful. It is not verbose, but it is under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (9 parameters, nested objects, no output schema, no annotations), the description is severely incomplete. It does not explain what a practice space is, how to configure it, or what the result is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no information about parameters. The input schema has 9 parameters with no descriptions, and the description does not compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a new practice space. It is specific with verb and resource, but does not differentiate from siblings like practiceSpaces_duplicate or practiceSpaces_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as practiceSpaces_duplicate or practiceSpaces_update. No context on prerequisites or conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_deleteC

Delete the practice space.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Lacks annotations. Description only states 'delete' without explaining side effects (e.g., cascading deletions of sessions/topics), permission requirements, or irreversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise but at the expense of informativeness. A single sentence is not enough to guide correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and a single parameter, the description should cover common concerns (e.g., error handling, prerequisites) but fails to do so.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no meaning to the 'id' parameter; the schema has 0% coverage with just a string type, leaving the agent to guess what value to provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Delete' and the resource 'practice space', which distinguishes it from sibling tools like practiceSpaces_create or practiceSpaces_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use or avoid this tool. No mention of prerequisites, side effects, or alternatives such as practiceSpaces_get for verification before deletion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_duplicateC

Duplicate the practice space's settings and topics to a new practice space.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description is the sole source for behavioral traits. It only states the action but fails to disclose side effects, permission requirements, or whether the operation is idempotent. No contradiction with annotations since none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no redundant words. However, it is too brief given the complexity and lacks parameter details, making it minimally acceptable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and incomplete parameter descriptions, the description leaves significant gaps. The agent cannot fully understand the tool's behavior, such as whether the original is modified or what the response contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not explain the parameters 'id' and 'name'. With 0% schema description coverage, the description adds no meaning beyond listing the parameters. The agent cannot deduce that 'id' is the source practice space and 'name' is for the new duplicate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the verb 'duplicate' and the resource 'practice space's settings and topics', with the outcome 'to a new practice space'. This clearly distinguishes it from siblings like practiceSpaces_create and practiceSpaces_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, such as practiceSpaces_create for creating an empty space or practiceSpaces_update for modifying an existing space. No exclusions or context provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_generateTopicFeedbackC

Get feedback for a specific topic in a practice space.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
topicIdNo
feedbackTypeNo

TDQS

C2.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only says 'Get feedback,' implying a read operation. It fails to disclose any behavioral traits such as authentication requirements, rate limits, or side effects. The description adds no value beyond the basic read assumption.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short (one sentence), which is concise, but it omits essential details. While brevity is good, the lack of substance hurts usability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has three parameters (one required, two optional with an enum) and no output schema, the description is insufficient. It does not mention that topicId or feedbackType are likely needed, nor does it describe the return value or behavior when parameters are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the three parameters (id, topicId, feedbackType). It does not clarify the meaning of feedbackType enum values 'progress' and 'challenges', nor the role of id and topicId.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get'), the resource ('feedback'), and the context ('for a specific topic in a practice space'). It sufficiently distinguishes from sibling tools like practiceSpaces_get (getting the space itself) and practiceSpaces_getProgress (overall progress).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, when not to use it, or how it relates to sibling tools like practiceSpaces_generateTopics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_generateTopicsC

Start generating mastery topics for a practice space.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description does not disclose key behavioral traits such as asynchronicity, job creation, required permissions, or side effects. The phrase 'start generating' implies initiation but lacks details on progress or cancellation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is concise but under-specified. It wastes no words but sacrifices informational value for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool likely initiates an asynchronous background process, the description omits critical context such as return value, progress tracking, or cancellation mechanism. This makes the tool poorly prepared for an agent to use effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description does not explain the meaning or format of the 'id' parameter. This leaves the agent without essential semantic context for parameter use.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('start generating') and the resource ('mastery topics for a practice space'), providing a specific verb and resource. However, it does not differentiate from the sibling tool 'practiceSpaces_cancelGenerateTopics' or indicate that the generation may be asynchronous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'practiceSpaces_cancelGenerateTopics' or other generation tools. There is no mention of prerequisites or context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_getC

Get a single practice space by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It only states 'Get,' implying read-only operation, but fails to disclose behaviors like permissions needed, error handling (e.g., when ID not found), or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words. While extremely short, it avoids unnecessary detail and is front-loaded with the essential action. It could arguably be more informative without sacrificing conciseness, hence a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and annotations, the description is too sparse. It does not explain what the tool returns, how errors are signaled, or any edge cases. A more complete description would add return value structure or common usage patterns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description merely says 'by its ID' without elaborating on the parameter's format, origin, or constraints. The agent gains no extra meaning beyond the schema's property name and type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Get a single practice space by its ID,' which clearly identifies the verb (get) and resource (practice space), and distinguishes it from siblings like practiceSpaces_list which retrieves multiple spaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as practiceSpaces_list or practiceSpaces_create. There is no mention of context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_getProgressC

Get the progress of a single practice space by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. The description only says 'Get', implying read-only, but does not mention permission requirements, rate limits, or side effects. This is minimal transparency for a simple retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently conveys the core purpose. Every word earns its place, with no wasted information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description is minimally adequate. However, it leaves gaps: it doesn't explain what 'progress' data looks like or whether any preconditions exist. The lack of output schema or further context makes it less complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter 'id' with no description, and the tool description does not add any meaning beyond the schema. Schema description coverage is 0%, but the description fails to explain what 'id' refers to (e.g., format, source). Thus, it adds no value to the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'progress of a single practice space' by ID, effectively distinguishing it from other practiceSpaces tools like practiceSpaces_get (which presumably gets the space itself) and practiceSpaces_list. However, it does not clarify what 'progress' specifically refers to (e.g., student progress, generation progress), leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as practiceSpaces_get or practiceSpaces_sessions_get. The description does not mention context, prerequisites, or exclusions, making it difficult for an agent to select the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_listA

Get a list of all practice spaces you have access to in this workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
folderNoOptional filter by folder ID, e.g. `folder_1234`.
roleNoOptional filter by permission role.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must disclose behavior. It states 'Get a list' indicating a read operation, but does not address edge cases, pagination, or side effects. Adequate for a simple list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no fluff. Front-loaded with purpose. Efficiently communicates essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, description could hint at return structure, but 'list' is sufficient. Parameters are optional, so using no filters returns all. Minor gap: not specifying whether results are sorted or limited.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers both parameters (folder and role) with descriptions. The tool description adds no extra meaning beyond the schema, placing it at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Get a list of all practice spaces you have access to in this workspace', specifying verb, resource, and scope. It distinguishes from sibling tools like practiceSpaces_get (single space) and practiceSpaces_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for use (listing accessible practice spaces) but does not explicitly mention when to avoid it or alternatives. Implicitly, filtering is supported via optional parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_questionsPreviewC

Get a preview of the top practice space questions by quality score.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must describe behavioral traits. It only says the tool returns a 'preview', implying it is read-only, but does not confirm this. It does not mention whether it returns limited data, how quality score is defined, or any side effects. The disclosure is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but it is too brief to be effective. Important details are omitted, making it under-specified rather than concise. A good concise description would front-load key information without sacrificing completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and parameter descriptions, the description is insufficient. For a tool that queries by ID, it should clarify what the preview contains (e.g., question text, score, number of results). The current text leaves too many gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter 'id' (string) with 0% description coverage. The description does not explain what 'id' refers to (e.g., practice space ID) or any constraints. It provides no value beyond the schema, which is already minimal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get a preview'), the resource ('practice space questions'), and a specific criterion ('by quality score'). It distinguishes from siblings like practiceSpaces_get (which probably gets the practice space itself) and practiceSpaces_list (lists practice spaces). However, the term 'preview' is vague and could be more specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like practiceSpaces_get or exams_questions_get. There is no mention of prerequisites, when-not-to-use, or intended workflow. With many sibling tools, this lack of context reduces utility.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_sessions_createC

Create a new practice space session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
promptNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description only indicates a mutation (create), but with no annotations, it fails to disclose behavioral traits like side effects, idempotency, authorization needs, or the result of the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—one sentence—but it effectively communicates the primary purpose without unnecessary words. However, it could add more detail without losing brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and only a minimal description, the tool's usage context is severely lacking. The agent cannot understand return values, error conditions, or how to construct valid requests.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 2 parameters with 0% description coverage, and the description does not mention any parameters or their meanings. The agent has no context for 'id' (likely a practice space ID) or optional 'prompt'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('create') and the resource ('new practice space session'), matching the tool's name. However, it does not explicitly distinguish from sibling tools like 'embedSessions_create' or explain what a session entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as 'practiceSpaces_sessions_get' or 'practiceSpaces_sessions_saveAnswer'. No context on prerequisites or typical use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_sessions_getC

Get a practice space session by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states the basic operation without disclosing behavioral traits such as read-only nature, authorization requirements, or error handling (e.g., what happens if the session does not exist). Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (one sentence) and front-loaded with the core purpose. However, it is overly brief and sacrifices critical details such as parameter roles and usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema and parameter descriptions, the tool description is incomplete. An agent may not know what to expect as a return value or how to correctly invoke the tool with the two ID parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has two parameters ('id' and 'sessionId') with no descriptions (0% coverage). The description mentions 'by its ID' but does not clarify which parameter is the session ID or how they relate. This could cause confusion for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the resource ('practice space session') with the method ('by its ID'). However, it does not explicitly differentiate from sibling tools like practiceSpaces_sessions_getMine, which might also retrieve sessions but without an explicit ID. The purpose is clear but lacks explicit distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention when not to use it or provide any context for selection among siblings like practiceSpaces_sessions_getMine or practiceSpaces_sessions_create.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_sessions_getMineC

Get the current user's past practice space session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only says 'get,' indicating a read operation. It does not disclose behavioral traits such as authentication requirements, response format, or error conditions beyond what is obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words, making it concise. However, it is too brief and lacks needed details, which is under-specification rather than efficient conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple parameter set (one required ID) and no output schema, the description should clarify the ID's purpose and what the tool returns. It fails to do so, leaving the agent with incomplete context for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema defines one required 'id' string parameter with 0% description coverage. The description does not explain what 'id' represents (e.g., session ID) or how it relates to the current user's past session, leaving the parameter's meaning ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool gets the current user's past practice space session, specifying the resource and scope. However, it does not explicitly distinguish from sibling tools like practiceSpaces_sessions_get, which likely retrieves any session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description implies it is for the current user's past sessions but lacks exclusions or context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_sessions_saveAnswerC

Save an answer for a specific question in a practice space session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
sessionIdYes
questionIdYes
questionNo
valueNo
completedNo
contextNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description implies mutation ('save') but provides no details on behavioral traits such as authentication needs, rate limits, side effects (e.g., overwriting existing answers), or what happens on error. With zero annotations, the description carries full burden but fails to disclose key behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but not sufficient. It is under-specified and essentially restates the tool's name. It lacks structure (e.g., not front-loaded with key info) and does not earn its place—there is room for more useful content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 parameters, no output schema, no annotations, many sibling tools), the description is highly incomplete. It does not explain return values, error conditions, parameter roles, or how to use the tool effectively. An agent would struggle to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description offers no explanation of the 7 parameters (id, sessionId, questionId, question, value, completed, context). For example, the patterns for 'question' (e.g., q_* ) are not mentioned, nor is the purpose of 'context' or 'completed'. The agent receives no semantic guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (save an answer) and the context (practice space session). It differentiates from sibling tools like practiceSpaces_sessions_create or exams_sessions_saveFeedback by specifying 'answer'. However, it adds little beyond the tool name, missing details like whether it creates or updates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., exams_sessions_editAnswer). No prerequisites, conditions, or exclusion criteria are provided. The agent must infer usage from context alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_students_getC

Get details about a specific practice space student.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
studentIdYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose any behavioral traits such as idempotency, side effects, or permissions. The agent cannot infer whether this is a safe read operation beyond the 'get' verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but lacks necessary detail. It under-specifies the tool's purpose and parameters, making it less helpful than it could be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and two undocumented parameters, the description is highly incomplete. It fails to provide sufficient context for an agent to correctly invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description does not explain what 'id' and 'studentId' represent. The agent has no semantic understanding of these required parameters, which is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a 'get' operation for a specific practice space student, with the verb 'Get' and the resource 'details about a specific practice space student'. It is specific and not a tautology, though 'details' is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus the sibling practiceSpaces_students_list, nor any context about prerequisites or alternatives. The description leaves the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_students_listC

Get a list of all students that practiced in a practice space.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should disclose behavioral traits beyond the basic operation. It only says 'Get a list', omitting details like read safety, rate limits, error handling, or pagination. This is minimal disclosure for a read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, very concise at 12 words. However, it sacrifices necessary detail like parameter explanation or output format. It is appropriately sized for a simple tool but lacks completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one required parameter and no output schema, the description should specify what the 'id' refers to, whether the list is paginated, and the format of returned data. It fails to provide these details, making it incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter 'id' with 0% description coverage. The description does not explain the parameter's meaning (e.g., that it represents the practice space ID) or provide any additional semantics beyond the schema's bare name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves a list of all students that practiced in a practice space, with a specific verb 'Get a list' and resource 'students'. It distinguishes from siblings like practiceSpaces_students_get (single student) and practiceSpaces_list (list spaces).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool vs alternatives, nor any prerequisites or exclusions. It simply states what it does without context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

practiceSpaces_updateC

Update an existing practice space.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo
settingsNo
folderNo
sourceMaterialIdsNo
masteryTopicsNo
studentNotesNo
publishedNo
metadataNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose any behavioral traits beyond 'update,' such as whether the operation is idempotent, what happens to existing values if omitted, or any side effects like requiring specific permissions. Since annotations are absent, the description carries the full burden, and it fails to provide meaningful transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (one sentence) but lacks structure. It does not front-load key information or use formatting like bullet points or sections. While brevity is positive, the description sacrifices substance for brevity, making it minimally useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the input schema (9 parameters, nested objects, enums) and no output schema, the description is critically incomplete. It does not explain the purpose of parameters like 'settings', 'folder', or 'published', nor does it describe return values or error conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention any of the 9 parameters, and the input schema lacks descriptions (0% coverage). The schema is complex with nested objects, but the description adds no semantic meaning to help the agent understand what each field does or how to use them correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Update') and the resource ('an existing practice space'), which distinguishes it from sibling tools like practiceSpaces_create or practiceSpaces_delete. However, it lacks additional context such as the specific fields that can be updated, which would enhance clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as practiceSpaces_create for new spaces or practiceSpaces_duplicate for copying. There are no prerequisites, examples, or context about typical use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionBank_createC

Add a question to the question bank in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description bears full responsibility. It only states 'add' without disclosing side effects, authorization requirements, or any constraints. No information about idempotency or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it lacks essential structure. For a tool with a complex input schema, more detail is warranted; the brevity here is under-specification rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the input (nested object with required fields) and no output schema, the description is severely incomplete. It omits required parameters, return value, and any behavioral context, leaving the agent with insufficient information to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 0%, so the description must compensate, but it provides no parameter details. The single parameter 'question' is a complex nested object; the description gives no hint about its structure or required fields like 'type' and 'settings'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('add') and the resource ('question to the question bank'), with a scope ('in the current workspace'). It distinguishes from sibling tools like questionBank_delete, questionBank_list, and questionBank_update by focusing on creation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description is too brief to provide usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionBank_deleteC

Delete a question bank item from the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must convey behavioral traits. It states 'delete' implying destruction, but lacks details on permanence, consequences, or required permissions. This is insufficient for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but overly brief. It sacrifices necessary parameter explanation, making it less helpful than it could be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description could be adequate if it explained the id. However, it leaves key information missing, so completeness is moderate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter 'id' with 0% description coverage. The description does not explain what 'id' represents, how to obtain it, or its format, leaving the agent with no guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Delete' and the resource 'question bank item', and specifies scope 'from the current workspace'. It distinguishes itself from sibling tools like questionBank_create, questionBank_list, and questionBank_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidelines are provided. The description does not indicate when to use this tool versus alternatives, nor does it mention prerequisites or scenarios to avoid.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionBank_listA

Get a list of question bank items that exist within the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, description carries full burden. It indicates a read operation ('Get a list') but provides no additional behavioral details such as pagination, rate limits, or authentication beyond workspace context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no fluff, front-loaded purpose. Every word is necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Minimal but sufficient for a simple list tool with no parameters and no output schema. However, lacks information about return format or any implicit filters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has no parameters, and schema coverage is 100%. The description implies workspace scoping, adding value beyond schema. Baseline for 0 params is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'Get', resource 'list of question bank items', and scope 'current workspace'. It distinguishes from sibling tools like questionBank_create or questionBank_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., exams_list, attributes_list). Does not mention filtering, pagination, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionBank_updateC

Update an existing question bank item in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
questionNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must convey behavioral traits. It only says 'update' which implies mutation but does not disclose whether the update is full replacement or partial, success/failure conditions, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that is front-loaded and contains no extraneous information. Every word serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested schema and lack of output schema, the description is insufficient. It does not specify update semantics (partial vs full), error handling, or return values, leaving the agent with significant ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% at the top level. The description adds no value beyond what is in the schema; it does not explain the 'id' or 'question' parameter semantics, nor does it clarify how to structure complex fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (update) and the resource (existing question bank item) and implies the scope (current workspace). It distinguishes from sibling tools like questionBank_create and questionBank_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, no prerequisites, and no context for appropriate usage. It simply states what it does.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_deleteA

Deletes a specific question type by its ID. Note that only the owner organization of the question type can delete it, and only if it has not been used in an exam.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description discloses key behavioral constraints (ownership and usage condition). However, it does not mention deletion reversibility, error scenarios, or response format, leaving gaps for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, each providing necessary information without redundancy. The action is stated first, followed by conditions in a clear note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deletion tool with conditions and no output schema or annotations, the description is incomplete. It lacks information on success/error responses, side effects, and any prerequisites like authentication, leaving the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description only says 'by its ID' for the 'id' parameter, adding minimal detail beyond the schema. It does not explain the format, source, or validation of the ID.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it deletes a question type by ID, using specific verb 'Deletes' and resource 'question type'. This clearly distinguishes it from sibling tools like questionTypes_enable, questionTypes_get, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear conditions for use: only owner organization can delete, and only if unused in exams. This gives important context on when the tool is applicable, though no explicit alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_disableB

Disable a question type for the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must cover behavioral traits. It fails to mention reversibility, effects on existing questions, permissions, or return value, leaving the agent uninformed about the tool's impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no fluff. However, it could include more context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description is bare minimum. It lacks details on prerequisites, return behavior, and side effects, making it incomplete for an agent to use confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%; the description does not explain the 'id' parameter. It adds no meaning beyond the schema, so the parameter's role is unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action (disable), resource (question type), and scope (current workspace). It distinguishes well from sibling tools like questionTypes_enable and questionTypes_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description implies usage when disabling a question type but lacks explicit guidance on when to use this tool versus enable or delete, and no when-not-to-use conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_enableC

Enable a question type for the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It only states 'enable' without disclosing what enabling entails (e.g., permissions required, reversibility, side effects). No behavioral details beyond the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no redundancy, but it is under-specified. It earns its brevity but at the cost of missing critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter descriptions, the description is insufficient. It does not explain the effect of enabling, return values, or how the 'id' parameter is used, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention the 'id' parameter at all. It adds no meaning beyond the schema, leaving the agent to guess what 'id' refers to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (enable) and resource (question type), and specifies the scope ('current workspace'). It is a specific verb-resource combination that is unambiguous, though it does not differentiate from sibling tools like questionTypes_disable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives (e.g., questionTypes_disable), no prerequisites, and no context for when enabling is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_getB

Retrieves a specific question type by its ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It only states 'Retrieves', offering no details about error responses, partial data, or whether the resource must exist. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (one sentence) and front-loaded, but it sacrifices useful detail. Could be improved with a few more words without harming conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple retrieval with one parameter and no output schema, the description is somewhat complete. However, missing error behavior and resource semantics (e.g., what constitutes a question type) limit completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no extra meaning to the 'id' parameter, such as expected format or example values. The parameter is self-explanatory but leaves room for ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (retrieves) and the resource (specific question type by ID). It effectively distinguishes from sibling tools that list, delete, or modify question types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like questionTypes_list or questionTypes_getQti3Pci. The purpose is implied but not differentiated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_getQti3PciC

Get JS module for a question type PCI for QTI 3 export.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits like read-only nature, authentication requirements, or failure modes. The description only restates the purpose without additional context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no unnecessary words. However, it could be expanded to include additional useful information without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, the description should hint at the return format or content. It mentions 'JS module' but lacks details on encoding, structure, or usage of the returned data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% coverage, and the description does not explain the meaning of the single parameter 'id'. The agent cannot infer what ID to provide (e.g., question type ID, PCI ID, module ID) from the description alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves a JS module for a question type PCI specifically for QTI 3 export. It uses a specific verb ('Get') and resource, and the context distinguishes it from generic 'questionTypes_get' and other sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as 'questionTypes_get'. The description does not mention any prerequisites, limitations, or intended scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_listA

Lists all question types, either those available by default or those created within the user's organization.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must handle behavioral transparency. It describes the basic listing function but omits potential details like pagination, ordering, or permission requirements. For a simple list with no parameters, this is acceptable but still lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is a single, clear sentence that front-loads the action. No unnecessary words, effectively concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool is simple with no parameters and no output schema. Description adequately covers the core function. However, it could mention if the list is paginated or if there are size limits, which would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has zero parameters with 100% coverage. Description does not add parameter semantics, but baseline for zero parameters is 4. No additional information needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool lists question types and distinguishes between default and organization-created ones. However, it does not explicitly differentiate from sibling tools like questionTypes_listPublic or questionTypes_get, which may serve similar purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description provides no guidance on when to use this tool versus alternatives, such as when to use questionTypes_listPublic instead. No context or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_listPublicA

Lists all public question types, which can be enabled in workspaces.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. Description implies a read-only list operation, but provides no details on pagination, limits, or possible behavioral traits beyond the basic listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, direct, and no wasted words. Ideal conciseness for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a parameterless list tool. Missing output schema description, but the purpose is clear. Slightly incomplete as it doesn't hint at response format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters in schema, so description need not explain any. Description does not mislead and adds no unnecessary info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool lists public question types and mentions their workspace enablement. Differentiates from siblings like 'questionTypes_list' or 'questionTypes_get' by specifying 'public', but could be more explicit about the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., 'questionTypes_list' for all types). Does not specify if there are prerequisites or context where this is preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionTypes_upsertA

Creates a new question type or updates an existing one. If the question type already exists, it will be updated; otherwise, a new one will be created.

ParametersJSON Schema
NameRequiredDescriptionDefault
$schemaNo
idNoUnique identifier for the question type (e.g. 'my-org.color-picker')
nameNoDisplay name for the question type, can be a simple string or translation object
descriptionNoDescription of the question type, can be a simple string or translation object
iconNoPath to the icon for this question type
shortcutNoKeyboard shortcut for quick access to this question type
generationNoOptions for AI question generation
gradingNoOptions for AI question grading
scanningNoOptions for AI question scanning
timeEstimateMinutesNoEstimated time in minutes to complete this question type
isAiNoWhether this question type uses AI functionality
backgroundColorNoBackground color for the question type
enforcePositionNoPosition enforcement for this question type
enforceTitleNoEnforced title for this question type
hideSettingsNoWhether to hide settings for this question type
isPreviewRefreshableNoWhether the preview can be refreshed for this question type
hasSimpleScoringNoWhether this question type has simple scoring
titlePlaceholderNoPlaceholder text for the title field
descriptionPlaceholderNoPlaceholder text for the description field
untitledPlaceholderNoPlaceholder text for untitled questions
componentsNoPaths to the components used by this question type
settingsNoConfiguration settings for the question type
publicNoWhether this question type is public and can be used by all users
translationsNoCustom translations for this question type
capabilitiesNoCapabilities of the question type for various operations
exportNoExport configuration for various formats
metadataNoOptional metadata for the question type, only used by external applications
stylesheetNo
indexNo
isDefaultNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the key behavioral trait: it creates or updates based on whether the question type already exists (implied by the 'id' field). However, it omits other critical details like return value, error conditions, or authorization requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, clear sentences. No filler or redundant information. Efficiently communicates the core function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 30 parameters, no output schema, and no annotations, the description is far too sparse. It does not explain what the tool returns, which fields are required for create vs. update, or how to verify the operation's success. For such a complex tool, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (87%), so the schema already documents most parameters. The description adds minimal value beyond the schema, only clarifying the existence check (which maps to the 'id' field). No new semantic details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates or updates a question type, using a standard upsert pattern. It distinguishes from siblings like 'questionTypes_delete' or 'questionTypes_get' by specifying the combined create/update behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly indicates use when you want to create or update a question type based on existence, but provides no explicit guidance on when to prefer this over separate create/update tools (which are not present as siblings). No alternatives or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sourceMaterials_createC

Add a source material and start processing it for later use in an exam.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
typeNo
nameNo
sizeNo
contentTypeNo
expiresNo
externalIdNoAn optional external identifier for the source material.

Output Schema

ParametersJSON Schema
NameRequiredDescription
orgYes
idYes
typeNo
createdByYes
createdAtNo
updatedAtNo
deletedAtNo
nameNo
originalSourceYes
convertedSourceV2No
geminiSourceNo
externalIdNo
parentSourceMaterialIdNo
cacheNameNo
summaryNo
factsCountNo
processingStatusNo
numberOfPagesNo
chapterMarkersNo
topicsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden of behavioral disclosure. It implies creation and processing, but does not explain side effects, required permissions, processing time, or error conditions. This under-disclosure leaves the agent uninformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it sacrifices necessary detail. It is not overly verbose, but the conciseness comes at the cost of completeness, earning a middling score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 7 parameters with minimal schema descriptions, no annotations, and a complex topic (source materials and processing), the description is far too brief. It does not explain what a source material is, what processing entails, or what the output contains, leaving the agent with insufficient context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (1 out of 7 parameters have a description). The description does not add any meaning to the parameters beyond the schema. It fails to compensate for the low coverage, leaving most parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (add) and resource (source material), and mentions 'start processing'. This distinguishes it from sibling tools like sourceMaterials_delete, sourceMaterials_list, etc. However, 'start processing' is somewhat vague but still acceptable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. There is no mention of prerequisites, conditions, or when not to use it. The description is a single sentence with no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sourceMaterials_deleteB

Deletes the specified source material.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only states 'deletes,' which implies destruction. It does not disclose whether the operation is reversible, what happens to associated data, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no extraneous information, and the essential action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (one parameter, no output schema, no annotations), the description is too sparse. It lacks details on how to obtain the 'id' and what the outcome of deletion entails, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, and the description does not explain the purpose or format of the 'id' parameter beyond stating it identifies the source material to delete. This adds no value over the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Deletes' and the resource 'the specified source material,' effectively distinguishing it from sibling tools like sourceMaterials_create and sourceMaterials_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for deleting source materials but provides no guidance on when to use it versus alternatives or any prerequisites. There is no mention of when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sourceMaterials_getB

Returns the current status and facts extracted from the specified source material.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
orgYes
idYes
typeNo
createdByYes
createdAtNo
updatedAtNo
deletedAtNo
nameNo
originalSourceYes
convertedSourceV2No
geminiSourceNo
externalIdNo
parentSourceMaterialIdNo
cacheNameNo
summaryNo
factsCountNo
processingStatusNo
numberOfPagesNo
chapterMarkersNo
topicsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates a read operation, but without annotations, it does not explicitly state safety, idempotency, or potential side effects. The terms 'current status and facts' are vague but imply retrieval.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, straightforward sentence of 12 words with no fluff. It is optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (1 parameter, simple get) and existence of an output schema, the description covers the essential purpose. It could be slightly more complete by stating that it retrieves by ID.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the tool description does not explain the single 'id' parameter beyond context. It does not specify the format or meaning of the ID, leaving the agent to infer.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns status and facts for a specified source material, using a clear verb 'Returns'. It distinguishes from sibling tools like create, delete, list, slice, and update, but could be more specific by mentioning retrieval by ID.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as sourceMaterials_list or sourceMaterials_get. There is no mention of prerequisites or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sourceMaterials_listB

Returns a list of source materials for the current organization.

ParametersJSON Schema
NameRequiredDescriptionDefault
externalIdNoFilter by externalId

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears full responsibility for behavioral disclosure. It only states a read operation, but omits important traits such as pagination behavior, response size limits, authorization requirements, or the fact that results are scoped to the current organization (already implied by the description).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the core purpose. Every word is necessary and there is no extraneous text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and the presence of an output schema, the description is minimally adequate. However, it could be improved by noting that the tool supports optional filtering via the externalId parameter and that it returns all source materials if no filter is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (the only parameter 'externalId' has a description). The tool description adds no additional meaning beyond the schema. According to guidelines, baseline is 3, and no adjustment is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Returns'), the resource ('a list of source materials'), and the scope ('for the current organization'). It effectively distinguishes itself from sibling tools like sourceMaterials_create and sourceMaterials_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any prerequisites, exclusions, or contextual usage rules. Among many list tools, this omission forces the agent to rely solely on the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sourceMaterials_sliceA

Create a new source material based on a specific page range of an existing source material.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
startPageNoStarting page number (1-based)
endPageNoEnding page number, inclusive (1-based)
nameNoOptional name for the new source material

Output Schema

ParametersJSON Schema
NameRequiredDescription
orgYes
idYes
typeNo
createdByYes
createdAtNo
updatedAtNo
deletedAtNo
nameNo
originalSourceYes
convertedSourceV2No
geminiSourceNo
externalIdNo
parentSourceMaterialIdNo
cacheNameNo
summaryNo
factsCountNo
processingStatusNo
numberOfPagesNo
chapterMarkersNo
topicsNo

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only states the basic action ('create') without disclosing side effects, authorization requirements, or whether the original is modified. This is insufficient for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-front-loaded sentence with no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the output schema may cover return values, the description lacks details on prerequisites (e.g., the existing material must exist), error conditions, and exact behavior (e.g., whether the page range is inclusive).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 75% schema coverage, the description adds no value beyond the schema. It doesn't explain parameter relationships (e.g., which id to use) or page numbering context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a new source material from a page range of an existing one, distinguishing it from sibling creation tools like sourceMaterials_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use (when needing a page-range slice) but provides no explicit guidance on when not to use or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sourceMaterials_updateC

Updates the specified source material.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNoThe new name of the source material.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It only says 'updates' without disclosing idempotency, return value, side effects, or required permissions. Lacks transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence is concise but overly terse. It provides minimal information; a slightly expanded description could maintain conciseness while adding clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of sibling tools for source materials, no output schema, and only partial parameter documentation, the description is incomplete. It does not specify what the tool returns or that it updates only the name field.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (name has description, id lacks it). The description adds no meaning beyond the schema: it restates 'updates' without elaborating on the id parameter or what the update does (e.g., partial vs full replacement).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Updates) and the resource (specified source material). It distinguishes from sibling tools like create, delete, get, list, slice. However, it could be more specific by indicating which fields are updatable (only name shown in schema).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like sourceMaterials_create or sourceMaterials_delete. The description does not mention context or prerequisites for updating.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

studentLevels_listA

Get a list of available default student levels.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description does not disclose any behavioral traits such as authentication requirements, rate limits, or read-only nature. Only a basic action is stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no unnecessary words, effectively conveying the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description is adequate for a simple list, it lacks details about return format, pagination, or any filtering capabilities, which are not provided by an output schema or annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so the description cannot add meaning beyond the schema. With 100% schema coverage trivially, the baseline of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get a list') and the specific resource ('available default student levels'), making the purpose unambiguous and distinct from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. However, the context of listing default levels is self-evident for a typical CRUD operation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

taxonomies_createC

Save a new taxonomy in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
descriptionNo
instructionsNo
levelsNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description only notes it's a save (create) operation. Lacks details on permissions, idempotency, error handling, or whether it returns the created object.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence of 9 words is under-specified, not concise. It fails to add value beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Description is extremely incomplete given 4 optional parameters and no output schema. Does not explain taxonomy structure, level semantics, or expected behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. Description does not mention any of the 4 parameters (name, description, instructions, levels), leaving the agent without guidance on how to populate them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Save' and the resource 'taxonomy' with scope 'current workspace'. It distinguishes from sibling tools like taxonomies_delete, taxonomies_list, and taxonomies_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. Does not mention prerequisites, conditions, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

taxonomies_deleteA

Delete a taxonomy from the current workspace. This fails if there are any exams associated with the taxonomy.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals a key constraint (failure if exams exist) but does not mention other aspects like destructiveness (implied), permission requirements, or the outcome of a successful deletion. This is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded with the action and immediately followed by a critical condition. Every sentence adds value without redundancy, making it efficient for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema, no annotations), the description covers the basic purpose and key behavior. However, it does not indicate what the tool returns on success (e.g., confirmation message) or mention permissions. It is mostly complete for a straightforward delete operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter 'id' with no description (0% coverage). The tool description does not explain what 'id' refers to (e.g., taxonomy ID). Since schema coverage is low, the description should compensate but fails to add meaning, leaving the agent to infer the parameter's role from the tool name only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Delete a taxonomy from the current workspace') and includes a specific failure condition ('This fails if there are any exams associated with the taxonomy.'). This distinguishes it from sibling tools like taxonomies_create, taxonomies_list, and taxonomies_update by specifying the operation and a unique constraint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly tells when to use the tool (to delete a taxonomy without associated exams) and when it will fail (if exams are associated). However, it does not explicitly provide alternatives or state when not to use it, such as suggesting to delete exams first. The context is clear but lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

taxonomies_listA

Get a list available taxonomies, including defaults and taxonomies created in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool returns a list including defaults and workspace taxonomies, but lacks details on ordering, pagination, or potential size limits. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence, front-loaded with the verb and resource. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description provides the general purpose but does not describe the structure of the returned list (e.g., fields, sorting). For a simple list with no parameters, it is minimally complete but could be more explicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so schema coverage is 100%. The description adds value by specifying the scope of the returned list (defaults + workspace). Baseline 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (get list), resource (taxonomies), and scope (defaults and workspace-created). It distinguishes from sibling tools like taxonomies_create, taxonomies_delete, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a list of taxonomies is needed, but does not provide explicit guidance on when not to use or mention alternatives. For a simple list tool, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

taxonomies_updateC

Update an existing taxonomy in the current workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
nameNo
descriptionNo
instructionsNo
levelsNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It only states 'update an existing taxonomy' without mentioning permissions, error handling, side effects, or whether it performs a partial or full update. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (7 words) but lacks necessary information. It is under-specified rather than concise; every sentence should earn its place, and this single sentence does not provide enough value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, no output schema, no annotations), the description is severely incomplete. It fails to explain return values, parameter semantics, or error behavior, leaving the agent with inadequate information to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the tool description adds no meaning to any of the 5 parameters. It does not explain id, name, description, instructions, or levels. The agent receives no guidance beyond the parameter names and types in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Update' and the resource 'existing taxonomy', with scope 'current workspace'. It is specific and distinguishes from sibling tools like taxonomies_create, taxonomies_delete, and taxonomies_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for modifying an existing taxonomy but provides no explicit guidance on when to use it vs alternatives, no prerequisites, and no exclusion criteria. The context is somewhat clear but lacks direct comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

users_createC

Create a new user in the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNo
nameNo
roleNoteacher
sendInviteNoWhether to send an invitation email. Defaults to false.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fails to disclose behavioral traits such as required permissions, idempotency, side effects, or error conditions. The minimal description does not compensate for the lack of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence, which is appropriate for a simple action, but it sacrifices completeness. It could include more context without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four parameters, no output schema, and no annotations, the description is insufficient. It lacks information about return values, required fields (though none are required), and common usage scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only sendInvite has a description). The description adds no meaning beyond the schema; it does not explain any parameter roles or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Create' and the resource 'new user' in the workspace, and it distinguishes from sibling tools like users_update, users_delete, and users_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description is too brief to offer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

users_deleteC

Delete a user from the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description lacks behavioral details beyond stating the delete operation. With no annotations, it should disclose consequences (e.g., irreversibility, cascading effects) but does not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, front-loading the purpose without unnecessary words, though it could include more detail without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the description is too brief; it omits important context like permanence, confirmation requirements, or return values, which are critical for a delete operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no meaning to the single 'id' parameter, and schema description coverage is 0%, so the parameter remains undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the action ('Delete') and the object ('a user from the workspace'), clearly distinguishing it from siblings like users_create, users_list, and users_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor any prerequisites or when-not-to-use scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

users_listA

Get a list of all users in the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavioral traits. It only states 'Get a list' without mentioning potential limitations such as pagination, filtering, or whether it returns all users. This lack of detail reduces transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words. It is appropriately sized and directly conveys the tool's function without any extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema), the description provides the essential context for a straightforward list operation. It could mention potential constraints like active users, but it is complete enough for its minimal complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema coverage is effectively 100%. The description does not need to add parameter details, and the baseline for zero parameters is 4. The description provides clear purpose, which is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get a list of all users in the workspace' uses a specific verb and resource, clearly indicating the tool's purpose. It distinguishes itself from sibling tools like users_create, users_delete, and users_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives. While the purpose is clear, there is no mention of when not to use it or any alternative tools for similar tasks, leaving some ambiguity for an AI agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

users_updateD

Update a user's role within the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits but fails to do so. It claims to update a role without specifying the role value input, side effects, or authentication needs. The description contradicts the schema, making behavior unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but it is underspecified and misleading. Conciseness without accuracy is not effective; it fails to earn its place due to incorrect content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one required parameter and no output schema or annotations, the description is critically incomplete. It omits what exactly is updated (role? but no role parameter), the expected return, and how it differs from other tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description does not clarify what the 'id' parameter refers to or how the role is specified. The description mentions a role but the schema lacks a role parameter, adding no semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'Update a user's role within the workspace', but the input schema only includes an 'id' parameter with no 'role' field. This is misleading and contradictory, as the tool's purpose as described does not match its actual parameters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus siblings like users_create, users_delete, or permissions_assign. The description does not mention any context, prerequisites, or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 177 tool updatesv0.0.3
    • Addedattributes_create
    • Addedattributes_delete
    • Addedattributes_list
    • Addedattributes_reorder
    • Addedattributes_update
    • RemoveddeleteExamsid
    • RemoveddeleteExamsidSessionssessionId
    • RemoveddeleteFoldersid
    • RemoveddeleteOrg
    • RemoveddeleteQuestion_bankid
    • RemoveddeleteQuestion_typesid
    • AddeddeleteRubricsid
    • RemoveddeleteSource_materialsid
    • RemoveddeleteTaxonomiesid
    • RemoveddeleteUsersid
    • AddedembedSessions_create
    • AddedembedSessions_get
    • AddedembedSessions_revoke
    • Addedexams_cancelGeneration
    • Addedexams_create
    • Addedexams_delete
    • Addedexams_duplicate
    • Addedexams_export_qti21Zip
    • Addedexams_export_qti3Zip
    • Addedexams_get
    • Addedexams_getContextSuggestions
    • Addedexams_import
    • Addedexams_list
    • Addedexams_print
    • Addedexams_questions_generate
    • Addedexams_questions_get
    • Addedexams_questions_import
    • Addedexams_questions_regenerate
    • Addedexams_sessions_acceptAllSuggestions
    • Addedexams_sessions_acceptSuggestion
    • Addedexams_sessions_delete
    • Addedexams_sessions_editAnswer
    • Addedexams_sessions_generateOverallFeedback
    • Addedexams_sessions_get
    • Addedexams_sessions_import
    • Addedexams_sessions_list
    • Addedexams_sessions_saveFeedback
    • Addedexams_sessions_saveOverallFeedback
    • Addedexams_sessions_scan
    • Addedexams_sessions_scanDocument
    • Addedexams_startGeneration
    • Addedexams_update
    • Addedfolders_create
    • Addedfolders_delete
    • Addedfolders_list
    • Addedfolders_update
    • RemovedgetExams
    • RemovedgetExamsid
    • RemovedgetExamsidContext_suggestions
    • RemovedgetExamsidSessions
    • RemovedgetExamsidSessionssessionId
    • RemovedgetFolders
    • RemovedgetMe
    • RemovedgetMediaUpload
    • RemovedgetOrg
    • RemovedgetOrgs
    • RemovedgetQuestion_bank
    • RemovedgetQuestion_types
    • RemovedgetQuestion_typesid
    • RemovedgetQuestion_typesPublic
    • AddedgetRubrics
    • RemovedgetSource_materials
    • RemovedgetSource_materialsid
    • RemovedgetStudent_levels
    • RemovedgetTaxonomies
    • RemovedgetUsers
    • Addedgroups_addMember
    • Addedgroups_create
    • Addedgroups_delete
    • Addedgroups_list
    • Addedgroups_listMembers
    • Addedgroups_removeMember
    • Addedgroups_update
    • Addedgroups_updateMember
    • Addedjobs_cancel
    • Addedjobs_get
    • Addedme_get
    • Addedme_update
    • Addedmedia_upload
    • Addedorg_customDomain_get
    • Addedorg_customDomain_update
    • Addedorg_delete
    • Addedorg_get
    • Addedorg_update
    • Addedorgs_create
    • Addedorgs_list
    • RemovedpatchMe
    • RemovedpatchOrg
    • RemovedpatchQuestion_bankid
    • AddedpatchRubricsid
    • RemovedpatchSource_materialsid
    • RemovedpatchUsersid
    • Addedpermissions_assign
    • Addedpermissions_autocomplete
    • Addedpermissions_delete
    • Addedpermissions_getActorRole
    • Addedpermissions_list
    • RemovedpostExams
    • RemovedpostExamsid
    • RemovedpostExamsidDuplicate
    • RemovedpostExamsidGenerate
    • RemovedpostExamsidGenerateCancel
    • RemovedpostExamsidPrint
    • RemovedpostExamsidQuestionsGenerate
    • RemovedpostExamsidQuestionsImport
    • RemovedpostExamsidQuestionsquestionIdGenerate
    • RemovedpostExamsidSessions
    • RemovedpostExamsidSessionsScan
    • RemovedpostExamsidSessionssessionIdFeedback
    • RemovedpostExamsidSessionssessionIdGenerate_overall_feedback
    • RemovedpostExamsidSessionssessionIdOverall_feedback
    • RemovedpostExamsidSessionssessionIdSuggestionsAccept_all
    • RemovedpostExamsidSessionssessionIdSuggestionsquestionIdAccept
    • RemovedpostExamsImport
    • RemovedpostFolders
    • RemovedpostFoldersid
    • RemovedpostOrgs
    • RemovedpostQuestion_bank
    • RemovedpostQuestion_types
    • RemovedpostQuestion_typesidDisable
    • RemovedpostQuestion_typesidEnable
    • AddedpostRubrics
    • AddedpostRubricsGenerate
    • RemovedpostSource_materials
    • RemovedpostSource_materialsidSlice
    • RemovedpostTaxonomies
    • RemovedpostTaxonomiesid
    • RemovedpostUsers
    • AddedpracticeSpaces_cancelGenerateTopics
    • AddedpracticeSpaces_create
    • AddedpracticeSpaces_delete
    • AddedpracticeSpaces_duplicate
    • AddedpracticeSpaces_generateTopicFeedback
    • AddedpracticeSpaces_generateTopics
    • AddedpracticeSpaces_get
    • AddedpracticeSpaces_getProgress
    • AddedpracticeSpaces_list
    • AddedpracticeSpaces_questionsPreview
    • AddedpracticeSpaces_sessions_create
    • AddedpracticeSpaces_sessions_get
    • AddedpracticeSpaces_sessions_getMine
    • AddedpracticeSpaces_sessions_saveAnswer
    • AddedpracticeSpaces_students_get
    • AddedpracticeSpaces_students_list
    • AddedpracticeSpaces_update
    • AddedquestionBank_create
    • AddedquestionBank_delete
    • AddedquestionBank_list
    • AddedquestionBank_update
    • AddedquestionTypes_delete
    • AddedquestionTypes_disable
    • AddedquestionTypes_enable
    • AddedquestionTypes_get
    • AddedquestionTypes_getQti3Pci
    • AddedquestionTypes_list
    • AddedquestionTypes_listPublic
    • AddedquestionTypes_upsert
    • AddedsourceMaterials_create
    • AddedsourceMaterials_delete
    • AddedsourceMaterials_get
    • AddedsourceMaterials_list
    • AddedsourceMaterials_slice
    • AddedsourceMaterials_update
    • AddedstudentLevels_list
    • Addedtaxonomies_create
    • Addedtaxonomies_delete
    • Addedtaxonomies_list
    • Addedtaxonomies_update
    • Addedusers_create
    • Addedusers_delete
    • Addedusers_list
    • Addedusers_update
  2. 62 tool updates
    • First observeddeleteExamsid
    • First observeddeleteExamsidSessionssessionId
    • First observeddeleteFoldersid
    • First observeddeleteOrg
    • First observeddeleteQuestion_bankid
    • First observeddeleteQuestion_typesid
    • First observeddeleteSource_materialsid
    • First observeddeleteTaxonomiesid
    • First observeddeleteUsersid
    • First observedgetExams
    • First observedgetExamsid
    • First observedgetExamsidContext_suggestions
    • First observedgetExamsidSessions
    • First observedgetExamsidSessionssessionId
    • First observedgetFolders
    • First observedgetMe
    • First observedgetMediaUpload
    • First observedgetOrg
    • First observedgetOrgs
    • First observedgetQuestion_bank
    • First observedgetQuestion_types
    • First observedgetQuestion_typesid
    • First observedgetQuestion_typesPublic
    • First observedgetSource_materials
    • First observedgetSource_materialsid
    • First observedgetStudent_levels
    • First observedgetTaxonomies
    • First observedgetUsers
    • First observedpatchMe
    • First observedpatchOrg
    • First observedpatchQuestion_bankid
    • First observedpatchSource_materialsid
    • First observedpatchUsersid
    • First observedpostExams
    • First observedpostExamsid
    • First observedpostExamsidDuplicate
    • First observedpostExamsidGenerate
    • First observedpostExamsidGenerateCancel
    • First observedpostExamsidPrint
    • First observedpostExamsidQuestionsGenerate
    • First observedpostExamsidQuestionsImport
    • First observedpostExamsidQuestionsquestionIdGenerate
    • First observedpostExamsidSessions
    • First observedpostExamsidSessionsScan
    • First observedpostExamsidSessionssessionIdFeedback
    • First observedpostExamsidSessionssessionIdGenerate_overall_feedback
    • First observedpostExamsidSessionssessionIdOverall_feedback
    • First observedpostExamsidSessionssessionIdSuggestionsAccept_all
    • First observedpostExamsidSessionssessionIdSuggestionsquestionIdAccept
    • First observedpostExamsImport
    • First observedpostFolders
    • First observedpostFoldersid
    • First observedpostOrgs
    • First observedpostQuestion_bank
    • First observedpostQuestion_types
    • First observedpostQuestion_typesidDisable
    • First observedpostQuestion_typesidEnable
    • First observedpostSource_materials
    • First observedpostSource_materialsidSlice
    • First observedpostTaxonomies
    • First observedpostTaxonomiesid
    • First observedpostUsers

TDQS

C2.5/5.0

Scored across 115 tools

Disambiguation4/5

Most tools are grouped by resource (e.g., exams_, groups_), making their purpose clear. However, some tool names like exams_import and sourceMaterials_create could be confused without descriptions, and the mix of singular/plural resource names (e.g., questionBank vs studentLevels) causes minor ambiguity.

Naming Consistency2/5

Naming conventions are inconsistent: some use snake_case (attributes_create), some use camelCase for the resource prefix (embedSessions_create), and a few use no separator (deleteRubricsid). This mix of styles makes the set feel unpolished.

Tool Count2/5

At 115 tools, the count is very high for an MCP server, likely overwhelming for agents. While the domain (exam management) is broad, many tools could be consolidated (e.g., separate tools for each exam session operation).

Completeness4/5

The tool surface covers the full lifecycle of exams, sessions, groups, rubrics, question banks, and more, including AI generation and export. Minor gaps exist (e.g., no bulk delete or analytics), but core workflows are well-supported.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers