Skip to main content
Glama

Caribooks (QuickBooks Online)

Server Details

QuickBooks Online in Claude and ChatGPT: 221 tools, full ledger, multi-company, Canada + US, FR/EN.

Ownership verified
Status
Healthy
Uptime
27.5% over 23 days
OAuth
Works in Glama
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.5/5.0

Scored across 220 tools

Disambiguation5/5

Every tool targets a distinct entity and action, following a clear verb_noun pattern. Even similar operations like attach_file, attach_receipt, and create_attachable are clearly separated by their descriptions.

Naming Consistency5/5

All tool names use snake_case verb_noun format, e.g., create_invoice, get_invoice, update_invoice, delete_invoice, search_invoices. Special tools like activate_loop and remember_rule follow the same pattern.

Tool Count1/5

With 220 tools, this server is far beyond any reasonable scope. Even for a comprehensive QuickBooks integration, the sheer number makes model selection impractical and overwhelming.

Completeness4/5

The tool set covers a very wide range of QuickBooks entities, reports, and actions, including CRUD for most core objects. However, some deletions are missing (e.g., customers, vendors, classes) and there is no void for purchases, leaving minor gaps.

Available Tools

220 tools
activate_loopStart a loopA
Destructive
Inspect

Activate an existing paused bookkeeping loop by id after the user approves its configuration. Sets the next run from its schedule. Depending on its approved configuration, a loop can change bookkeeping records or email invoice reminders to external recipients. Preview mode requires approval of proposed actions. The Caribooks portal provides pause controls and approval of proposed changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
loop_idYesThe loop's id, as create_task and list_loops show it.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this is a mutating operation. The description adds valuable context by noting that activation sets the next run from the schedule, and that depending on configuration the loop can change bookkeeping records or email invoice reminders to external recipients. It also mentions preview mode requiring approval, which is useful behavioral context. While it doesn't detail side effects like exactly what records change or email content, the annotations plus description sufficiently convey the risk profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise at about four sentences, with the core action front-loaded in the first sentence. Later sentences add relevant context about schedule, effects, and approval requirements. It is not overly verbose, though the last sentence about the Caribooks portal is somewhat tangential and could be trimmed without losing essential meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter, no output schema, and clear annotations, the description covers the essential context: what it does, when it's appropriate, the approval requirement, and potential side effects. It doesn't describe the return value, but without an output schema, that's less critical. The main gap is that it doesn't explicitly state prerequisites like 'loop must be paused' (though implied by 'existing paused') or what happens if the loop is already active, but overall it's fairly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter loop_id is well described in the schema itself ('The loop's id, as create_task and list_loops show it.'). The description doesn't add much beyond the schema, but it does contextualize that the loop must be paused and existing, which implies the ID refers to an existing paused loop. Since the schema already handles the parameter description well, a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool activates an existing paused bookkeeping loop by id after user approval, which is a specific verb-resource pair. It mentions setting the next run from its schedule and distinguishes between preview mode requiring approval and normal activation, but doesn't explicitly differentiate it from sibling tools like 'list_loops' or 'create_task' that might be related to loop management.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: to activate a paused loop after user approval, and mentions preview mode requires approval of proposed actions. However, it doesn't explicitly state when not to use it or name alternative tools for managing loops (e.g., if the user wants to modify a loop instead of activating it), though the context implies activation is the specific purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attach_fileAttach a file to a transactionA
Destructive
Inspect

Attach an original PDF, image, spreadsheet, CSV or Word file to an existing QuickBooks transaction. upload_id identifies a file in the Caribooks document box, populated through portal uploads or the account's receiving email address. For ChatGPT native attachments, file accepts the original file descriptor supplied by ChatGPT. content_base64 accepts the actual bytes of a small file. Maximum file size: 20 MB. Requires full access on the connection. Does not create a transaction or render a transcription.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNoA file the user attached in this conversation, passed by the assistant's app. Use this or upload_id or content_base64, not two.
txn_idYesQuickBooks Id of that transaction.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
filenameNoFile name with its extension, e.g. 'facture-2026-08.xlsx'. Shown in QuickBooks. Required with content_base64; with upload_id it defaults to the uploaded file's name.
txn_typeYesQuickBooks transaction type the file documents.
upload_idNoThe document id from list_receipt_inbox. Use this or content_base64, not both.
content_typeNoMedia type of the file, e.g. application/pdf, image/jpeg, text/csv, application/vnd.openxmlformats-officedocument.spreadsheetml.sheet. With content_base64 only; optional when it is a data URL that names it.
content_base64NoThe file's bytes encoded as base64 (a data URL is accepted too), for a small file you hold.
allow_duplicateNoThe call fails if the transaction already has an attachment. Only set true after the user explicitly confirms they want an additional document attached.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already indicating mutability and destructiveness, the description adds valuable context: the 20 MB size limit, the full-access permission requirement, and the fact that this action does not create a transaction or transcription. This goes beyond the structured hints without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well organized: opening with the core action, then clarifying the three file sources, then constraints, permissions, and exclusions. Every sentence earns its place and no schema information is needlessly repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation tool with no output schema, the description covers the key invocation concerns: what files are accepted, how to supply them, size limits, permissions, and side-effect boundaries. The only minor gap is that it does not describe the call's return value or confirmation behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all parameters, so the baseline is 3. The description adds meaningful extra semantics by explaining the three file-source modes: upload_id from the Caribooks document box, file for ChatGPT native descriptors, and content_base64 for small inline bytes. This helps agents choose the right mutually exclusive parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: attach a file to an existing QuickBooks transaction, and enumerates supported file types. It also clarifies that it does not create a transaction or render a transcription, which separates it from the create_* family. However, it does not explicitly differentiate itself from the closely named sibling attach_receipt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful context: it targets existing transactions, requires full access, and explicitly excludes creating transactions or rendering transcriptions. It does not name alternative tools such as attach_receipt or state when to prefer them, so the usage guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attach_receiptAttach a receipt to a transactionAInspect

Render user-provided receipt text as a PDF transcription and attach it to a QuickBooks transaction. Intended for receipts that exist only as text, such as an email body or copied order confirmation. The PDF is visibly marked as a transcription, not an original document. Requires confirmed_by_user: true after the user explicitly requests a transcription, and full access on the connection. The receipt fields contain the original text and, when applicable, the actual email sender and subject. Does not transfer original uploaded files.

ParametersJSON Schema
NameRequiredDescriptionDefault
txn_idYesQuickBooks Id of that transaction.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
receiptYesThe receipt content to render into the attached PDF.
filenameNoAttachment file name. Defaults to '<date> <sender>.pdf', or '<date> receipt.pdf' without a sender.
txn_typeYesQuickBooks transaction type the receipt documents.
allow_duplicateNoThe call fails if the transaction already has an attachment. Only set true after the user explicitly confirms they want an additional document attached.
confirmed_by_userNotrue only after the user explicitly asked for a rendered transcription, knowing it is not the original document. Without it the call is refused.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide readOnlyHint=false and destructiveHint=false, so the description carries the behavioral burden and fulfills it well. It discloses that the PDF is 'visibly marked as a transcription, not an original document,' that it requires explicit confirmation, that it fails on duplicates unless confirmed, and that it never transfers original uploaded files. This far exceeds the minimal annotation signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each serving a distinct purpose: what it does, when to use it, behavioral caveat, and explicit non-transfer. It is front-loaded with the core action and wastes no words. The structure is tight and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 parameters, a nested object, no output schema, and mutation semantics, the description covers all critical aspects: purpose, target scenario, prerequisites (confirmation and access), duplicate handling, and non-transfer of originals. Combined with the 100% schema descriptions, an agent has everything needed to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful context for receipt fields: 'The receipt fields contain the original text and, when applicable, the actual email sender and subject.' It also clarifies the filename default and emphasizes the confirmation gates on confirmed_by_user and allow_duplicate. That extra guidance compensates for the nested-object complexity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Render user-provided receipt text as a PDF transcription and attach it to a QuickBooks transaction.' It clearly distinguishes from siblings by narrowing to text-only receipts and explicitly stating 'Does not transfer original uploaded files,' which separates it from attach_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance ('Intended for receipts that exist only as text, such as an email body or copied order confirmation'), states prerequisites ('Requires confirmed_by_user: true... and full access on the connection'), and implies the alternative for original files with 'Does not transfer original uploaded files.' No ambiguity remains about when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_accountcreate accountCInspect

Create an account in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Account to create, as a QuickBooks Online API v3 Account object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. AccountType is required; AccountSubType picks the specific kind.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false and destructiveHint=false, which already establish that this is a mutating but non-destructive write operation. The description adds no behavioral traits beyond that, such as side effects, permissions, or impact on QuickBooks data. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundant wording. It is concise, though it forgoes the opportunity to add useful context; still, it earns its place as a clear statement of the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives the core operation, and the schema's entity description covers the important field-level restrictions and the required AccountType. However, the overall definition is minimal: it doesn't clarify the exact kind of account, provide sibling differentiation, or describe what happens after creation. This is adequate but not rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the required 'entity' object and optional 'company' parameter. The description adds no parameter-level meaning, but because the schema is thorough, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Create an account in QuickBooks Online.' The word 'account' plus the surrounding sibling tools (get_account, update_account, search_accounts) makes the target resource identifiable. It is slightly less specific than it could be because it doesn't say 'chart of accounts account,' but it is not tautological or misleading.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives like update_account or search_accounts, nor any mention of prerequisites or setup. The description only states the action without contextual direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_attachablecreate attachableAInspect

Create a text note attached to a QuickBooks record. Does not upload a file or render a PDF.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Attachable to create, as a QuickBooks Online API v3 Attachable object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. A note: Note (the text) plus AttachableRef: [{ EntityRef: { type: 'Bill', value: '<Id>' } }] to pin it to a record. Files are not created this way; attach_file uploads a file and attach_receipt renders receipt text as a PDF.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false, so the description doesn't need to restate those. It adds a useful boundary ('does not upload a file or render a PDF') but does not disclose auth requirements or side effects. Adequate for a create operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences, zero filler. The description earns its place by stating the core function and the key exclusion in the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a detailed schema and the boundary note about file vs. PDF, the agent has everything needed to call this tool correctly. No output schema exists, but a create operation's return value is not necessary for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter description is rich enough to explain the required structure (Note plus AttachableRef). The tool description adds no extra parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: create a text note attached to a QuickBooks record, and explicitly differentiates from file upload and PDF rendering. This makes it immediately clear what the tool does and how it differs from attach_file and attach_receipt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'Does not upload a file or render a PDF,' implying the alternatives, and the input schema explicitly names attach_file and attach_receipt as the tools for those actions. An agent gets clear routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_billcreate billAInspect

Create a bill in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Bill to create, as a QuickBooks Online API v3 Bill object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. An ItemBasedExpenseLineDetail line only works if the item itself carries an ExpenseAccountRef; otherwise use an AccountBasedExpenseLineDetail line naming the expense account.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false (write) and destructiveHint=false (non-destructive), so the description doesn't need to restate that. The description adds critical behavioral context: it instructs to avoid sending government identifiers, birth dates, and card-processing details, and clarifies the constraint for ItemBasedExpenseLineDetail lines. These go beyond annotations and are valuable for safe usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is a single, front-loaded sentence that delivers the core purpose immediately. It's concise, though it could incorporate a brief usage-hint or differentiation. The entity parameter description is embedded in the schema, so the main description itself is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a nested, complex parameter and no output schema, the description effectively covers the essential intent. The schema handles the detailed object structure. The description adds necessary guardrails about sensitive fields and line-detail conditions. No critical information for invoking the tool correctly seems missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The entity parameter description in the schema already explains the object. The tool description adds meaning beyond the schema by warning about sensitive data and explaining the line-detail rule, which are not in the schema's generic description. This enriches the parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'Create' and resource 'a bill' in 'QuickBooks Online.' This directly distinguishes it from siblings like create_bill_payment, create_purchase, and create_vendor_credit. The action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance is provided. The description does not mention alternatives or conditions for choosing this tool over create_purchase/invoice. However, the name and the singular resource strongly imply this is for creating a bill, and the parameter guidance implies the required structure. Usage is implied but not explicitly differentiated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_bill_paymentcreate bill paymentAInspect

Create a bill payment in QuickBooks Online. An accounting record of a payment to a vendor. It does not initiate a bank or card payment.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe BillPayment to create, as a QuickBooks Online API v3 BillPayment object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. PayType (Check|CreditCard) plus the matching CheckPayment.BankAccountRef or CreditCardPayment.CCAccountRef, and Line[].LinkedTxn: [{ TxnId, TxnType: 'Bill' }].
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations only indicate this is not read-only and not destructive. The description adds meaningful behavioral context by stating this is a bookkeeping record and does not trigger a real bank or card transfer, which is important for an agent to set user expectations correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler. The first states the action, the second gives context, and the third clarifies a key boundary. Every sentence earns its place and the key distinction is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested-object creation tool, the description plus rich entity schema is largely sufficient for correct invocation. The only notable gap is that no return value or failure behavior is described, but this is minor given the clear schema and boundary clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents the entity and company parameters well. The entity description details PayType, bank/card references, and LinkedTxn requirements. The tool description itself adds no extra parameter-level meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Create a bill payment in QuickBooks Online'), clarifies it is an accounting record of a payment to a vendor, and explicitly distinguishes it from actual payment initiation. This sets it apart from related tools like create_payment and create_credit_card_payment_txn.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: use this when recording an accounting payment to a vendor, not when disbursing funds. It does not name sibling alternatives explicitly, but the 'does not initiate a bank or card payment' boundary gives useful when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_budgetcreate budgetAInspect

Create a budget in QuickBooks Online. A profit-and-loss budget with one amount per account per period.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Budget to create, as a QuickBooks Online API v3 Budget object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Name, StartDate, EndDate, BudgetType: 'ProfitAndLoss', BudgetEntryType (Monthly|Quarterly|Annually) and BudgetDetail: [{ BudgetDate, Amount, AccountRef }] with one entry per account per period.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a mutation (readOnlyHint=false), and the description/schema add useful behavioral constraints: the budget must be P&L with one amount per account per period, and sensitive data like SSNs, tax IDs, birthdays, and card-processing details must not be sent. It does not describe response shape or duplicate-handling behavior, but the added constraints go beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is two concise sentences with the core action first and the qualifying scope second. The longer entity parameter description is dense but every phrase contributes required field, enum, array-shape, or privacy information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition provides enough information to construct a valid request: entity fields, acceptable enum values, company selection, and a privacy warning are all covered. There is no output schema and the nested AccountRef/date formats are not fully spelled out, leaving a small gap, but nothing essential for invoking the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters have schema descriptions (100% coverage), and the entity parameter description adds substantial meaning beyond the bare object type. It enumerates the required budget fields, allowed BudgetType and BudgetEntryType values, the BudgetDetail array shape, and a clear privacy restriction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: "Create a budget in QuickBooks Online." The second sentence narrows the scope to "a profit-and-loss budget with one amount per account per period," which clearly distinguishes this from the get/search/update/delete budget siblings and from other create_* entity tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is explicit: call this when creating a new QuickBooks budget, specifically a profit-and-loss budget. It does not explicitly mention when not to use it or point to update_budget/search_budgets as alternatives, but the create verb and the P&L constraint make the selection context unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_classcreate classCInspect

Create a class in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Class to create, as a QuickBooks Online API v3 Class object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds no behavioral detail beyond what the annotations already convey (readOnlyHint=false, destructiveHint=false). It does not disclose return behavior, idempotency, validation, or effects on QuickBooks data. The sensitive-data warning appears only in the schema, not in the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundant wording. It is efficient, though so minimal that it leaves usage and behavioral context to other fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward create operation, the description plus a fully documented schema gives an agent the essential action and parameters. However, with no output schema and no mention of return values, duplicate handling, or how this relates to update_class/search_classes, the definition is adequate but not rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters ('entity' and 'company') are documented in the input schema, including the restriction on sending government identifiers and card-processing details. The description itself adds no parameter-level meaning, but the high schema coverage justifies the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create') and resource ('a class') in the QuickBooks Online context, making the primary function clear. It does not explicitly differentiate from sibling tools like update_class or search_classes, but the verb and resource are specific enough to avoid core ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to choose this tool over update_class, search_classes, get_class, or other create_* tools. The verb 'Create' implies the use case, but the description provides no explicit context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_company_currencycreate company currencyAInspect

Create a company currency in QuickBooks Online. The currencies the company transacts in; only meaningful with multicurrency enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe CompanyCurrency to create, as a QuickBooks Online API v3 CompanyCurrency object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Code (ISO 4217, e.g. 'USD'). Requires multicurrency to be enabled in the company.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already signal this is not read-only (readOnlyHint=false) and not destructive (destructiveHint=false), so the 'Create' wording is consistent rather than additive. The description adds the multicurrency prerequisite and clarifies what a company currency represents, but it does not disclose behaviors like duplicate-currency errors, idempotency, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short, direct sentences with no filler. The primary action is front-loaded, and the additional context about multicurrency and the resource meaning earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple create-with-nested-object tool, the combination of description and rich schema covers the prerequisite, the object shape, and the sensitive-data restriction. It does not describe return values or error conditions, but no output schema exists and the create action is predictable enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the entity object, the company selector, and the ISO 4217 Code requirement. The main description adds little parameter-level detail, but that is acceptable because the schema carries the semantic load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Create a company currency') and a specific resource ('in QuickBooks Online'), which is unambiguous. It also distinguishes itself from sibling tools like update_company_currency, delete_company_currency, get_company_currency, and search_company_currencies by being the create operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context about when this is applicable: 'only meaningful with multicurrency enabled' and the schema reinforces the prerequisite. It does not explicitly name alternatives or exclusion cases, but the multicurrency precondition is a useful and clear usage signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_credit_card_payment_txncreate credit card payment txnAInspect

Create a credit card payment txn in QuickBooks Online. An accounting record of a bank payment toward a credit card balance. It does not move money.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe CreditCardPaymentTxn to create, as a QuickBooks Online API v3 CreditCardPaymentTxn object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. TxnDate, Amount, BankAccountRef (the paying bank account) and CreditCardAccountRef (the card being paid down) are required.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, meaning the tool writes but is not destructive. The description adds a key behavioral detail: 'It does not move money,' clarifying that this is purely an accounting record with no real-world fund movement. This is beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place. The first names the action and target, the second defines the concept, and the third negates a common misunderstanding. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple two-parameter schema with exhaustive parameter descriptions and the behavioral note about money not moving, the description is complete enough for an agent to invoke correctly. It does not describe return values, but no output schema is provided and none is required for this kind of creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter description already details required fields (TxnDate, Amount, BankAccountRef, CreditCardAccountRef) and restrictions on sensitive data. The tool description adds no parameter-level information; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the verb 'Create' and the resource 'credit card payment txn in QuickBooks Online'. Clarifies the concept with 'accounting record of a bank payment toward a credit card balance' and distinguishes it from money movement with 'It does not move money.' This is specific enough to separate from sibling tools like create_payment or create_transfer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: recording a bank payment toward a credit card balance. It does not explicitly name alternatives or state when to use vs create_payment/create_bill_payment, but the 'does not move money' clause gives a hint that it is not for actual transfers. This is implicit guidance, not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_credit_memocreate credit memoBInspect

Create a credit memo in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe CreditMemo to create, as a QuickBooks Online API v3 CreditMemo object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is not read-only and not destructive, but the description adds no behavioral context beyond the basic 'create' action. It does not disclose effects on QuickBooks, required permissions, validation behavior, or what response to expect, so the description contributes little beyond what the annotations already imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence with no filler, tautology, or redundant detail. It is appropriately front-loaded and easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema covers parameters well, but the overall definition lacks sibling differentiation, usage context, and behavioral expectations. An agent can likely invoke the tool in a straightforward case, but it is not fully equipped to choose this tool correctly or understand side effects without additional inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully documents the entity object and the optional company parameter. The description adds no parameter-level detail, but the baseline of 3 applies because the structured schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Create'), a clear resource ('a credit memo'), and the platform ('QuickBooks Online'), so an agent can identify what the tool does. It does not, however, differentiate this from closely related create_* siblings such as create_refund_receipt or create_vendor_credit, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives like send_credit_memo, update_credit_memo, create_refund_receipt, or create_credit_card_payment_txn. There is no mention of prerequisites, when a credit memo is the right transaction type, or when to choose another create tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_customercreate customerCInspect

Create a customer in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Customer to create, as a QuickBooks Online API v3 Customer object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=false and destructiveHint=false, so the description aligns by stating 'create'. However, the description adds no behavioral details beyond that—no mention of side effects, required permissions, or what happens on creation. It does not contradict annotations, but it provides no extra value beyond the annotation fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it is under-specified rather than appropriately sized. It lacks any structuring or front-loading of key info. While it is short, it does not earn its place because it provides minimal value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with two parameters and no output schema, the description is incomplete. It doesn't mention what the return value will be (likely the created customer object), any constraints on the entity, or how to select the company. The schema covers parameters, but the description provides no additional context that would help an agent call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (entity and company) in detail. The description itself adds no parameter information, so the baseline of 3 is appropriate. It does not compensate for any gaps because there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (create) and resource (customer) with the context 'in QuickBooks Online'. It is distinct from other create_* tools because the resource is customer, and it is not a tautology since it adds context. However, it doesn't mention alternatives like update_customer or search_customers, but that's not required for this dimension.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. No mention of prerequisites, when to use update_customer instead, or any conditions. The description is solely a statement of purpose, providing zero direction for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_departmentcreate departmentBInspect

Create a department in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Department to create, as a QuickBooks Online API v3 Department object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate a non-read-only write operation (readOnlyHint=false), and the description adds no further behavioral detail about side effects or permissions. However, the entity parameter description includes a useful compliance caveat—do not send government identifiers, birth dates, or card-processing details—which goes beyond the annotations by warning about data handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words, stating the verb, resource, and target system immediately. It is appropriately terse, though it omits usage guidance that would make it fully self-sufficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition is minimally viable: it names the operation clearly and the schema covers parameters, but the entity is a freeform object with no documented required fields. With no output schema, the expected result is undisclosed, and the agent must rely on external QuickBooks API knowledge to construct a valid Department object. The compliance warning helps, but overall the context is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The entity description identifies the expected QuickBooks API v3 Department object and adds a compliance warning, while the company parameter is clearly described with an optional-when-single-company note. The tool description itself adds no parameter meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create'), resource ('department'), and target system ('QuickBooks Online'), making it clear what the tool does. It is easily distinguished from update_department, get_department, and search_departments, though it does not explicitly contrast with sibling create_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like update_department or create_class. There is no mention of prerequisites, such as checking whether the department already exists, or any exclusions that would route an agent to another tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_depositcreate depositBInspect

Create a deposit in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Deposit to create, as a QuickBooks Online API v3 Deposit object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. DepositToAccountRef is required; each line needs DepositLineDetail.AccountRef.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a mutation (readOnlyHint=false) and not destructive. The description adds a useful compliance constraint by warning against sending SSN/tax IDs, birth dates, or card-processing details, but it does not disclose return behavior, side effects, or failure modes. Transparency is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one front-loaded sentence with no filler, immediately naming the action and target system. Every word earns its place, though the brevity does rely on the schema for additional detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create mutation with no output schema, the description is thin: it does not state what a successful call returns, how errors surface, or prerequisites such as an existing deposit account. The rich schema and mutation annotations make the tool callable, but the description leaves important operational context unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity property already documents the required Deposit object shape, DepositToAccountRef, per-line DepositLineDetail.AccountRef, and sensitive-data restrictions. The tool description itself adds no parameter meaning, so it stays at the high-coverage baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource, 'Create a deposit in QuickBooks Online,' so an agent can tell it apart from sibling create_* tools at the resource-name level. It does not elaborate on what a deposit involves or why it differs from similar transaction-creation tools, so it is clear but not maximally informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use create_deposit versus alternatives such as create_payment, create_credit_card_payment_txn, or update_deposit. No exclusions, prerequisites, or routing hints are provided, leaving the agent to infer usage solely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_employeecreate employeeCInspect

Create an employee in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Employee to create, as a QuickBooks Online API v3 Employee object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false, so the write behavior is known. The description adds no behavioral context beyond that: no mention of required permissions, idempotency, duplicate handling, or what effects creation has in QuickBooks. With minimal annotations, the description shoulders more burden but does not carry it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is concise and easy to parse, though its brevity comes at the cost of informative substance. As far as structure and economy, it earns a high score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no mention of return values, success criteria, or post-create state. The nested entity object is documented in the schema, but the description provides no completeness for what an agent should expect after calling this tool. For a mutation with only sparse annotations, this is a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description itself adds no parameter details, but the input schema covers 100% of parameters with meaningful descriptions, including the entity object restrictions and company selection. With full schema coverage, the baseline of 3 applies; the description does not need to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create an employee in QuickBooks Online.' This distinguishes it from sibling create_customer, create_vendor, and other create_* tools. However, it lacks any additional scoping or differentiation beyond the resource name, so it is clear but not strongly differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to use this tool versus alternatives, when not to use it, or any prerequisites. The single sentence implies the obvious use case but offers no decision support for an agent choosing among the many create_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_estimatecreate estimateBInspect

Create an estimate in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Estimate to create, as a QuickBooks Online API v3 Estimate object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description only says 'Create an estimate' and does not disclose any side effects, permissions, or data handling behavior. The annotations are minimal (readOnlyHint=false, destructiveHint=false) and do not add context. Although the entity parameter warns against sending sensitive data, that is a parameter constraint, not a tool behavioral trait. The description carries the burden and does not satisfy it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no redundant phrases or extra details. It is concise and front-loaded with the core operation, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested entity object, no output schema), the description is minimal but the schema carries substantial detail including required fields and data constraints. However, there is no guidance on expected return values or when to use this over related tools, leaving some gaps. Still, the overall definition is enough for a competent agent to invoke correctly with the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both entity and company parameters have descriptive text, with the entity description warning about government identifiers and card-processing details. The tool description adds no parameter information beyond what the schema already provides, which meets the baseline for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Create an estimate in QuickBooks Online', which is a clear, specific verb+resource statement. It distinguishes this from the many other create_* tools in the sibling list by naming 'estimate' as the resource. It is not a tautology and is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives like create_invoice, create_sales_receipt, or the related search_estimates/update_estimate. There are no prerequisites, exclusions, or conditions stated, so the agent must infer usage from the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_inventory_adjustmentcreate inventory adjustmentAInspect

Create an inventory adjustment in QuickBooks Online. Changes the quantity on hand of inventory items and books the valuation difference; US companies on Plus or Advanced only.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe InventoryAdjustment to create, as a QuickBooks Online API v3 InventoryAdjustment object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. TxnDate, DocNumber, AdjustAccountRef (the inventory shrinkage/adjustment account) and Line: [{ DetailType: 'ItemAdjustmentLineDetail', ItemAdjustmentLineDetail: { ItemRef, QtyDiff (signed change) or NewQty } }].
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate this is a write operation (readOnlyHint false) and not destructive (destructiveHint false). The description adds the specific behavioral detail that it books a valuation difference, and the plan/region constraint, which are not covered by annotations. It does not disclose potential side effects like requiring specific permissions or creating an irreversible transaction, but the added context is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the action and then efficiently adds the effect and the usage constraint. There is no redundancy or fluff; every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description includes the critical region/plan constraint, and the schema thoroughly covers the complex nested entity parameter. There is no output schema, so no return-value details are needed. It does not mention prerequisites like having inventory items set up, but the core information an agent needs to decide and invoke the tool is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are fully documented in the schema. The entity parameter has a detailed description explaining the structure and required fields. The tool description does not add any parameter-specific guidance beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (create an inventory adjustment), the resource (inventory adjustment in QuickBooks Online), and the specific effect (changes quantity on hand and books valuation difference). It also includes a scope constraint (US companies on Plus or Advanced only), which distinguishes it from other create tools and clarifies its applicability.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a concrete usage condition (US companies on Plus or Advanced) and implies the purpose of adjusting inventory quantities. It does not explicitly mention alternatives like update_inventory_adjustment, but the create vs. update distinction is implicit. The constraint is a strong usage guideline, though exclusions are not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_invoicecreate invoiceAInspect

Create an invoice in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Invoice to create, as a QuickBooks Online API v3 Invoice object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the agent knows this is a mutating operation. The description adds no behavioral details beyond 'create' – no mention of required permissions, side effects, or what happens on success/failure. The entity parameter description does add a useful constraint (do not send government identifiers, birth dates, card-processing details), which is a behavioral guardrail. However, the core description itself is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that front-loads the core purpose. It is appropriately concise. The schema's entity description is longer but earns its place by providing critical data-handling constraints. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create operation with a nested object parameter and no output schema, the description is adequate but not rich. It tells the agent what to create and warns about sensitive fields, but doesn't describe the expected response, error conditions, or any prerequisites (e.g., customer must exist). Given the complexity of a QuickBooks Invoice object, more guidance on required sub-fields or common pitfalls would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The entity parameter description adds meaningful guidance beyond the schema: it explicitly warns against sending sensitive data (SSN, tax IDs, birth dates, card-processing details) and directs managing those fields in QuickBooks. The company parameter is also clearly explained. This exceeds the baseline 3 for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create an invoice in QuickBooks Online.' This clearly identifies the operation and target system. It doesn't explicitly differentiate from sibling tools like create_credit_memo or create_sales_receipt, but the resource 'invoice' is distinct enough among the create_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this when you need to create an invoice in QuickBooks. It doesn't explicitly state when not to use it or mention alternatives like create_sales_receipt or create_credit_memo. The schema's entity description adds some context about what to send, but no explicit routing guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_itemcreate itemAInspect

Create an item in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Item to create, as a QuickBooks Online API v3 Item object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Type is required: Service or NonInventory need Name and IncomeAccountRef (ExpenseAccountRef too if purchased); Inventory additionally needs AssetAccountRef, ExpenseAccountRef, TrackQtyOnHand: true, QtyOnHand and InvStartDate; Category needs only Name (nest with ParentRef); Group (a bundle) needs ItemGroupDetail.ItemGroupLine.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a write operation (readOnlyHint=false) and not destructive, so the description correctly stays consistent. It adds meaningful behavioral context by instructing the agent not to send government identifiers, birth dates, or card-processing details and to manage those fields directly in QuickBooks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is a single clear sentence, and the entity parameter description is dense but every clause earns its place. It packs type-based requirements and a privacy warning into a compact, readable format without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex nested create operation, the definition covers the entity shape, required references, confidential-data restrictions, and how to target a company. It does not mention return values or error/validation behavior, and there is no output schema, leaving a minor gap for the agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema description coverage is 100%, the description adds substantial meaning beyond the field names. It enumerates type-specific required fields (Service/NonInventory, Inventory, Category, Group) and explains what company identifiers are accepted, giving an agent enough detail to construct a valid entity without external lookups.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create') and resource ('an item in QuickBooks Online') without ambiguity. It does not explicitly differentiate from the many sibling create_* tools, but the QBO Item resource is well-defined enough that an agent can tell it from create_account, create_invoice, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as update_item, get_item, or search_items. No conditions, prerequisites, or exclusions are provided; any usage signal is only implicit in the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_journal_entrycreate journal entryBInspect

Create a journal entry in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe JournalEntry to create, as a QuickBooks Online API v3 JournalEntry object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Lines must balance: equal totals of PostingType Debit and Credit.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a write operation (readOnlyHint=false) and non-destructive (destructiveHint=false). The description adds no further behavioral context, such as side effects, idempotency, posting behavior, or result handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is appropriately concise given the input schema carries the detailed parameter guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple create tool but omits what a successful call returns and any operational context (e.g., whether the journal entry is posted or saved as a draft). The schema covers parameters, so the main gap is the lack of invocation-level detail beyond creation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed parameter descriptions for entity and company. The tool description itself adds no parameter-level meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Create') and a specific resource ('journal entry') in a named context ('QuickBooks Online'), which clearly distinguishes it from sibling create_* tools. It goes beyond a bare tautology by specifying the platform.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as create_invoice, create_bill, or update_journal_entry. Any usage signal is purely inferred from the tool name, not explicitly described.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_paymentcreate paymentAInspect

Create a payment in QuickBooks Online. An accounting record of a customer payment already received. Payment processing is not supported; ProcessPayment must be omitted or false.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Payment to create, as a QuickBooks Online API v3 Payment object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Link the invoices it settles with Line[].LinkedTxn: [{ TxnId, TxnType: 'Invoice' }].
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a write operation that is not destructive. The description adds a meaningful behavioral boundary: this tool records a payment rather than executing payment processing. It does not detail side effects or idempotency, but the key non-obvious behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core action is front-loaded, and the critical caveat about payment processing is stated immediately afterward.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description, combined with the fully documented schema, covers the essential semantics and the most important restriction for this create operation. It does not explicitly describe the return value or side effects after creation, but the annotations clarify the write/non-destructive profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents `entity` and `company`. The description adds one actionable field-level constraint beyond the schema—`ProcessPayment` must be omitted or false—and clarifies the semantic meaning of the payment object as an accounting record rather than a processing action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create a payment in QuickBooks Online.' It further clarifies that this is an accounting record of an already-received customer payment, which distinguishes it from payment-processing or payment-method creation tools among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'customer payment already received' gives clear context for when to use this tool, and 'Payment processing is not supported; ProcessPayment must be omitted or false' is an explicit when-not. However, it does not name an alternative tool for payment processing, so it stops short of full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_payment_methodcreate payment methodBInspect

Create a payment method in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe PaymentMethod to create, as a QuickBooks Online API v3 PaymentMethod object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The tool description adds no behavioral nuance beyond the annotations (readOnlyHint=false, destructiveHint=false). It simply restates the action. Notably, the rich data-handling warning ('Do not send government identifiers...') lives in the schema property description, not in the tool description, so it does not count toward behavioral disclosure here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler: 'Create a payment method in QuickBooks Online.' It is appropriately front-loaded and compact for a simple create operation, with zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The input side is thoroughly covered by the schema, but the description omits expected return behavior (no output schema exists) and fails to disambiguate from create_payment. For a simple create with no annotations beyond readOnly/destructive hints, this is adequate but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The entity parameter description is strong: it identifies the object type, references QuickBooks Online API v3, and warns against sensitive fields. The company parameter is also well explained. The tool description itself adds no additional parameter meaning, so it does not exceed baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource: 'Create a payment method in QuickBooks Online.' This identifies the operation unambiguously at a glance. However, it does not explicitly distinguish itself from the similarly named sibling create_payment, so the differentiation burden falls on the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives like create_payment, update_payment_method, or search_payment_methods. The general purpose is implied by the name, but no explicit context or exclusion criteria are given to prevent misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_purchasecreate purchaseBInspect

Create a purchase in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Purchase to create, as a QuickBooks Online API v3 Purchase object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. PaymentType (Cash|Check|CreditCard) and AccountRef (the paying account) are both required.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false, so the agent knows this is a mutation that is not read-only. The description adds no behavioral context beyond the act of creating, such as side effects on company books, irreversibility, permission requirements, or response behavior. It does not contradict the annotations, but it also does not disclose anything beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single direct sentence with no filler or redundancy. Every word contributes to stating the action and target system, and the most important information is front-loaded. This is an appropriately concise definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema carries substantial parameter context, covering the nested entity structure and company parameter. However, the tool operates on a complex object with no output schema, and the description offers no high-level context about return behavior or how purchase differs from related transaction types. It is minimally viable but would benefit from a brief note on what a purchase represents in QuickBooks.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the input schema already provides detailed meaning for both parameters, including the required fields, PII handling warning, and company selection behavior. The description itself adds no parameter-level information, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create a purchase in QuickBooks Online.' This clearly identifies the operation and the target system, and the noun 'purchase' distinguishes it from sibling tools like create_bill, create_purchase_order, or create_vendor_credit. It lacks explicit sibling differentiation, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool versus alternatives such as create_bill, create_purchase_order, or create_payment. The description simply names the action without explaining its typical use case, prerequisites, or exclusions. An agent must infer which financial document is intended.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_purchase_ordercreate purchase orderBInspect

Create a purchase order in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe PurchaseOrder to create, as a QuickBooks Online API v3 PurchaseOrder object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. An ItemBasedExpenseLineDetail line only works if the item itself carries an ExpenseAccountRef; without one QuickBooks answers 'Select an account for this transaction'. Use an AccountBasedExpenseLineDetail line instead, or set the item's expense account first.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal that this is a mutating operation (readOnlyHint=false, destructiveHint=false), and the description adds no behavioral context beyond that. It does not mention prerequisites, whether duplicate POs are allowed, or what side effects occur in QuickBooks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is appropriately sized for a tool whose detailed parameter guidance already lives in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema covers entity complexity and company selection well, and annotations cover the basic safety profile. However, the overall definition omits what the tool returns, prerequisites, and when to choose a purchase order over sibling purchase/bill tools, leaving some gaps for a create operation with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter description is unusually detailed, including sensitive-data warnings and a specific QuickBooks error scenario. However, the tool description itself adds no additional parameter meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: it creates a purchase order in QuickBooks Online. This is unambiguous and distinguishes it from get/search/update/delete siblings, though it does not explicitly contrast it with similar create tools such as create_purchase or create_bill.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus create_purchase, create_bill, or create_vendor_credit. There are no alternatives, exclusions, or contextual conditions, so the agent must infer usage solely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_recurring_transactioncreate recurring transactionB
Destructive
Inspect

Create a recurring transaction in QuickBooks Online. A template QuickBooks uses to create a transaction on a schedule, or on request.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe RecurringTransaction to create, as a QuickBooks Online API v3 RecurringTransaction object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. The template's transaction under its own type, carrying RecurringInfo: e.g. { Invoice: { ...invoice fields..., RecurringInfo: { Name, RecurType: 'Automated'|'Reminded'|'UnScheduled', Active: true, ScheduleInfo: { IntervalType: 'Monthly', NumInterval: 1, DayOfMonth, StartDate, NextDate, ... } } } }.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate that this is a write operation (readOnlyHint: false) and potentially destructive (destructiveHint: true). The description adds no behavioral context beyond that—no mention of side effects, permission requirements, or what happens to existing data. With annotations present, the bar is lower, but the description still offers no added transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise—two sentences that front-load the purpose and define the term 'recurring transaction'. It is efficiently written with no wasted words, though it could benefit from more context without harming conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex (nested RecurringInfo structure, multiple fields), and there is no output schema. The description does not explain what the response looks like, what errors might occur, or any prerequisites (like company selection). The agent is left to infer behavior from the schema alone, which is insufficient for a create operation with this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema documents both parameters. However, the description adds valuable semantics beyond the schema: it provides a concrete example of the RecurringInfo structure and explicitly warns against sending sensitive data (SSN, tax IDs, birth dates, card-processing details). This enhances the agent's understanding of the entity parameter beyond the schema's generic description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a recurring transaction in QuickBooks Online') and further clarifies that it's a template used to create transactions on a schedule or on request. This clearly distinguishes it from sibling tools like create_invoice or create_bill, which create actual transactions rather than recurring templates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide any guidance on when to use this tool versus alternatives. It does not mention exclusions, prerequisites, or suggest when to prefer other create_* tools. The purpose is implicit but no explicit usage context is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_refund_receiptcreate refund receiptBInspect

Create a refund receipt in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe RefundReceipt to create, as a QuickBooks Online API v3 RefundReceipt object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. DepositToAccountRef names the account the refund was paid from.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is minimal, providing no behavioral details beyond the fact that it creates a refund receipt. Even though annotations indicate readOnlyHint=false (so the agent knows it's a write operation), the description adds no information about what happens on creation (e.g., whether it affects inventory, requires specific account settings, or if it triggers any side effects). For a mutation tool, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the primary action. It wastes no words and is easily scannable. However, it could have been slightly more informative without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool involves a nested object parameter ('entity') and has no output schema, but the description doesn't explain what response to expect or any validation constraints. It also doesn't mention required fields within the entity object (though that's partly covered in the schema description, but not fully). Given the complexity of QuickBooks entities, the description is thin and omits important context like what constitutes a valid refund receipt.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides detailed descriptions for both parameters. The 'entity' parameter is well-explained, including what to include or exclude (government identifiers, birth dates, etc.) and naming the DepositToAccountRef field. The 'company' parameter is also adequately described. Because schema coverage is 100%, the description itself doesn't need to add much, but the schema's rich descriptions raise the score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource ('Create a refund receipt in QuickBooks Online'). It is specific and distinct from tools like create_sales_receipt or create_credit_memo by naming the specific receipt type. However, it doesn't explicitly differentiate it from the sibling create_sales_receipt, which might be confused, but the name itself is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like create_sales_receipt or create_credit_memo. The description doesn't mention any prerequisites, such as having an existing customer or the appropriate permissions. An agent would have to infer usage from the tool name and context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_sales_receiptcreate sales receiptCInspect

Create a sales receipt in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe SalesReceipt to create, as a QuickBooks Online API v3 SalesReceipt object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. DepositToAccountRef names the account the money landed in.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a mutating operation (readOnlyHint=false, destructiveHint=false), and the description merely restates that by saying 'Create'. It adds no detail about side effects, required permissions, idempotency, or how the created receipt is returned, so it contributes no behavioral transparency beyond structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no filler, and the key verb and resource are front-loaded. While very short, it is structurally clean and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested entity and many create_* siblings, the description is too sparse to be contextually complete: it does not mention the document type's typical use, what happens after creation, or any relationship to send_sales_receipt. The structured schema fills parameter details, but selection context and workflow are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the entity property is richly documented in the schema, including the warning about government identifiers. The tool description adds no parameter-level meaning of its own, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the operation as creating a sales receipt in QuickBooks Online, giving a specific verb and resource. However, it makes no reference to sibling tools like create_invoice or create_credit_memo, so it does not help an agent distinguish when a sales receipt is the right document type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose this tool over alternatives such as create_invoice, create_credit_memo, or create_refund_receipt, nor any prerequisites or exclusions. The only usage signal is the tool name itself, which is not enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_taskCreate a task or a standing loopAInspect

Create a paused bookkeeping loop from a user-requested task, in the user's language. Suitable for recurring work on a schedule or after documents reach the Caribooks document box. The loop defines its data access, permitted actions and approval requirements. Returns the saved task and its proposed configuration for review. Creating a task does not start it or write to QuickBooks. New loops start in preview mode.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
run_nowNotrue to start the loop and run it once immediately, only when the user asked for the work now.
scheduleNorecurring for standing work, once for a single errand. Defaults to recurring when the sentence names a rhythm (every morning, chaque lundi) or waits on arriving documents, and once otherwise.
sentenceYesWhat the user wants done, in their own words and their own language, one or two sentences. Do not translate or tidy it.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses useful behaviors beyond the annotations (paused, no QuickBooks write, preview mode, returns configuration for review). However, it states 'Creating a task does not start it' without qualification, yet the run_now parameter can start the loop immediately. This overgeneralization is misleading and could cause an agent to ignore run_now when the user asks for work now. The annotations (readOnlyHint=false, destructiveHint=false) do not contradict, but the internal inconsistency reduces transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is six sentences, each carrying meaningful information: purpose, suitability, internal behavior, return value, safety, and preview mode. It is front-loaded with the main purpose and avoids fluff, though the final two sentences could be tightened. Overall, it is efficient for the amount of context it provides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers return value ('Returns the saved task and its proposed configuration') and safety, which helps given there is no output schema. However, the absolute claim that creating a task does not start it conflicts with the run_now parameter, leaving a gap for agents handling immediate-execution requests. It also doesn't elaborate on when company selection is needed, though the schema covers that. These gaps prevent full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with parameter descriptions, so the baseline is 3. The description adds semantic value by framing the sentence as the user's own words and language, and by explaining the loop is paused and returns a configuration for review. This extra context helps an agent understand the intent behind the parameters, especially sentence and run_now.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create a paused bookkeeping loop from a user-requested task'. It also explicitly distinguishes this tool from the many sibling create_* tools for QuickBooks entities by noting it does not write to QuickBooks. The title 'Create a task or a standing loop' further clarifies it handles both one-off and recurring work, making its purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear when-to-use signal: 'Suitable for recurring work on a schedule or after documents reach the Caribooks document box.' It also implies safety constraints (paused, no QuickBooks write, preview mode). However, it does not explicitly name alternatives like activate_loop or mention run_now as the path for immediate execution, and it lacks an explicit 'when not to use' against other create_* tools, though the loop concept is distinct enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_tax_agencycreate tax agencyA
Destructive
Inspect

Create a tax agency in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe TaxAgency to create, as a QuickBooks Online API v3 TaxAgency object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Creating a tax agency cannot be undone: QuickBooks also creates payable and suspense accounts that can never be deleted. Only create one when the user has asked for it.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though destructiveHint=true already signals destructiveness, the description adds valuable behavioral context: creating a tax agency cannot be undone and automatically creates payable and suspense accounts that can never be deleted. It also warns against sending government identifiers, birth dates, and card-processing details, which is critical safety information beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is a single efficient sentence, and the parameter descriptions pack important warnings into a compact, front-loaded structure. Every sentence earns its place without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition covers the irreversible consequences, sensitive data exclusions, and the one required complex object. It does not describe the return value, but there is no output schema and the safety-critical context is sufficiently complete for an agent to invoke the tool responsibly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly, including the entity object shape and the optional company field. The tool description itself adds no parameter-level semantics, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Create a tax agency in QuickBooks Online.' It is clear and unambiguous about the core action. However, it does not explicitly differentiate from the sibling create_tax_service, so an agent must infer the distinction from the resource name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage guidance is buried in the entity parameter: 'Only create one when the user has asked for it.' This implies the tool should not be used casually, but there is no explicit mention of when to prefer search_tax_agencies/get_tax_agency or how to distinguish this from create_tax_service.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_tax_servicecreate tax serviceA
Destructive
Inspect

Create a sales tax code in QuickBooks Online, with new or existing tax rates attached to it. Cannot be undone; only on an explicit request.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe TaxService to create, as a QuickBooks Online API v3 TaxService object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. { TaxCode: '<new code name>', TaxRateDetails: [{ TaxRateName, RateValue (0-100), TaxAgencyId (an existing TaxAgency Id), TaxApplicableOn: 'Sales'|'Purchase' }] }, or TaxRateId to reuse an existing rate. A tax code or rate cannot be deleted once created, only made inactive. Only create one when the user has asked for it.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as destructive, and the description reinforces that with 'Cannot be undone' and 'only on an explicit request.' This adds meaningful behavioral context beyond the generic destructiveHint flag. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both purposeful. The core action is front-loaded and the irreversible side effect and authorization condition are stated immediately afterward with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create operation with a nested object, the entity schema supplies the necessary structural detail, and the description supplies the key caveats: irreversibility and explicit-request-only usage. It does not describe the return value, but no output schema exists and this gap is minor for a create action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the entity parameter is richly documented, including privacy restrictions, required fields, and inactive-vs-delete semantics. The tool description itself does not need to add parameter detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Create a sales tax code in QuickBooks Online, with new or existing tax rates attached to it.' This differentiates it from the sibling create_tax_agency and from read-only tax tools like get_tax_code and search_tax_codes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when invocation is appropriate: 'only on an explicit request' and 'Cannot be undone.' It lacks an explicit mention of alternatives, but the purpose is unambiguous enough that an agent can decide when to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_termcreate termCInspect

Create a term in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Term to create, as a QuickBooks Online API v3 Term object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal a write operation (readOnlyHint: false) and non-destructive intent (destructiveHint: false). The description adds little beyond restating the mutation: it does not mention persistence, side effects, validation behavior, or any other operational traits that the annotations do not already cover. There is no annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and free of filler, but it is under-specified to the point of being barely informative. It is concise but not well-rounded; it earns its place only as a minimal restatement of the tool's name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and an open nested object parameter, the agent needs more context about what a Term contains and what a successful creation returns. The description does not provide that context, nor does it reference related tools like get_term or update_term that could help. It is incomplete for a tool that creates a domain-specific object.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema provides robust parameter detail: the entity is described as a 'QuickBooks Online API v3 Term object' with explicit warnings about sensitive data, and the company parameter is explained. The tool description itself adds no parameter meaning, but the schema carries the burden adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Create a term in QuickBooks Online.' It unambiguously identifies the target resource, distinguishing it from sibling create_* tools. However, it does not explain what a QBO 'term' actually is, leaving some ambiguity for an agent unfamiliar with QuickBooks terminology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as update_term, get_term, or search_terms. There is no context about prerequisites, exclusions, or conditions that would help an agent select this tool over a sibling. The usage context is only implied by the verb 'create.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_time_activitycreate time activityCInspect

Create a time activity in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe TimeActivity to create, as a QuickBooks Online API v3 TimeActivity object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. NameOf ('Employee' or 'Vendor') with the matching EmployeeRef or VendorRef, plus Hours and Minutes.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false, so the agent knows this is a mutating but non-destructive operation. The description adds no behavioral context beyond that – nothing about side effects, required related entities, permissions, or rate limits. Since it discloses nothing beyond what annotations already convey, it earns a low score without constituting a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero wasted words. It efficiently states the core operation and the system context. It is slightly too terse to earn a 5, as it omits any additional value-add, but it is well-structured for its brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with a nested entity object and no output schema, the description is minimal. The schema compensates with rich parameter documentation, but the description itself does not mention prerequisites (e.g., existing employee/vendor), the NameOf distinction, or what a successful response contains. This is adequate for basic selection but leaves gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters: entity (with detailed guidance on TimeActivity structure and sensitive-field restrictions) and company. The description adds no parameter-level meaning, but because the schema carries the burden, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a time activity in QuickBooks Online'), which clearly identifies the operation and the object type. It is distinguishable from sibling tools like update_time_activity or get_time_activity by the verb and resource. However, it does not explicitly call out sibling alternatives, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, such as update_time_activity for modifying existing entries or search_time_activities for retrieval. There are no prerequisites, exclusions, or context to help the agent decide. The only implied usage is the literal meaning of 'create'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_transfercreate transferAInspect

Create a transfer in QuickBooks Online. An accounting entry recording a transfer between accounts. It does not move money between bank accounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Transfer to create, as a QuickBooks Online API v3 Transfer object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. FromAccountRef, ToAccountRef and Amount are required, and the two accounts must differ.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a key behavioral fact beyond the annotations: that the tool creates an accounting entry rather than moving money. This is crucial because the name alone might suggest a bank transfer. The annotations already indicate a non-read-only, non-destructive write operation, so the description's contribution is valuable but not comprehensive (e.g., no mention of permissions or side effects).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the main action. The three sentences each add essential information: what it does, what kind of entry it is, and what it does NOT do. There are no wasted words, though it could be slightly more compact by merging the first two sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description, combined with the schema, provides enough information for an agent to invoke the tool correctly. It covers the crucial non-obvious behavior (not moving money) and the schema covers parameter constraints. It omits any mention of permissions or the response format, but for a simple create operation with full schema coverage, this is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with detailed documentation of the entity object (required FromAccountRef, ToAccountRef, Amount; accounts must differ) and the company parameter. The tool description adds no additional parameter-level information, so it relies fully on the schema. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Create') and resource ('transfer in QuickBooks Online'), and adds a clarifying definition ('An accounting entry recording a transfer between accounts'). It explicitly distinguishes itself from actual bank transfers, which is a common source of confusion and helps differentiate it from other tools. This is a specific, unambiguous purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit exclusion: 'It does not move money between bank accounts.' This tells the agent when NOT to use this tool and implies the tool is for accounting entries only. However, it doesn't name a specific alternative tool for actual bank transfers, so the guidance is clear but incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_vendorcreate vendorBInspect

Create a vendor in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe Vendor to create, as a QuickBooks Online API v3 Vendor object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false, so the description correctly implies a mutation. It adds important context about not sending government identifiers and managing those fields directly in QuickBooks, which is valuable beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, direct sentence with no wasted words. The crucial note about government identifiers is left to the schema, but the description itself is appropriately terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 params, no output schema) and the rich schema descriptions, the description is complete enough for an agent to use it. However, it lacks any note about prerequisites like company connectivity, which might be inferred from the schema but could be more explicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides full descriptions for both parameters, including the structure of the 'entity' object and the 'company' parameter. The description adds no additional parameter-level information, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (create) and resource (vendor) in QuickBooks Online. It is distinct from siblings like create_customer or create_bill, but doesn't explicitly differentiate from create_vendor_credit, which is a related but different action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for creating a vendor but doesn't provide context on when to prefer this over create_vendor_credit or other vendor-related tools. No explicit exclusions or alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_vendor_creditcreate vendor creditBInspect

Create a vendor credit in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe VendorCredit to create, as a QuickBooks Online API v3 VendorCredit object. Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds no behavioral context beyond what the annotations already convey: it simply says 'Create,' matching the readOnlyHint=false annotation. It does not disclose side effects on accounting records, whether the operation is reversible, idempotency, or error conditions. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no filler. It communicates the essential operation and platform efficiently, and every word contributes to the meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple create operation with a detailed input schema, this is minimally adequate. The description lacks usage guidance, expected return behavior, and side-effect awareness, but the schema covers the critical entity details. Given no output schema and a nested object parameter, more context would help an agent invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool description provides no parameter-level meaning, but the input schema has 100% coverage and already describes the VendorCredit entity and the optional company parameter. The entity schema adds useful context about prohibited sensitive fields, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the operation: 'Create a vendor credit in QuickBooks Online.' It identifies a specific verb and resource and adds the platform context. However, it does not differentiate this tool from related create tools like create_credit_memo or create_bill, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites (e.g., vendor must exist), and no mention of related read/update/delete operations. An agent must infer usage solely from the tool name and generic create pattern.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_attachabledelete attachableA
Destructive
Inspect

Permanently delete an attachable from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Attachable to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the destructive nature is covered. The description adds valuable behavioral context by stating the deletion is permanent and that explicit user confirmation is required. This goes beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, each earning its place: one states the action and scope, the other the required confirmation. There is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple destructive operation with a rich sibling list and clear annotations, the description is sufficient. It communicates irreversibility and the safety requirement. It could mention what happens to associated file relationships, but this is not essential for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are already documented. The description reinforces that confirmation must be provided but does not add new parameter details. Baseline 3 is appropriate since the schema carries the semantic burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Permanently delete'), a specific resource ('an attachable'), and the system ('QuickBooks Online'). The word 'permanently' disambiguates from update or detach operations. It clearly distinguishes this from siblings like create_attachable, update_attachable, and search_attachables.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear what the tool does and that it deletes (not creates or updates) an attachable, so usage context is implied. However, it does not explicitly state when not to use it or name alternatives such as update_attachable for reversible changes. Guidance is adequate but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_billdelete billA
Destructive
Inspect

Permanently delete a bill from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Bill to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true and readOnlyHint=false. The description adds value by stating 'Permanently delete' and 'Requires explicit user confirmation', which are behavioral traits beyond the annotations. There is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the action and the key prerequisite. Every word earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple deletion tool with a fully documented schema and no output schema, the description provides sufficient context: it names the action, the target, and the confirmation requirement. No critical information is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters (id, company, confirm) are already documented. The description does not add extra semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'delete', the resource 'bill', and the system 'QuickBooks Online'. It is specific enough to distinguish from other delete_* tools (e.g., delete_bill_payment, delete_credit_memo) by explicitly naming 'bill'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes the usage requirement 'Requires explicit user confirmation', which tells the agent when it is appropriate to invoke the tool. It does not explicitly contrast with alternatives, but given the resource-specific naming and the presence of many delete tools, the intended use case is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_bill_paymentdelete bill paymentA
Destructive
Inspect

Permanently delete a bill payment from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the BillPayment to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, but the description adds meaningful behavioral context: the deletion is permanent and requires explicit user confirmation. This tells the agent this is irreversible and gated on user consent, going beyond the annotation alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The primary action and key safety requirement are front-loaded, and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, destructive operation with three well-documented parameters, the description is nearly complete. It conveys irreversibility and the confirmation safeguard; only minor details about response behavior or effects on related records are omitted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents id, company, and confirm. The description's confirmation requirement reinforces the confirm parameter but does not add meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Permanently delete a bill payment from QuickBooks Online.' The word 'permanently' also helps distinguish this from the sibling void_bill_payment, making the tool's purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys that this is for permanent deletion and requires explicit user confirmation, which implies appropriate usage. However, it does not explicitly compare against alternatives like void_bill_payment or state when one should be preferred over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_budgetdelete budgetA
Destructive
Inspect

Permanently delete a budget from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Budget to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, and readOnlyHint=false, establishing the operation as destructive. The description adds the important behavioral details that the deletion is permanent and requires explicit user confirmation, which goes beyond the annotation. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two tight sentences with no filler. The first sentence front-loads the core action and resource, and the second adds a critical safety prerequisite. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward delete operation with fully documented parameters and a destructiveHint annotation, the description is complete. It conveys the essential context: permanence, the need for explicit consent, and the target system. The absence of an output schema is acceptable because delete operations typically don't require return-value documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents all three parameters (id, company, confirm) with 100% coverage, so the baseline is 3. The description's mention of 'explicit user confirmation' mirrors the confirm parameter's schema description but adds no new parameter semantics. The description does not elaborate on id or company beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('delete'), resource ('budget'), and system ('QuickBooks Online'), with the key qualifier 'permanently'. This clearly distinguishes it from sibling delete tools and from update_budget. The confirmation requirement adds further semantic precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear usage context: use when a budget must be permanently removed, and only after explicit user confirmation. It does not explicitly name alternatives like update_budget for non-destructive changes, but the confirmation prerequisite and the delete semantics make the intended use unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_company_currencydelete company currencyA
Destructive
Inspect

Deactivate a company currency in QuickBooks Online. Existing transaction references remain intact. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CompanyCurrency to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint: true and readOnlyHint: false, so the safety profile is known. The description adds useful behavioral nuance beyond that: 'Existing transaction references remain intact' and the requirement for explicit user confirmation. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler. The core action is front-loaded, and each sentence earns its place: what it does, a key behavioral consequence, and the required precondition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter deletion tool, the description covers the action, the key behavioral effect on existing references, and the required user confirmation. No output schema exists, so return-value details are not described, but that is not a significant gap for this operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented. The description's confirmation requirement is largely redundant with the schema's confirm parameter description ('Must be true. Only set after the user has explicitly confirmed this deletion.'). Thus the description adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific action ('Deactivate a company currency') and a specific resource, clearly distinguishing this tool from siblings like create_company_currency, get_company_currency, search_company_currencies, and update_company_currency. 'Deactivate' also clarifies the semantic difference from a hard delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the core use case and adds a critical precondition: 'Requires explicit user confirmation.' It does not explicitly enumerate when to avoid using it or name alternatives, but the operation is distinct enough from sibling tools that the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_credit_card_payment_txndelete credit card payment txnA
Destructive
Inspect

Permanently delete a credit card payment txn from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CreditCardPaymentTxn to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds the crucial context that deletion is permanent and requires explicit user confirmation. This goes beyond the annotation by informing the agent of irreversibility and the need for human consent, which is vital for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that conveys the core action, scope, and critical prerequisite (confirmation) with no unnecessary words. It is front-loaded with the verb and object.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive delete operation, the description covers permanence and confirmation, and the schema covers all parameters. It does not mention cascading effects or prerequisites (e.g., whether the transaction must be voidable first), but given the annotations and simple parameter set, this is largely sufficient. The lack of an output schema is not an issue here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides full descriptions for all three parameters (100% coverage), including the confirm flag's requirement to be true. The tool description reinforces the confirmation requirement but does not add new meaning beyond the schema. Baseline of 3 is appropriate given high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (permanently delete) on a specific resource (credit card payment txn) within QuickBooks Online, which clearly distinguishes it from other delete tools for different entities. The addition of the confirmation requirement further clarifies the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose is clear, and the resource type is unambiguous. However, it does not explicitly mention alternatives or exclusions, though the name and description make it evident this tool is for credit card payment transactions only. There is no need to list every other delete tool, but a note like 'for other transaction types, use the corresponding delete tool' would be explicit. It's adequate but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_credit_memodelete credit memoA
Destructive
Inspect

Permanently delete a credit memo from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CreditMemo to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as destructive and non-read-only, and the description consistently reinforces this with 'permanently delete' and 'requires explicit user confirmation.' It adds context beyond the annotations by emphasizing irreversibility and the need for consent, though it does not detail cascade effects or recovery limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The core action and the critical confirmation requirement are front-loaded, and every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema, the description provides the essential behavioral facts: permanent deletion, target system, and user-confirmation requirement. It could mention side effects or prerequisites for deletion, but the annotations plus description cover the main safety-critical context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters (id, company, confirm) are already fully documented. The description adds no new parameter-level meaning beyond restating the confirmation requirement already present in the confirm field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('delete'), resource ('credit memo'), and system ('QuickBooks Online'), clearly distinguishing it from sibling delete_* tools like delete_invoice or delete_vendor_credit. The word 'permanently' adds important scope beyond the title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to choose this tool versus alternatives such as voiding, updating, or searching credit memos. The only condition mentioned is explicit user confirmation, which is a prerequisite rather than a usage-vs-alternative guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_depositdelete depositA
Destructive
Inspect

Permanently delete a deposit from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Deposit to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal destructiveHint=trueaine, but the description adds meaningful context by emphasizing that deletion is permanent and requires explicit user confirmation. This goes beyond the annotation's binary destructive flag and helps the agent treat the operation as irreversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no wasted words. The core action is front-loaded, and the safety-critical confirmation requirement is stated immediately afterward.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a delete operation, the essential contextual information is the permanence of the action and the need for user confirmation, both explicitly stated. With full schema coverage for parameters and destructiveHint annotations, nothing else is needed to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all three parameters. The description reinforces the confirm requirement but does not add new meaning beyond what the schema provides; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Permanently delete a deposit from QuickBooks Online.' This clearly distinguishes delete_deposit from the many other delete_* siblings and identifies both the target entity and the system.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly communicates the key usage precondition: 'Requires explicit user confirmation.' This gives the agent actionable guidance about when it is appropriate to invoke the tool, though it does not name alternatives or explicitly state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_estimatedelete estimateA
Destructive
Inspect

Permanently delete an estimate from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Estimate to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, so the description's 'permanently delete' reinforces that. But the description adds critical behavioral context: the deletion is permanent (irreversible), and explicit user confirmation is required. This is exactly the kind of high-stakes behavior an agent needs to know before invoking. It also implies real-world consequences (data loss) that annotations alone don't convey. No contradiction with annotations; the description aligns with destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The most important facts — permanent deletion and user confirmation — are front-loaded. Could potentially add more behavioral context, but for what it contains, it's efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for a destructive operation: it states permanent deletion and the confirmation requirement. The output schema is absent, but for a delete operation the return value is usually a success indicator, which the agent can infer. The sibling list includes void_* tools (void_bill_payment, void_invoice, void_sales_receipt) which are alternative non-permanent operations, and the description doesn't explicitly guide the agent to them; that's a minor gap but not critical for a permanent delete tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the schema already thoroughly documents 'id' as the QuickBooks Id, 'company' as which connected company, and 'confirm' as must be true. The description doesn't add much parameter-level meaning beyond the schema. The 'require explicit user confirmation' in the description adds slight emphasis to the 'confirm' parameter's importance, but the schema already covers it. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('permanently delete'), a specific resource ('an estimate from QuickBooks Online'), and includes the key requirement of explicit user confirmation. The title is just 'delete estimate', but the description adds the permanent, irreversible nature and the QuickBooks context, making the tool's purpose unambiguous. The sibling list includes many delete_* tools, but the description names the exact resource (estimate), which distinguishes it from delete_bill, delete_invoice, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says the delete is permanent and requires explicit user confirmation, which gives clear guidance on when NOT to call it (without user confirmation) and implies it should only be used after user consent. It doesn't name an alternative (e.g., void vs delete), but the context signals are clear enough — this is a permanent delete, versus void_* siblings which might be the non-permanent alternative. The 'requires explicit user confirmation' is a strong usage guideline even though it doesn't enumerate alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_inventory_adjustmentdelete inventory adjustmentA
Destructive
Inspect

Permanently delete an inventory adjustment from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the InventoryAdjustment to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true, and the description adds value by stating the deletion is permanent and that explicit user confirmation is required. This extra context reinforces the confirm parameter's role and warns the agent that the action is irreversible, which is important behavioral guidance beyond the annotation alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, both informative and front-loaded. The first names the operation and its scope; the second provides the critical safety condition. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-entity delete tool with 100% schema coverage, destructive annotation, and no nested objects, the description covers the essential operational and safety context. It could add detail about what the response contains or how to confirm success, but the tool is simple enough that this is not a critical gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (id, company, confirm) are already documented in the schema. The description's statement that the action 'Requires explicit user confirmation' aligns with the confirm parameter but adds little new semantic meaning beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('delete'), a specific resource ('inventory adjustment'), and the target system ('QuickBooks Online'), and makes the destructive nature explicit with 'Permanently'. This clearly distinguishes it from the many sibling delete_* tools, which act on different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this tool when the user wants to delete an inventory adjustment. However, the description does not explicitly say when not to use it or mention any alternative such as void_inventory_adjustment or get_inventory_adjustment for verification. The confirmation requirement hints at the decision gate but does not provide deeper routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_invoicedelete invoiceA
Destructive
Inspect

Permanently delete an invoice from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Invoice to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint: true and readOnlyHint: false, so the destructive nature is covered. The description adds 'Permanently' and 'Requires explicit user confirmation,' which go beyond the annotation flags. This informs the agent that the deletion is irreversible and requires explicit user consent, adding meaningful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of two short sentences with no unnecessary words. It front-loads the core action and includes the critical requirement of confirmation. Every word serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a delete tool with a simple schema and no output schema, the description is adequate but not complete. It omits any mention of alternatives (like void_invoice) or potential side effects on linked records. Given the destructive nature and the existence of a void alternative, the description could be more informative. It meets the minimum viability but leaves gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (id, company, confirm) have their own descriptions. The tool description does not add any additional parameter semantics beyond what the schema already provides. Baseline of 3 is appropriate because the schema carries the full burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'delete' and the resource 'invoice', and adds 'Permanently' to convey irreversibility. This distinguishes it from the sibling void_invoice, which likely voids rather than deletes. The purpose is unambiguous and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide guidance on when to use this tool versus alternatives. It does not mention void_invoice or update_invoice as alternatives for less permanent actions. The only requirement mentioned is user confirmation, which is a precondition, not a usage guideline. No when-not or alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_journal_entrydelete journal entryA
Destructive
Inspect

Permanently delete a journal entry from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the JournalEntry to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds 'permanently delete' and 'requires explicit user confirmation,' which are meaningful behavioral traits beyond the annotations' destructiveHint and readOnlyHint. It reinforces irreversibility and user-consent requirements, though it does not mention potential side effects or permission prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load the core purpose ('Permanently delete a journal entry') and then state the confirmation requirement. No filler words or redundant information; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple deletion tool with full schema coverage and destructiveHint annotation, the description adequately covers purpose and key behavioral requirements. It does not explain return values or error behavior, but the lack of an output schema and the simplicity of the operation make this a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (id, company, confirm) are already documented. The description's mention of explicit user confirmation essentially mirrors the schema's confirm parameter description and adds no additional parameter-specific meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific verb 'delete', the resource 'journal entry', and the system 'QuickBooks Online', with 'permanently' clarifying the destructive nature. This clearly distinguishes it from siblings like create_journal_entry, update_journal_entry, get_journal_entry, and search_journal_entries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as update_journal_entry for modifications or void operations for reversible actions. The only hint is the verb 'delete', which implies purpose but does not explicitly state selection criteria or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_paymentdelete paymentA
Destructive
Inspect

Permanently delete a payment from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Payment to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, and the description reinforces and extends this by stating the deletion is permanent and requires explicit user confirmation. It adds the human-consent guardrail beyond what the annotations alone convey, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, information-dense sentences. The core action is front-loaded, and the safety requirement follows immediately. Every word contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-resource delete operation, the description, combined with full schema coverage, covers the essential facts: what is being deleted, that it is permanent, and that confirmation is mandatory. The lack of usage-vs-alternative guidance is the main gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a clear meaning in the schema. The description's mention of explicit user confirmation aligns with the 'confirm' parameter but adds no new semantic detail beyond what the schema documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('delete'), a specific resource ('a payment'), and a specific environment ('QuickBooks Online'). The word 'Permanently' distinguishes it from sibling tools like void_payment, and the payment resource distinguishes it from delete_bill_payment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explain when to choose this tool over alternatives such as void_payment or search_payments. It mentions the confirmation requirement, but that is a safety prerequisite, not guidance about when this tool should be used versus its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_purchasedelete purchaseA
Destructive
Inspect

Permanently delete a purchase from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Purchase to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint: true, and the description adds 'Permanently delete,' which clarifies irreversibility beyond the annotation's generic destructiveness. It also states the confirmation requirement, adding behavioral context. This aligns with annotations and adds meaningful detail, so the description adds value beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary action and scope, followed by the critical confirmation requirement. Every word contributes value; there is no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete operation with three parameters and no output schema, the description covers the essential purpose and a key prerequisite (confirmation). It does not specify return values, but that is not expected without an output schema. It is sufficiently complete for an agent to invoke correctly, though it could mention the irreversibility explicitly (it does via 'permanently').

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter (id, company, confirm) is documented. The description mentions 'explicit user confirmation,' which reinforces the confirm parameter but does not add new semantic meaning for id or company. The baseline for 100% coverage is 3, and the description provides marginal additional context about confirmation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Permanently delete a purchase from QuickBooks Online.' This is precise, identifies the target entity (purchase), and includes the scope (QuickBooks Online). It clearly distinguishes from other delete_* tools by naming the resource type, so an agent can select it without confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives one key prerequisite: 'Requires explicit user confirmation.' This is a clear condition for usage. However, it does not provide guidance on when to use this tool versus alternatives (e.g., void vs. delete for other entities) or mention any exclusions. The usage context is implied by the purpose, but not explicitly differentiated from siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_purchase_orderdelete purchase orderA
Destructive
Inspect

Permanently delete a purchase order from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the PurchaseOrder to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false, and the description adds that the deletion is permanent and requires explicit user confirmation. This is valuable behavioral context beyond the structured annotations, and there is no contradiction with destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core action and target are front-loaded, and the confirmation requirement is stated immediately after, so an agent can parse the essential information quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter destructive tool with fully documented parameters and a destructiveHint annotation, the description provides everything needed to select and invoke it: the action, resource, platform, permanence, and required user confirmation. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so id, company, and confirm are already fully explained in the schema. The tool description's mention of user confirmation mirrors the confirm parameter's schema description but adds no new parameter-level detail, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb, a concrete resource ('a purchase order'), and the platform ('QuickBooks Online'), so there is no ambiguity about the operation. The resource is named exactly as in the tool, which distinguishes it from sibling delete tools for other object types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly implies when the tool is appropriate: when a purchase order must be permanently removed in QuickBooks Online. The explicit-confirmation requirement is a concrete precondition that tells the agent to gate the call on user approval. It does not explicitly contrast with alternatives, but the resource-specific name makes the primary alternative choice obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_recurring_transactiondelete recurring transactionA
Destructive
Inspect

Permanently delete a recurring transaction from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the RecurringTransaction to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the description is not burdened by stating basic safety. It adds valuable behavioral context by emphasizing permanence and requiring explicit user confirmation, which are important for an irreversible mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, direct sentences with no filler. The key information (permanence and confirmation requirement) is front-loaded and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple destructive operation with full schema coverage and strong annotations, the description captures the essential prerequisites (permanence, user confirmation). It does not describe success/error outcomes, but no output schema exists and the operation's core behavior is sufficiently communicated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents id, company, and confirm. The description's mention of 'explicit user confirmation' reinforces the confirm parameter's requirement, but it does not add new parameter-level meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('delete') and resource ('recurring transaction from QuickBooks Online'), clearly distinguishing it from the many sibling delete_* tools. Adding 'Permanently' clarifies the operation's scope and irreversibility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states the operation targets a recurring transaction and that explicit user confirmation is required before deletion. However, it does not explicitly mention alternatives (e.g., deactivating rather than deleting, or using search/get tools first) or give when-to-use versus when-not-to-use guidance among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_refund_receiptdelete refund receiptA
Destructive
Inspect

Permanently delete a refund receipt from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the RefundReceipt to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal destructiveHint=true and readOnlyHint=false. The description adds valuable context beyond the annotations: the deletion is permanent and requires explicit user confirmation. This enriches the safety profile for the agent without contradicting the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler. The action is stated first, followed by the key safety constraint. Every word earns its place, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with full schema coverage, destructive annotations, and no output schema, the description covers the essential operational facts: permanent deletion, the entity, and the confirmation requirement. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: id, company, and confirm all have descriptive text in the schema. The description adds no additional parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb ('delete') and resource ('refund receipt') and adds the critical qualifier 'permanently'. This clearly distinguishes the operation from send/void/update variants and makes the tool's purpose unambiguous even without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the core condition for use: the user must have explicitly confirmed deletion. It also communicates that this is a permanent removal. However, it does not explicitly name alternatives (e.g., voiding) or state when not to use the tool, so it falls short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_sales_receiptdelete sales receiptA
Destructive
Inspect

Permanently delete a sales receipt from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the SalesReceipt to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the description adds value by specifying 'permanently' and emphasizing 'requires explicit user confirmation'. This goes beyond the annotation by disclosing the irreversibility and the confirmation requirement, which is not captured elsewhere. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundant information. The action and key constraint (permanence and confirmation) are front-loaded. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple delete operation given the annotations and full schema coverage. However, it lacks any reference to the alternative void_sales_receipt and does not explain the consequences of deleting (e.g., impact on linked transactions). This is a notable gap for a destructive action with a closely related sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% description coverage for all parameters, so the description adds little beyond that. The phrase 'explicit user confirmation' hints at the confirm parameter but does not elaborate on its semantics beyond what the schema already states. This meets the baseline for well-documented schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'delete' and resource 'sales receipt', and adds 'permanently' to indicate severity. However, it does not explicitly distinguish from the sibling void_sales_receipt, which is a common alternative for sales receipts. The purpose is clear but lacks explicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like void_sales_receipt. It only states a requirement (explicit user confirmation) but does not explain the conditions under which deletion is appropriate or when voiding is preferred. There is no mention of alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_time_activitydelete time activityA
Destructive
Inspect

Permanently delete a time activity from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the TimeActivity to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds value by specifying that deletion is permanent and requires explicit user confirmation, which goes beyond the mere destructive hint and clarifies the irreversible nature of the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler. The primary action is front-loaded, and the confirmation requirement is stated immediately. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a delete tool with annotations and fully described parameters, the description covers the essential behavior: what is deleted, that it is permanent, and that confirmation is required. It does not explain return values, but no output schema exists and it is not critical for a delete operation. It is adequate but could mention potential side effects, though 'permanently' largely covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all three parameters, so the baseline is 3. The description does not add extra meaning beyond the schema; it simply reiterates the confirmation requirement, which is already documented in the confirm parameter description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('delete'), a specific resource ('time activity'), and the context ('from QuickBooks Online'). It also adds 'Permanently' and 'Requires explicit user confirmation,' which distinguishes it from update or void operations and from sibling tools like update_time_activity and create_time_activity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (when you want to permanently delete a time activity) and includes a precondition ('Requires explicit user confirmation'), but it does not explicitly compare with alternatives or state when not to use it. There is no mention of when to prefer delete over update or void, so guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_transferdelete transferA
Destructive
Inspect

Permanently delete a transfer from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Transfer to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true and readOnlyHint=false, and the description adds meaningful beyond-annotation context by stating the deletion is permanent and requires explicit user confirmation. This gives the agent crucial operational framing without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two short sentences with no filler. The core behavior and the critical confirmation precondition are front-loaded, and every word contributes to the agent's ability to invoke the tool safely.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter destructive operation, the description, combined with the full schema coverage and annotations, is sufficiently complete. It covers the essential deletion semantics and the confirmation gate, though it does not describe success/failure response behavior since no output schema is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with clear descriptions for id, company, and confirm. The description reinforces the confirmation requirement but adds no syntax, format, or parameter-specific details beyond what the schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Permanently delete') on a specific resource ('a transfer from QuickBooks Online'), making the tool's purpose unmistakable. It distinguishes the tool from create_transfer, update_transfer, and search_transfers by focusing on deletion and permanence. No tautology or ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the appropriate usage context: when a transfer must be permanently removed and the user has confirmed the action. However, it does not explicitly name alternatives or state when this tool should not be used. The confirmation requirement provides useful procedural context but not full alternative-routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_vendor_creditdelete vendor creditA
Destructive
Inspect

Permanently delete a vendor credit from QuickBooks Online. Requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the VendorCredit to delete.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this deletion.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The destructiveHint annotation already marks the tool as destructive, and the description adds meaningful context beyond it: the deletion is permanent and requires explicit user confirmation. This gives an agent important behavioral information about irreversibility and the need for a confirmation step.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with zero wasted words. The primary action and permanence are front-loaded, and the confirmation requirement is stated immediately after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple destructive operation with three well-documented parameters and the destructiveHint annotation, the description is largely complete. It states the resource, permanence, and confirmation prerequisite; a fully complete description might also note that the deletion cannot be undone, but 'Permanently' already implies this.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a clear description, especially `confirm`, which states it must be true and only set after explicit user confirmation. The tool description itself adds no additional parameter-level meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('delete'), a specific resource ('vendor credit'), and a clear scope ('from QuickBooks Online'). The word 'Permanently' adds an important precision that distinguishes this destructive action from a void or reversible operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear, explicit prerequisite: this should only be invoked after the user has explicitly confirmed the deletion. It does not enumerate alternatives or exclusions, but the resource-specific naming and confirmation gate make the intended usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discard_inbox_fileDiscard a file from the receipt inboxA
Destructive
Inspect

Permanently remove a user-selected file from the Caribooks document box without attaching it to QuickBooks. Intended for unwanted documents, duplicates and unrelated files. Requires confirm: true after the user's approval. Retains the document's filing-history row marked discarded.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue once the user has agreed to drop this file.
upload_idYesThe file's id, as listed by list_receipt_inbox.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses that removal is permanent, that the filing-history row is retained and marked discarded, and that user confirmation via confirm: true is required. This gives the agent meaningful behavioral context for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences front-load the core action, then add use cases, the confirmation requirement, and the retention side effect. Every sentence contributes useful information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter destructive tool, the description is complete: it explains what is removed, when to use it, what must be provided, and what residual record remains. The schema covers parameter source for upload_id, and no output schema is needed for this operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both upload_id and confirm. The description repeats the confirm requirement but adds no parameter-specific detail beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Permanently remove a user-selected file from the Caribooks document box.' It also distinguishes itself from related flows by noting it happens 'without attaching it to QuickBooks' and by naming the intended cases: unwanted documents, duplicates, and unrelated files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool ('Intended for unwanted documents, duplicates and unrelated files') and the required approval gate ('Requires confirm: true after the user's approval'). It does not explicitly name alternatives or when-not-to-use conditions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_expenses_without_attachmentsFind expenses without attachmentsA
Read-only
Inspect

List purchases and bills with no document attached in QuickBooks. Compares transactions against all attachments on the server. Returns compact rows with type, id, date, amount, currency, payee, payment account, expense category, memo and document number, newest first, plus documented and undocumented counts. Some categories, such as payroll, tax remittances, transfers and bank fees, may not require receipts. A window too wide to fit is truncated from the oldest end; from and to limit the date range.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoLatest transaction date, YYYY-MM-DD. Defaults to no upper bound.
fromNoEarliest transaction date, YYYY-MM-DD. Defaults to twelve months ago.
typesNoPurchase (anything paid from a bank or card account, including bank-feed lines), Bill (payables), or both. Defaults to both.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only and non-destructive, and the description adds substantial behavioral detail: it compares against all attachments, returns compact rows with a specific field list, orders newest first, reports documented/undocumented counts, and truncates oversized windows from the oldest end. This goes well beyond the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized: purpose first, then methodology, output shape, a practical caveat, and truncation behavior. Every sentence adds distinct value and there is no redundant or fluffy wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, the description compensates by enumerating returned fields, ordering, counts, and truncation behavior. It also warns about legitimate receipt-free categories. Combined with complete schema parameter docs and safety annotations, nothing essential is missing for an agent to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes all four parameters with 100% coverage. The description adds only that 'from' and 'to' limit the date range, which is already implied by the schema descriptions. It does not introduce new parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List purchases and bills with no document attached in QuickBooks.' It clearly distinguishes itself from sibling search tools like search_purchases and search_bills by focusing on missing attachments across two transaction types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the tool's context clear: it is for finding expenses without receipts, and it adds a caveat that some categories may not require receipts. It does not explicitly name alternatives or say when not to use it, but the unique attachment-filtering purpose is evident from the description and name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_accountget accountA
Read-only
Inspect

Get a single account from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Account.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds no further behavioral context (e.g., auth requirements, error behavior, or return shape), but for a simple read-only get operation, it does not contradict the annotations. It meets the baseline but adds no extra value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence that is front-loaded with the verb and resource, then the identifier criterion. Zero waste, ideal structure for a simple get tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read-only tool with full schema coverage and safety annotations, the description is complete. It could mention the return object explicitly, but that is implicit in a 'get' operation. Minor gap: no mention of not-found behavior, but that is likely handled by the platform. Overall adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both id and company already described. The description merely repeats 'by Id' without adding syntax, format, or interplay details beyond the schema. Per the rubric, baseline 3 applies when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a specific resource ('account'), and a specific criterion ('by Id'), which clearly distinguishes it from sibling tools like get_account_list (list all) and search_accounts (search). No ambiguity about what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you have a specific account Id, but it does not explicitly mention alternatives like get_account_list or search_accounts, nor when to avoid this tool. The distinction is inferable from the name and description, but not stated outright, so guidance is present but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_account_listget account listA
Read-only
Inspect

Generate the QuickBooks account list report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: account_type, account_status (Deleted|Not_Deleted), start_moddate/end_moddate or moddate_macro, createdate_macro; columns (account_name, account_type, detail_acc_type, account_bal, account_desc, account_cur, create_date, last_mod_date), sort_by, sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds substantial behavioral detail beyond that: row types, depth, column alignment, label vs value id alignment, omission and counting of all-zero rows, filter confirmation semantics, and the no_data flag. This gives an agent an unusually clear model of what the tool returns and how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place, covering purpose, output structure, edge cases, and behavioral nuances. It is front-loaded with the primary action and contains no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by detailing the return object, row semantics, id alignment, hidden-zero handling, filter confirmation, and no_data behavior. Combined with the input schema and annotations, an agent has everything needed to call the tool and interpret its response correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the parameters. The description adds param-relevant meaning by explaining that the returned filters list only includes parameters QuickBooks actually applied, and that ignored parameters are absent. This helps an agent interpret parameter effects even though the description does not re-explain individual params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks account list report.' This clearly distinguishes it from sibling tools like get_account (single account) and search_accounts (search results), and the return shape confirms it is a report-producing tool, not a simple lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: call this when you need the QuickBooks account list report. However, it does not explicitly state when to prefer this over alternatives such as get_account or search_accounts, nor does it provide exclusion conditions. The usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_aged_payable_detailget aged payable detailA
Read-only
Inspect

Generate the QuickBooks aged payable detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); aging_method (Report_Date|Current), aging_period (days per bucket, default 30), num_periods (number of buckets, default 4), past_due (minimum days overdue); start_duedate, end_duedate (YYYY-MM-DD); vendor, term, shipvia; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only and non-destructive, but the description adds rich behavioral detail: return shape, row types, hidden zero-row handling, filters that QuickBooks confirms, and no_data semantics. This goes well beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: report purpose, return structure, row semantics, id alignment, hidden-zero accounting, filter confirmation, and no-data handling. It front-loads the main purpose and then structures the behavior logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by explaining the return value, row types, id alignment, omitted rows, filter confirmation, and empty-period behavior. Nothing necessary to invoke or interpret the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers parameters thoroughly (100% coverage), so the baseline is 3. The description adds value beyond the schema by explaining that the output's filters field reveals which parameters QuickBooks actually applied, and that ignored parameters are absent. This gives agents useful feedback about parameter behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks aged payable detail report.' It clearly identifies what the tool produces and distinguishes it from sibling report tools by naming the exact report and report type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like get_aged_payables or get_aged_receivable_detail. The description states what the tool does but does not mention when-not-to-use or provide a routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_aged_payablesget aged payablesA
Read-only
Inspect

Generate the QuickBooks aged payables report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); aging_method (Report_Date|Current); vendor, department.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without an output schema, the description fully explains the return shape: columns, column_types, rows, row types, depth, id alignment, hidden_zero_rows, filters, and no_data. It also discloses behavior like omitted zero rows and how errors/ignored parameters surface (absent from filters). This goes far beyond the readOnly hint and is exception detail for a reporting tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main action is front-loaded, and every sentence contributes operational detail. The description is dense but not verbose; it packs the essential output contract into a compact paragraph without filler. Structured enough for an agent to parse the key facts in order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description provides a complete contract: it explains the row schema, value types, id semantics for summary vs detail, zero-row handling, filter verification, and no_data edge case. An agent can invoke this tool and interpret its response without needing to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the params object already names report_date, date_macro, aging_method, vendor, department. The description adds no new meaning about how to use these parameters beyond what the schema states; the extra details about filters relate to output behavior, not parameter semantics. Baseline 3 is appropriate because the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb-resource pair ('Generate the QuickBooks aged payables report') and goes on to detail the output structure. It distinguishes this tool from related siblings like get_aged_payable_detail by describing both summary and detail row semantics, and the name itself disambiguates payables from receivables.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells you what the tool does but does not explicitly state when to choose it over get_aged_payable_detail or other aged reports. The sibling list hints at an alternative, but there is no direct guidance such as 'use this for the summary; use get_aged_payable_detail for line-level detail.' Usage context is implied from the name and report type but not made explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_aged_receivable_detailget aged receivable detailA
Read-only
Inspect

Generate the QuickBooks aged receivable detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); aging_method (Report_Date|Current), aging_period (days per bucket, default 30), num_periods (number of buckets, default 4), past_due (minimum days overdue); start_duedate, end_duedate (YYYY-MM-DD); customer, term, shipvia; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal read-only and non-destructive behavior, and the description adds substantial context beyond that: row types, depth, value alignment, id semantics, hidden_zero_rows, filters that QuickBooks confirms applied, and no_data behavior. This is rich, honest behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it opens with the tool's purpose, then explains the return format, and then covers edge cases. Every sentence carries useful information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining the return shape, and it does so thoroughly: columns, row types, ids, hidden zero rows, filter confirmation, and no_data. Combined with the schema's parameter documentation, an agent has what it needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds no parameter-level meaning beyond what the schema provides, which meets the baseline but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks aged receivable detail report.' This distinguishes it from siblings like get_aged_receivables and get_aged_payable_detail by naming the exact report being produced.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to choose this tool over alternatives. It does not compare itself to get_aged_receivables, get_aged_payable_detail, or other report tools, leaving selection largely to inference from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_aged_receivablesget aged receivablesA
Read-only
Inspect

Generate the QuickBooks aged receivables report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); aging_method (Report_Date|Current); customer, department.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only, and the description adds substantial behavioral context: exact return shape, hidden_zero_rows counting, filters listing only confirmed applied parameters, and the no_data flag. It also discloses that ignored parameters are absent from filters, which is valuable beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded and the description is dense with useful return-format details. It's slightly run-on and the 'detail reports' aside could confuse, which prevents a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description fully specifies the return contract: columns, column_types, rows, row types, depth, money representation, id semantics, hidden rows, filters, and no_data. The only small gap is explaining aging_method values, which is minor given the schema's description coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by explaining that a parameter QuickBooks ignored will be absent from filters, clarifying the effective behavior of customer/department/aging_method parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Generate') and resource ('QuickBooks aged receivables report'). It doesn't explicitly distinguish itself from siblings like get_aged_receivable_detail or get_aged_payables, so it's clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus the many report siblings. It mentions detail reports but doesn't instruct the agent to use get_aged_receivable_detail for detail, and it provides no exclusions or prerequisite conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_attachableget attachableA
Read-only
Inspect

Get a single attachable from QuickBooks Online by Id. A file attachment's TempDownloadUri is a short-lived download link to the original file.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Attachable.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this as read-only and non-destructive. The description adds value by disclosing that TempDownloadUri is a short-lived download link to the original file, which is a meaningful behavioral detail an agent should know before relying on the URI. No contradictions with annotations are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The first sentence identifies the exact operation and scope, and the second adds a relevant caveat about the download URL. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only fetch-by-id tool, the description, annotations, and complete schema coverage are sufficient for an agent to invoke it correctly. The TempDownloadUri note adds the most important return-behavior context; a fuller response shape description would be nice but is not essential here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents both `id` and `company`. The description's mention of 'by Id' aligns with the required parameter but does not add semantic detail beyond the schema, which is acceptable at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), resource ('a single attachable'), and lookup method ('by Id'), making it clear this is a point-read operation. It also adds useful domain context about TempDownloadUri, and the singular 'single attachable' distinguishes it from search_attachables and create/update/delete siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly conveys when to use this tool: when you already have a QuickBooks Attachable Id and need the single attachable record. It does not explicitly name alternatives like search_attachables for listing/filtering, so it falls short of a 5, but the 'single ... by Id' framing provides clear usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_balance_sheetget balance sheetA
Read-only
Inspect

Generate the QuickBooks balance sheet report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); adjusted_gain_loss (true|false).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description richly discloses behavior beyond the readOnlyHint/destructiveHint annotations: row types, depth, Money cell representation, id alignment on summary vs detail reports, hidden_zero_rows, the filters field reflecting confirmed/applied parameters, and no_data semantics. This gives the agent a strong model of what the tool returns and how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's purpose and then packs complex return-format details into two dense sentences with no filler. Every clause adds needed information about row structure, money values, ids, hidden rows, filters, and no_data. It is appropriately sized for a report with no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description thoroughly explains the return shape and edge cases, including id alignment, hidden_zero_rows, filters, and no_data. Combined with 100% parameter schema coverage and safety annotations, the description provides everything needed to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds context about the returned filters field showing which parameters were applied, but it does not add new meaning to the input parameters themselves. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Generate the QuickBooks balance sheet report.' It clearly states what the tool does. However, it does not explicitly distinguish itself from the sibling tools get_balance_sheet_detail and get_balance_sheet_summary, so it is clear but lacks direct sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by 'Generate the QuickBooks balance sheet report,' but the description does not state when to use this tool versus alternatives like get_balance_sheet_detail or get_balance_sheet_summary. There are no exclusions or alternative-selection conditions, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_balance_sheet_detailget balance sheet detailA
Read-only
Inspect

Generate the QuickBooks balance sheet detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); account, account_type; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly/destructive hints; the description goes well beyond them by explaining the row type/depth structure, how ids align to label cells vs values, hidden_zero_rows counting omitted rows, the filters confirmation behavior including ignored parameters, and the no_data edge case. This is rich behavioral context that significantly informs invocation and result interpretation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place: it front-loads the return shape before diving into row semantics, id alignment, hidden rows, and filter/no_data behavior. No filler or repetition is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries full responsibility for explaining the return value, and it does so thoroughly: columns, row types, depth, id alignment, hidden rows, filter reporting, and no_data. For a read-only report tool with simple parameters, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents the parameters thoroughly at 100% coverage, so the baseline is 3. The description adds extra semantic value by explaining that the filters output reflects only parameters QuickBooks actually applied and that an ignored filter parameter will be absent, plus what no_data means for the report period. This goes beyond the schema's static parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and resource ('the QuickBooks balance sheet detail report'), and the return-shape details make the tool's purpose concrete. However, it does not explicitly distinguish itself from siblings like get_balance_sheet or get_balance_sheet_summary, leaving the agent to infer the difference from names and output details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance about when to use this tool versus alternatives, such as get_balance_sheet_summary for summaries or get_trial_balance for a different view. The agent must infer the appropriate context solely from the tool name and the output description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_balance_sheet_summaryget balance sheet summaryA
Read-only
Inspect

Generate the QuickBooks balance sheet summary report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal read-only safety, but the description goes well beyond that by disclosing return shape, row types, the meaning of id, the difference in id alignment between summary and detail reports, hidden all-zero rows, filters QuickBooks actually applied, and the no_data flag. These are meaningful behaviors an agent would otherwise have to guess.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries necessary information. It front-loads the core purpose and then packs return-format details into a compact, structured set of clauses with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining return values, and it does so thoroughly: columns, row types, depth, value alignment, ids, hidden_zero_rows, filters, and no_data. Nothing essential for correctly invoking and interpreting this report tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents start_date, date_macro, accounting_method, summarize_column_by, filters, and company. The description adds no parameter-level detail beyond mentioning the filters field as part of the return object, so it earns the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and a specific resource ('the QuickBooks balance sheet summary report'). It clearly identifies this as the summary variant, and even contrasts summary-report behavior with detail-report behavior, differentiating it from siblings like get_balance_sheet_detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool's purpose as the summary version of the balance sheet is clear, and the mention of 'summary reports' vs 'detail reports' implies when each variant applies. However, it never explicitly says 'use this for summaries, use get_balance_sheet_detail for detailed line-item data,' so the usage guidance remains implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_billget billA
Read-only
Inspect

Get a single bill from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Bill.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safe read-only nature is covered. The description adds that it operates on QuickBooks Online and returns a single bill, but it does not disclose not-found behavior, authorization needs, or whether the full bill object is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundant qualifiers. Every word contributes to identifying the tool's action, target, and criteria.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter getter with 100% schema coverage and safety annotations, the description is largely sufficient. It does not explicitly describe the return value or error cases, but the tool's name and 'Get a single bill' make the primary outcome evident.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, with clear descriptions for 'id' and 'company'. The description adds no new parameter-level meaning beyond reiterating the Id-based lookup, so the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get'), a specific resource ('a single bill'), and a precise retrieval key ('by Id'). This clearly distinguishes it from sibling tools like search_bills, which perform broader lookups, and get_bill_payment, which targets a different entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by Id' implies this tool is for retrieving a known specific bill, but the description does not explicitly say when to use it versus search_bills or other getters. There is no explicit guidance about what to do when the Id is unknown or which company must be specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_bill_paymentget bill paymentA
Read-only
Inspect

Get a single bill payment from QuickBooks Online by Id. An accounting record of a payment to a vendor. It does not initiate a bank or card payment.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the BillPayment.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the bar is lower. The description adds valuable domain-specific context: it is an accounting record and does not initiate a bank or card payment. This goes beyond the generic read-only annotation by clarifying the real-world side effect that is NOT triggered, which helps prevent misuse.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary purpose is front-loaded, and the clarifying exclusions follow immediately. Every sentence earns its place—nothing extraneous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-Id tool with fully documented schema parameters and no output schema, the description provides all essential context: what it retrieves, how it identifies it, and what it does not do. Combined with annotations indicating safety, an agent has enough to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'id' and 'company' already described. The description only restates that retrieval is 'by Id', adding no new meaning to the parameters. Per the baseline rule, when the schema fully documents parameters, score 3, and the description does not elevate it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence uses a specific verb and resource: 'Get a single bill payment from QuickBooks Online by Id.' It clearly identifies the object (a bill payment, as opposed to a general payment) and method (by Id). The second sentence adds identity as 'an accounting record of a payment to a vendor,' distinguishing it from customer payments and from tools like create_bill_payment or search_bill_payments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states 'It does not initiate a bank or card payment,' which is an explicit exclusion that routes agents away from using this tool for payment initiation. It also implies retrieval use via 'Get ... by Id' without naming alternatives like search_bill_payments. This is clear context but lacks an explicit 'use X instead when' statement, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_budgetget budgetA
Read-only
Inspect

Get a single budget from QuickBooks Online by Id. A profit-and-loss budget with one amount per account per period.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Budget.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the structural detail about the budget being profit-and-loss with one amount per account per period, but does not address not-found behavior, permissions, or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the first sentence leads with verb and resource, while the second adds a valuable structural clarification. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id tool with annotations covering read-only safety and a fully described schema, the description is nearly sufficient. It could mention what happens when no budget is found or note that company selection is optional, but the missing details are minor given the low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters have meaningful descriptions. The description's 'by Id' reinforces the id parameter but adds no new meaning beyond what the schema already provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), resource ('a single budget from QuickBooks Online'), and key ('by Id'), and clarifies the budget kind, which distinguishes it from search_budgets and get_budget_vs_actuals. It leaves no doubt about what entity is returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for fetching one budget when its QuickBooks Id is known, but it does not specify when to prefer this over search_budgets or get_budget_vs_actuals, nor any exclusions. It gives usable context without explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_budget_vs_actualsget budget vs actualsA
Read-only
Inspect

Generate the QuickBooks budget vs actuals report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters. All four are required: budget (a Budget Id, from search_budgets), start_date, end_date (YYYY-MM-DD) and rowaxis ('primary' for accounts). Optional: summarize_column_by (Total|Month|Quarter|Year|Classes|Departments|Customers), accounting_method (Cash|Accrual). QuickBooks answers 'Permission Denied' when the budget id is missing or unknown, so read it with search_budgets first.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint annotation, disclosing the exact return shape, row types, depth, value alignment, Money-cell handling, label id behavior, omitted zero rows with hidden_zero_rows, filter confirmation semantics, and the no_data flag. This is exemplary transparency for an agent invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries meaningful behavioral or usage information. It front-loads the report purpose and then efficiently packs return format, edge cases, filter semantics, and no-data handling without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and absence of an output schema, the description is remarkably complete. It covers return shape, row/value semantics, ignored-parameter detection via filters, zero-row handling, no-data behavior, required parameters, and the prerequisite search_budgets call. An agent has everything needed to invoke it correctly and interpret its response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters and their requirements. The tool description itself mostly explains output behavior rather than adding new parameter semantics. It reinforces the budget id prerequisite and error behavior, but those details already appear in the schema description, so the description adds no significant beyond-schema parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks budget vs actuals report.' This clearly differentiates the tool from siblings like get_budget, get_cash_flow, and the other report getters, with no ambiguity about what the report contains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong context: required parameters, date formats, rowaxis expectations, and the prerequisite to use search_budgets first to obtain a valid budget id. It also warns about the 'Permission Denied' behavior when the budget id is missing or unknown. It does not explicitly name an alternative tool or state when not to use it, but the context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cash_flowget cash flowA
Read-only
Inspect

Generate the QuickBooks cash flow report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the readOnly/destructive annotations by disclosing the exact return structure, row type semantics, id alignment behavior, omission of zero rows, hidden_zero_rows count, filters behavior, and no_data flag. This gives the agent a detailed and honest model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description opens with the core purpose in one short sentence, then packs essential return-format details into a dense but well-ordered paragraph. Every sentence contributes meaningful information, especially valuable given there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by explaining columns, column_types, rows, row types, depth, values, id semantics, hidden_zero_rows, filters, and no_data. Nothing critical is missing for an agent to understand what the tool returns and how to interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both params and company. The description adds some context about how filters appear in the response, but it does not add significant new meaning to the input parameters themselves.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Generate the QuickBooks cash flow report.' The report type, clearly reflected in the tool name, differentiates it from sibling report tools such as get_balance_sheet or get_profit_and_loss.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states that this tool produces the cash flow report, giving the agent a direct basis for selecting it. It does not explicitly list when-not-to-use conditions or name alternatives, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_changesget changesA
Read-only
Inspect

List everything created, updated or deleted in QuickBooks Online since a given moment, for the named entities. Returns QuickBooks' CDCResponse, in which a record carrying status: Deleted was removed. QuickBooks only looks back 30 days.

ParametersJSON Schema
NameRequiredDescriptionDefault
sinceYesReturn everything changed since this moment (YYYY-MM-DD or an ISO timestamp).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
entitiesNoQuickBooks entity names to watch. Defaults to Invoice, Bill, Payment, BillPayment, Purchase, Estimate, SalesReceipt, JournalEntry, Customer, Vendor.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds valuable context: it explains the return type (CDCResponse) and that records with status: Deleted were removed, clarifying how deletions are represented. It also discloses the 30-day lookback limit. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose. It efficiently includes the return format, the deletion status nuance, and the 30-day limit without any fluff. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description explains the return value as CDCResponse and the meaning of the Deleted status. It also covers the time constraint and entity selection. While it doesn't mention pagination or record limits, these are likely not critical for this read-only change-listing tool, so the description is adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides descriptions for all three parameters (since, company, entities) with 100% coverage, so the schema does the heavy lifting. The description adds minimal extra meaning beyond restating 'since a given moment' and 'named entities,' which are already in the schema. It does not introduce any new parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'List everything created, updated or deleted in QuickBooks Online since a given moment, for the named entities.' This is a specific verb (list), resource (QuickBooks changes), and scope (since a time, for entities). It distinguishes itself from the many get_* and search_* siblings by focusing on change events rather than static data retrieval, and it explicitly mentions the CDCResponse format.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when you need to see changes since a moment, for specific entities. It also notes a key constraint: 'QuickBooks only looks back 30 days.' However, it does not explicitly state alternatives or when not to use it, but given the tool's unique nature among siblings, this is minor.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_classget classA
Read-only
Inspect

Get a single class from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Class.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds no behavioral context beyond the operation itself—nothing about not-found behavior, response shape, or company scoping—so it provides minimal added transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly worded sentence that front-loads the action and object. There is no filler, repetition, or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-record fetch with a fully described schema and read-only annotations, the description is nearly sufficient. The absence of an output schema and lack of not-found/error behavior are minor gaps but do not prevent a capable agent from invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear descriptions for both 'id' and 'company'. The description only restates the 'by Id' concept and does not add additional meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and names the exact resource ('a single class from QuickBooks Online by Id'), clearly distinguishing it from related siblings like search_classes and get_class_sales. An agent can immediately understand what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this when you have a QuickBooks class Id and need one specific class. However, the description does not explicitly state when not to use it or mention alternatives like search_classes for finding classes by criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_class_salesget class salesA
Read-only
Inspect

Generate the QuickBooks class sales report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the tool as read-only and non-destructive, so the description does not need to repeat that safety profile. It adds significant behavioral detail beyond annotations: it describes the exact return shape (columns, column_types, rows), edge cases (all-zero rows omitted, hidden_zero_rows), how ids align on summary vs detail reports, and the meaning of filters and no_data. This goes well beyond the minimal expectations and provides deep transparency, though it could clarify whether pagination or limits exist, which is not mentioned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense paragraph that front-loads the primary purpose and then efficiently details the return format. Each sentence contributes substantive information about behavior or output structure. It is not overly long and there is minimal wasted phrasing, though it could be broken into smaller sentences for easier scanning. Still, it is well within acceptable length and avoids redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (a report with nested structures), the description covers essential aspects: the action, the input parameters (via schema), the output format, edge cases, and error-like conditions. The absence of an output schema makes this explanation critical, and it does provide it. However, it lacks explicit mention of pagination limits or potential throttling, and doesn't clarify how errors (e.g., invalid date) are surfaced, but for this context it is quite complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides a comprehensive description of the params object, covering date formats, macros, accounting_method, summarize_column_by, and filters, with 100% coverage. The description does not add new parameter-specific details beyond what's already in the schema. It references 'filters' but without explaining them further, so it adds minimal value. Given the schema's thoroughness, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Generate the QuickBooks class sales report') and a specific resource (class sales). It distinguishes itself from related report tools (like get_customer_sales, get_item_sales) by starting with 'class sales,' making the purpose unambiguous. The detailed return format description also reinforces what the tool produces, setting it apart from simply listing data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions that filters are applied and that ignored parameters are absent from the output, which implies how to use it (e.g., if a filter isn't applied, it won't appear). However, it doesn't explicitly state when to use this tool over alternatives (e.g., versus get_customer_sales, get_item_sales) or provide exclusions. The sibling list shows many similar report tools, but no guidance on selecting among them is given, leaving the agent to infer based on the 'class sales' focus.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_company_currencyget company currencyA
Read-only
Inspect

Get a single company currency from QuickBooks Online by Id. The currencies the company transacts in; only meaningful with multicurrency enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CompanyCurrency.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds useful domain context about multicurrency, but it does not describe edge behavior such as an invalid Id or what happens when multicurrency is disabled. This is acceptable given the read-only annotation, but not richly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The primary action is front-loaded, and the domain clarification about multicurrency earns its place by preventing misuse in non-multicurrency companies.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id tool with two parameters and no nested objects, the description plus schema is sufficient for correct invocation. The lack of an output schema means return-value details are not described, but the resource semantics and identification-by-Id are clear enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'id' and 'company' clearly described in the input schema. The description adds minimal parameter-level meaning beyond restating lookup 'by Id,' so it meets the baseline for high schema coverage but does not elevate it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get a single company currency from QuickBooks Online by Id.' It also distinguishes itself from sibling search tools by emphasizing 'single' and lookup by Id, so an agent can tell it apart from search_company_currencies without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when the tool is relevant: CompanyCurrency is 'the currencies the company transacts in' and is 'only meaningful with multicurrency enabled.' It does not explicitly name alternatives like search_company_currencies, but the 'single ... by Id' phrasing implies the line between direct lookup and search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_company_infoget company infoA
Read-only
Inspect

Get a single company info from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoDefaults to the connected company.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description is consistent with those. However, the description adds no behavioral context beyond the annotations—no mention of error cases, return payload contents, or what happens when no id is provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundancy. It conveys action, resource, source system, and lookup key efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, low-complexity tool with no required parameters and fully described schema fields, the description plus schema is mostly sufficient. However, with no output schema, the description does not clarify what fields the returned 'company info' contains, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both id and company already described meaningfully (id defaults to the connected company; company selects which connection). The description adds little beyond 'by Id,' so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('company info') and scopes it precisely as 'a single company info ... by Id.' This clearly distinguishes it from sibling tools like search_company_infos (which lists/searches) and update_company_info (which mutates).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'by Id' phrasing implies this tool is for fetching one known company rather than searching, but it never explicitly names alternatives or conditions—such as when to prefer search_company_infos or how to handle the default company. Usage guidance is only implied, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_credit_card_payment_txnget credit card payment txnA
Read-only
Inspect

Get a single credit card payment txn from QuickBooks Online by Id. An accounting record of a bank payment toward a credit card balance. It does not move money.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CreditCardPaymentTxn.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the read-only, non-destructive nature. The description adds useful context beyond annotations: it explains the record type (accounting record of a bank payment) and explicitly states 'It does not move money,' which reinforces the read-only behavior and prevents misuse as a financial transaction tool. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences. The first sentence states the action and target, the second provides domain context and a critical clarification. No fluff, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id tool with two well-documented parameters and annotations covering the safety profile, the description is complete. It clarifies the domain meaning and the non-monetary nature. There is no output schema, but that is not a gap for this simple fetch operation. The only slight omission is an explicit pointer to alternatives, but that is not critical for a basic getter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – both 'id' and 'company' are clearly described in the schema. The description does not add any parameter-specific information beyond what the schema already provides. Since the schema carries the full burden, baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), resource ('credit card payment txn from QuickBooks Online'), and scope ('by Id'). It also clarifies the domain meaning as 'an accounting record of a bank payment toward a credit card balance' and adds 'It does not move money,' distinguishing it from financial operations and siblings like create/update/delete and search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Get a single ... by Id' implies this tool is for retrieving one specific record when an Id is known, but it does not explicitly mention alternatives such as search_credit_card_payment_txns for multiple records or filtering. The clarification that it does not move money hints at when not to use it for financial actions, but no direct comparison to siblings is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_credit_memoget credit memoA
Read-only
Inspect

Get a single credit memo from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CreditMemo.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds only that this is a single-record lookup by Id, which is useful but not deeply behavioral; it does not mention error behavior or return format, though for a simple read this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to the core purpose and retrieval mechanism.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-Id operation with fully documented parameters and read-only annotations, the description is nearly complete. It does not describe missing-id behavior, but that is a minor omission for such a well-scoped lookup tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both id and company. The description's 'by Id' adds no new meaning beyond what the schema states, so a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get'), a clear resource ('a single credit memo from QuickBooks Online'), and a retrieval method ('by Id'). It effectively distinguishes itself from sibling search_credit_memos, which is for finding credit memos without requiring a known Id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the tool is for retrieving one specific credit memo when an Id is already known. It does not explicitly name alternatives or exclusions, but the 'by Id' wording provides enough context to route the agent appropriately alongside search_credit_memos.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customerget customerA
Read-only
Inspect

Get a single customer from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Customer.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is established. The description adds the scope of 'a single customer' and the lookup-by-Id behavior, but it doesn't disclose any additional behavioral traits (e.g., what happens if the Id doesn't exist, whether it returns deleted customers, or response shape). With annotations carrying most of the burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one concise sentence that completely covers the core purpose. There is zero wasted text and the essential scope ('single customer', 'by Id') is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-by-Id tool with two well-documented parameters and annotations declaring it safe, the description plus schema is nearly complete. The only missing context is clarification of edge-case behavior (e.g., not found, error handling) or how 'company' behaves when multiple are connected, but the schema already covers the connection ambiguity. This is complete for practical invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with the 'id' parameter documented as 'QuickBooks Id of the Customer' and 'company' documented clearly. The description adds no new parameter semantics beyond restating 'by Id'. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a specific resource ('a single customer'), and the system ('QuickBooks Online'), and identifies the lookup key ('by Id'). This clearly distinguishes it from broader tools like search_customers or list-like tools, and from customer-related aggregate tools like get_customer_balance or get_customer_sales.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the tool is for retrieving a specific customer when the Id is known, which is a clear context. It doesn't explicitly name alternatives (e.g., search_customers for finding by name) or state exclusions, so it misses the explicit 'when not to use' guidance, but the context is clear enough for an agent to select it for id-based lookups.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customer_balanceget customer balanceA
Read-only
Inspect

Generate the QuickBooks customer balance report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); customer, department, arpaid (All|Paid|Unpaid).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the call read-only and non-destructive, but the description adds substantial behavioral nuance: hidden_zero_rows accounting, filters echoing only parameters QuickBooks actually applied, and no_data signaling an empty period. It also documents row type/depth/value semantics and how id relates to label cells. This goes well beyond what the annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and no sentence is padding. The return-format sentence is dense and somewhat run-on, but given the lack of an output schema, every clause earns its place. It could be better structured, but it is not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and nested return objects, the description covers the essential contract: top-level keys, row types, depth, value alignment, id meaning, hidden-zero handling, filter echo, and empty-period behavior. This is sufficient for an agent to invoke the tool and correctly interpret its results. The only notable gap is usage positioning, which is already accounted for in the usage_guidelines dimension.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all parameters thoroughly, including date formats, date_macro choices, and enums for accounting_method, summarize_column_by, and arpaid. The description adds little about how to set parameters; its filter-echo note is about output behavior rather than parameter syntax. With 100% schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Generate the QuickBooks customer balance report.' This clearly distinguishes it from create/update/delete siblings and from unrelated report tools like get_cash_flow or get_profit_and_loss. However, it never explicitly positions itself against the closely related get_customer_balance_detail, so some sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to choose this tool over alternatives. It names no siblings such as get_customer_balance_detail, get_customer_sales, or get_customer_income, and it states no exclusions or when-not-to-use conditions. An agent must infer selection entirely from the tool name and nearby context signals.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customer_balance_detailget customer balance detailA
Read-only
Inspect

Generate the QuickBooks customer balance detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); aging_method (Report_Date|Current); start_duedate, end_duedate (YYYY-MM-DD); customer, department, term, shipvia, arpaid (All|Paid|Unpaid); columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds valuable behavioral details: how rows are structured (type, depth, values), the meaning of 'id' and its alignment on detail vs summary reports, the omission of all-zero rows (counted in hidden_zero_rows), the 'filters' field reflecting applied parameters, and 'no_data' flag. This goes beyond annotations and clarifies output behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence but packs in essential details about the return format and behavioral quirks. It is front-loaded with the purpose, then explains the output structure. While it could be broken into shorter sentences for readability, it contains no fluff and each clause adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (many parameters, a detailed report) and the absence of an output schema, the description explains the return structure thoroughly: columns, column_types, rows with type/depth/values, id alignment, hidden_zero_rows, filters, and no_data. This is sufficient for an agent to understand what the tool returns. Minor gaps exist (e.g., does not describe what the report contains beyond 'balance detail'), but the core context is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, meaning all parameters (report_date, date_macro, aging_method, customer, columns, etc.) are already documented in the input schema. The description does not add any parameter-specific semantics beyond what the schema provides. Since the schema carries the full burden, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate the QuickBooks customer balance detail report.' This is a specific verb (generate) and resource (customer balance detail report), distinguishing it from summary reports like get_customer_balance by the word 'detail.' The purpose is unambiguous and matches the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives. It focuses on the output structure and behavior but provides no guidance on when to choose this over get_customer_balance, get_aged_receivable_detail, or other report tools. Usage is implied by the name and the fact it generates a specific report, but there is no explicit context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customer_incomeget customer incomeA
Read-only
Inspect

Generate the QuickBooks customer income report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); term.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral detail beyond the annotations: row types, depth and value alignment, id semantics for summary vs detail reports, hidden_zero_rows counting, filter-confirmation behavior, and the no_data flag. This gives an agent a strong model of what the call will return and how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the primary purpose and then packs highly relevant output semantics into a compact second sentence. Every clause earns its place; there is no filler or repetition of annotation data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, the description takes on the burden of explaining return values, and it does so thoroughly: columns, rows, types, depth, id alignment, hidden_zero_rows, filters, and no_data. Combined with a fully documented input schema, this is complete enough for an agent to invoke the tool and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description does not need to restate parameter meanings. It adds one useful nuance by explaining that filters only appear in the response when QuickBooks confirms it applied them, but it otherwise contributes little parameter-specific semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: 'Generate the QuickBooks customer income report.' This is specific enough to identify the operation. However, it does not explicitly distinguish this report from sibling tools like get_customer_sales or get_customer_balance, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or conditions that would route an agent to get_customer_income rather than other report tools. The only usage context is implied by the tool name and title.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customer_salesget customer salesA
Read-only
Inspect

Generate the QuickBooks customer sales report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate read-only and non-destructive, and the description adds valuable behavioral details: it explains the output format, handling of zero rows, and the semantics of filters and no_data flag. It goes beyond annotations by clarifying the shape of the response and edge cases, which is crucial for understanding tool behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but extremely dense, packing a lot of information into a few sentences. It's structured and front-loads the main purpose, but the output format explanation is lengthy and might be considered as lacking in readability, though still efficient for the technical audience.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool and the lack of an output schema, the description does an admirable job explaining the return format, including the role of each field and edge cases. The output schema is absent, so the description compensates fully; the only minor omission is examples of date_macro values, but the ellipsis implies more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 100% of parameters with descriptions, so baseline is 3. The description adds context about filters (that unconfirmed filters are absent), but the schema already explains the parameters and their types. The description reinforces the filter behavior but doesn't introduce new parameter details beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates the QuickBooks customer sales report, a specific verb and resource. It distinguishes itself from siblings by specifying the report type and providing detailed output structure, which sets it apart from other get_* reports.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

While it doesn't explicitly say when not to use it or name alternative tools, the description clearly implies it's for customer sales reports, and the sibling list includes related tools like get_customer_income or get_item_sales, suggesting an agent can infer usage. It provides report-specific details but lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customer_typeget customer typeB
Read-only
Inspect

Get a single customer type from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CustomerType.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description accurately reflects a simple read operation and adds the 'single ... by Id' scoping, but offers little behavioral detail beyond that, such as not-found behavior or return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every word is necessary and the core operation, resource, and lookup method are all stated efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id tool, the schema and annotations together provide enough invocation context: required id, optional company, and read-only safety. The lack of output schema or usage alternatives is a minor gap given the tool's low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters well. The description's reference to 'by Id' aligns with the required id parameter but does not add meaningful meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Get a single customer type from QuickBooks Online by Id.' It communicates the scope (single entity, by Id) and distinguishes from broader search tools, though it does not name a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as search_customer_types. It implies a direct-by-Id lookup but gives no context about when that is appropriate or when a search tool should be selected instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_departmentget departmentA
Read-only
Inspect

Get a single department from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Department.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the mild context that a single department is returned, but provides no additional behavioral details such as error handling or what happens if the ID is invalid. It is consistent with annotations; no contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler. It front-loads the action and target, and includes the key qualifier 'by Id' immediately. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-ID tool with two parameters and no output schema, the description covers the essential use case. It does not explain the return structure or failure modes, but these are largely inferable for a standard 'get' operationcase. The interplay of the optional 'company' parameter is documented in the schema, so the description is complete enough for selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the input schema already documents both 'id' and 'company' with clear descriptions. The tool description only reiterates 'by Id' without adding semantics about the 'company' parameter or ID format beyond what the schema provides. This meets the baseline for fully described schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Get'), the resource ('a single department'), and the identifying mechanism ('by Id'). This distinguishes it from sibling tools like get_department_sales (sales data) and search_departments (search without a known ID). No ambiguity remains about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'by Id' phrasing implies the tool is appropriate when the caller already has a QuickBooks department ID, but it does not explicitly point to alternatives like search_departments for lookup-by-name or list scenarios. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_department_salesget department salesA
Read-only
Inspect

Generate the QuickBooks department sales report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide readOnlyHint=true and destructiveHint=false; the description goes far beyond by documenting row types, depth, Money-as-number values, id alignment on summary vs detail reports, hidden_zero_rows, the filters echo with ignored params omitted, and the no_data flag. This gives an agent an accurate model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries a distinct fact about the return shape or edge-case behavior, with the core purpose front-loaded. The description is dense but not bloated, and the structured enumeration of return fields reads efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description manually documents columns, column_types, rows, row type and depth, id alignment, hidden_zero_rows, filters, and no_data. Combined with the self-describing input schema and safety annotations, an agent has enough to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the schema already documents params and company. The description adds valuable semantic context: ignored filter parameters silently disappear from the returned filters list, and no_data explains a period with nothing to report. It doesn't restate syntax, but it clarifies interpretation beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Generate the QuickBooks department sales report'), and the detailed return-shape explanation makes it unmistakable from sibling tools like get_class_sales or get_customer_sales. Even without naming alternatives, it fully identifies what the tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The report name implies when to use it for department-level sales reporting, but there is no explicit statement of when to prefer it over sibling sales-report tools or what inputs select it. It provides clear context, but not exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_depositget depositA
Read-only
Inspect

Get a single deposit from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Deposit.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. The description adds no additional behavioral context such as return format, error handling, or auth requirements. It is consistent with annotations but does not provide value beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with zero waste. The verb and object are front-loaded, and there is no redundant phrasing. It is appropriately minimal for a simple getter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read operation, the description is adequate. Annotations cover safety, schema covers parameters, and the return value (a deposit) is implied. The only missing element is an explicit note about the return format, but that is not critical for a get-by-ID tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage, with both parameters (id and company) described. The description adds no further meaning to the parameters; it relies entirely on the schema. Baseline of 3 is appropriate since schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a clear resource ('a single deposit'), and the identifying mechanism ('by Id'). This distinguishes it from search_deposits (which searches by criteria) and create_deposit (which creates a new one), leaving no ambiguity about the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives like search_deposits. It is implied that it should be used when a deposit ID is known, but there is no explicit guidance on when not to use it or mention of the search alternative. The usage context is only inferred from the tool name and sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_employeeget employeeA
Read-only
Inspect

Get a single employee from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Employee.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. The description adds the 'by Id' scoping constraint, which is useful. However, it doesn't disclose behavior like 404 handling, whether the employee must be active, or response shape. With annotations covering the safety profile, a 3 is appropriate – the description adds some value but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with zero waste. The core action, resource, system, and lookup key are all front-loaded. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a simple read-by-id tool whose annotations already carry the safety profile. The only minor gap is not explicitly routing to search_employees when an Id is unavailable, but the schema and sibling list make that inferable. No output schema exists, but for a single-entity getter the return value is self-evident.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds 'by Id' which reinforces the id parameter's role, but doesn't add meaning beyond what the schema provides. Baseline 3 is correct when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a specific resource ('a single employee'), and the system ('QuickBooks Online'), and identifies the lookup key ('by Id'). It is clear and distinguishes from search_employees (which searches) and create_employee/update_employee (which mutate). It doesn't explicitly name a sibling, but the verb+resource+scope is enough to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this when you have an employee Id and need a single employee record. It does not explicitly state when to use search_employees instead (e.g., when you don't have an Id or need multiple employees), nor does it mention that the company parameter is optional when only one company is connected. The schema covers the company parameter, but the description provides no explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_estimateget estimateA
Read-only
Inspect

Get a single estimate from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Estimate.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is established. The description adds minimal behavioral context beyond ID-based retrieval and does not discuss not-found behavior or response details; this is acceptable for a simple read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It efficiently states the action, resource, system, and lookup key.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with annotations covering safety and a schema fully describing the parameters, the description is mostly complete. It does not describe the return shape or error behavior, but those are not essential for selecting and invoking this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with descriptions for both id and company. The description reinforces that retrieval is by ID, but it adds no meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is specific and unambiguous: 'Get a single estimate from QuickBooks Online by Id.' It clearly identifies the verb, resource, and scope, and stands out from sibling tools like search_estimates or update_estimate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context that this tool is for retrieving one estimate by ID. It does not explicitly mention when to prefer search_estimates over this tool, but the by-ID scope is a sufficient and clear usage signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_exchange_rateget exchange rateA
Read-only
Inspect

Get the exchange rate QuickBooks uses for a currency into the company's home currency on a date (multicurrency companies only).

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
currencyYesSource currency code, e.g. USD (the rate into the company's home currency).
as_of_dateNoYYYY-MM-DD; defaults to today.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is established. The description adds that the rate is the one QuickBooks uses and is date-specific, which is useful context. However, it does not disclose behavior for non-multicurrency companies, missing rates, or the return shape, so it adds only modest behavioral transparency beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that communicates the operation, direction, date dimension, and multicurrency restriction without any filler. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only lookup with fully documented parameters, the description plus schema covers the essential invocation details. The main gaps are the lack of an explicit pointer to search_exchange_rates for plural/search scenarios and no note about what happens if the company is not multicurrency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and all three parameters already have meaningful descriptions: company, currency, and as_of_date. The description reinforces the currency/home-currency direction and date relevance, but it does not add significant semantic detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: get the exchange rate QuickBooks uses for a source currency into the company's home currency on a date. It also adds the important scope restriction that this applies to multicurrency companies only. However, it does not explicitly distinguish itself from the sibling search_exchange_rates, so an agent might not immediately know when to use this getter versus the search tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear precondition (multicurrency companies only) and defines the operation's context: one currency, home-currency direction, and a date. It does not name alternatives such as search_exchange_rates or explain when to prefer one over the other, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_general_ledgerget general ledgerA
Read-only
Inspect

Generate the QuickBooks general ledger report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); account (Account Ids, comma-separated), account_type (Bank, AccountsReceivable, Income, Expense, ...), source_account; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool readOnly and non-destructive, and the description goes well beyond them. It documents the full return contract: row type/depth semantics, Money cells as numbers, id alignment for summary vs detail reports, hidden_zero_rows, confirmed applied filters, and no_data. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the purpose is front-loaded and the dense second sentence packs only high-value behavioral facts. It is not perfectly scannable due to the long semicolon-heavy structure, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description serves as the sole documentation of the return value and covers its important quirks: row types, value alignment, id semantics, hidden zero rows, filter confirmation, and no_data. Input parameters are fully covered by the schema, so nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema documents all parameters at 100% coverage, so the baseline is 3. The description adds one valuable parameter-related behavior: the returned filters array only includes parameters QuickBooks actually applied, so an absent filter means an ignored input. It avoids restating parameter formats already covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Generate the QuickBooks general ledger report.' It clearly identifies the exact report and the returned object shape. However, it never contrasts with the many sibling report tools (e.g. get_journal_report, get_trial_balance), so differentiation relies on the report's name rather than explicit scoping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance and no reference to alternative report tools. Given the large sibling set, an agent must infer from the tool name alone when to pick this instead of get_journal_report or get_transaction_detail_by_account. This is the clearest gap in the definition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_inventory_adjustmentget inventory adjustmentC
Read-only
Inspect

Get a single inventory adjustment from QuickBooks Online by Id. Changes the quantity on hand of inventory items and books the valuation difference; US companies on Plus or Advanced only.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the InventoryAdjustment.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotation contradiction: the description claims the operation 'Changes the quantity on hand of inventory items and books the valuation difference' while readOnlyHint=true declares the tool is read-only. The description does not clarify that these change semantics describe the inventory adjustment record itself, not the GET call. It does add the useful US Plus/Advanced eligibility constraint, but the contradiction is a critical failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is appropriately front-loaded and concise, but the second sentence is an ambiguous dangling fragment that introduces the annotation contradiction. Not every sentence earns its place, and the unclear subject structure is a readability problem.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and the description does not describe what the GET returns. The read-only behavior is undercut by the contradictory second sentence, and although the US Plus/Advanced eligibility note is useful, the description is not complete enough for an agent to safely invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so id and company are already fully documented in the JSON schema. The description adds only the eligibility constraint and no meaningful parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb ('Get'), resource ('inventory adjustment'), and retrieval key ('by Id'), making it easy to distinguish from sibling create/update/delete and get-valuation tools. The second sentence is confusing but does not obscure the primary purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to prefer this tool over alternatives, such as get_inventory_valuation_summary, get_physical_inventory_worksheet, or search tools, and does not state when not to use it. 'By Id' is already captured by the required id parameter and adds no real selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_inventory_valuation_detailget inventory valuation detailA
Read-only
Inspect

Generate the QuickBooks inventory valuation detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); start_svcdate/end_svcdate or svcdate_macro, group_by; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and destructiveHint annotations, the description discloses the exact row shape, row-type/depth semantics, id alignment behavior, omission of all-zero rows with hidden_zero_rows count, and the fact that filters contains only QuickBooks-confirmed filters. This is substantial non-obvious behavior an agent needs to interpret output correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the action and return shape, then uses compact clauses to convey row types, id alignment, zero-row handling, filter confirmation, and no_data. No sentence is wasted and the structure mirrors the data hierarchy it describes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining return values, and it covers columns/types/rows, row kinds, id semantics, hidden zero rows, filter confirmation, and empty-period signaling. For a read-only report tool, remaining gaps such as auth and error details are minor and not needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100% and the params property already documents date macros, group_by, columns, sort_by, and sort_order. The description adds meaning by explaining that unconfirmed/ignored parameters are absent from the returned filters list and that no_data signals an empty period, which helps agents reason about parameter effects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Generate the QuickBooks inventory valuation detail report.' The 'detail' wording and row semantics distinguish it from get_inventory_valuation_summary, but the description never names that sibling or states the selection criterion, so differentiation is implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance or alternative routing in the description. The only signal is the report's name and the word 'detail', so usage is implied for line-level inventory valuation rather than clearly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_inventory_valuation_summaryget inventory valuation summaryA
Read-only
Inspect

Generate the QuickBooks inventory valuation summary report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); item.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and non-destructive, so description need not repeat that. The description adds rich behavior: return format details (columns, rows with types/depth, Money cells as numbers), the id alignment difference between summary and detail reports, omission of zero-rows with a counter, filters list inclusion, and no_data flag. This is valuable beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one dense paragraph, front-loaded with the report generation purpose. It packs a lot of details, but could be broken into bullet points for easier scanning. However, every sentence adds value, so it remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the return structure and no output schema, the description does a good job explaining the format, including edge cases like zero-rows and no_data. It lacks details on pagination or max depth but is sufficient for an agent to interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and includes descriptions for report_date, date_macro, summarize_column_by, and company. The description adds only a bit about filters and no_data, but baseline is 3 due to high coverage; the added detail about filter behavior pushes it to 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates a QuickBooks inventory valuation summary report. It distinguishes itself from the sibling get_inventory_valuation_detail by focusing on 'summary' and describing a summary-specific structure. However, it doesn't explicitly compare to the detail tool by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for summary-level inventory valuation reports, and its focus on summary report structure differentiates it from detail reports. However, it does not explicitly state when to prefer this over the detail counterpart or other inventory reports, and there are no exclusion conditions or alternative tool names mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_invoiceget invoiceA
Read-only
Inspect

Get a single invoice from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Invoice.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds no additional behavioral context such as error handling, rate limits, or authentication requirements. Since the description carries minimal burden given annotations, it adds no extra value beyond the obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence that conveys the essential action and distinguishing factor. No fluff or redundancy, and the key qualifier 'by Id' is front-loaded. Perfectly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with only 2 parameters, no output schema, and no nested objects. The description covers the action and scope but does not mention the return value structure or behavior when the invoice is not found. While annotations cover safety, a brief note on the response or error case would make it more complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'id' and 'company' are already documented. The description does not add further meaning about parameter formats or relationships. With full schema coverage, baseline 3 is correct; the description adds no extra nuance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get', the resource 'a single invoice', and the key qualifier 'by Id'. This distinguishes it from sibling search tools like search_invoices and list tools like get_open_invoices, making the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when an invoice ID is available, and the schema requires 'id'. However, it does not explicitly name alternatives or state when not to use this tool (e.g., when searching without an ID). Guidance is implied rather than explicit, so a 3 is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_itemget itemA
Read-only
Inspect

Get a single item from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Item.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds the context that this is a QBO Item lookup, but does not disclose behavior such as error handling when the item is not found, or whether it returns the full item representation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource, with no filler. Every word contributes to purpose or scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-by-Id tool, the description plus annotations and schema cover the essential calling context. The absence of an output schema is slightly mitigated by 'Get a single item', but there is no mention of what happens if no item matches the Id, leaving a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in structured form. The description adds no additional parameter-level nuance beyond the mention of 'by Id', which is consistent with the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get'), a specific resource ('a single item from QuickBooks Online'), and a clear lookup key ('by Id'). It is immediately distinguishable from sibling tools like search_items or get_item_sales even without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives. It does not mention that search_items should be used when the Id is unknown, or explain how this differs from get_item_sales. The agent must infer usage from the phrase 'by Id'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_item_salesget item salesA
Read-only
Inspect

Generate the QuickBooks item sales report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); start_duedate, end_duedate (YYYY-MM-DD).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a wealth of behavioral detail beyond the annotations: the exact return structure (columns, column_types, rows), row semantics (section/data/total, depth, values alignment), ID alignment behavior, handling of all-zero rows via hidden_zero_rows, the filters field reflecting applied parameters, and the no_data flag. This goes far beyond the readOnlyHint and destructiveHint annotations and provides critical operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose and then provides a dense but structured explanation of the output. It is one long sentence plus a brief final clause, but it packs a large amount of necessary information without redundancy. It could be more concise, but the structure is logical and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully explains the expected return value: columns, rows, types, depth, ID alignment, hidden zero rows, filters, and the no_data flag. It covers edge cases and the meaning of each field, making the tool effectively usable without needing to guess at the response format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description covers 100% of parameters with detailed explanations for start_date, end_date, date_macro, accounting_method, summarize_column_by, filters, and company. The tool description adds no parameter-level meaning, so a baseline 3 is appropriate given the schema already does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates the QuickBooks item sales report, which is a specific verb and resource that distinguishes it from sibling sales reports like customer sales or department sales. Though it doesn't explicitly name alternatives, the resource is unambiguous and the phrase 'item sales report' is distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus other sales report tools. It only states what it does, without mentioning exclusions, alternatives, or the typical scenario for which it is intended. Among many sibling sales reports, an agent gets no help selecting this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_journal_entryget journal entryA
Read-only
Inspect

Get a single journal entry from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the JournalEntry.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds only that the operation reads from QuickBooks Online by Id, which is essentially the core behavior already evident from the name. No extra behavioral details such as error handling or response format are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence that front-loads the operation, resource, and lookup key. There is zero redundancy or irrelevant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-by-Id tool with full schema coverage and annotations declaring the operation read-only and non-destructive, the description is nearly complete. The only minor gap is that, with no output schema, the return payload is not described, but the semantics of a 'get' tool make this predictable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the 'id' and 'company' parameters are already documented in the schema. The description's 'by Id' phrasing reinforces the importance of the id parameter but adds no new semantic value beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a specific resource ('a single journal entry'), and a precise scope ('from QuickBooks Online by Id'). It clearly distinguishes this tool from search_journal_entries (which searches multiple entries) and get_journal_report (which produces a report).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by Id' strongly implies the tool should be used when the agent already knows the specific QuickBooks Id of the journal entry. However, it does not explicitly name alternatives like search_journal_entries or state when not to use this tool, so some inference is required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_journal_reportget journal reportA
Read-only
Inspect

Generate the QuickBooks journal report report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/destructiveHint annotations, the description extensively documents runtime behavior: row types/depth, numeric Money cells, id alignment differences between summary and detail reports, omitted zero rows counted in hidden_zero_rows, filter confirmation semantics, and the no_data flag. This is substantial added behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but economical, packing the entire return contract into two sentences with no filler. However, the second sentence is a long run-on that could be clearer as structured bullets, and 'journal report report' is a minor typo.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description fully compensates by specifying the exact return shape, row semantics, id behavior, zero-row handling, filter confirmation, and no_data case. Combined with the 100% schema coverage for parameters, an agent has enough context to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input parameters are already fully described. The tool description adds no extra meaning about start_date, date_macro, columns, or sort parameters; it only describes result behavior. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks journal report report,' and details the returned structure. It doesn't explicitly contrast the journal report with sibling report tools like get_general_ledger or get_trial_balance, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to choose this report over the many sibling report/list tools, nor are any exclusions or alternatives named. The only implied usage is the report name itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_open_invoicesget open invoicesA
Read-only
Inspect

Generate the QuickBooks open invoices report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); customer, department, term; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the readOnlyHint/destructiveHint annotations by detailing the return structure (columns, column_types, rows), row types, id semantics, hidden_zero_rows, filters, and no_data flag. This gives the agent significant behavioral insight, though it could mention potential size of results or any caveats about performance, which is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but packed with essential information. It front-loads the core purpose and return structure, then provides detail on edge cases. While it's long, every sentence provides value, and the structure is logical, though it could be slightly more streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested objects, no output schema), the description comprehensively explains the return format, edge cases (all-zero rows), and filtering behavior. It covers everything an agent needs to interpret the results correctly, making it complete for its purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has near-complete coverage of parameters with descriptions of date options, field lists, and company selection. The description does not add additional parameter-level details beyond what's in the schema, so a baseline score of 3 is appropriate given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that it generates the QuickBooks open invoices report, using a specific verb ('generate') and a specific resource ('QuickBooks open invoices report'). It distinguishes itself from sibling tools like search_invoices and get_invoice by focusing on the report format rather than individual invoice retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for reporting scenarios, and the detailed explanation of report structure and filters helps an agent understand when to use it. However, it does not explicitly state when NOT to use it or direct to alternatives like search_invoices for non-report needs, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_paymentget paymentA
Read-only
Inspect

Get a single payment from QuickBooks Online by Id. An accounting record of a customer payment already received. Payment processing is not supported; ProcessPayment must be omitted or false.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Payment.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the call as read-only and non-destructive, so the description does not need to restate that. It adds useful behavioral nuance by warning that payment processing is unsupported and ProcessPayment must be omitted or false, which prevents an agent from attempting write-like behavior through this read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded and the one non-obvious constraint placed second. There is no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple single-resource lookup with only two parameters and no output schema, so the description supplies enough operational context. It might have named search_payments as the companion for finding payments without an Id, but that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so id and company are already documented. The description adds only that lookup is 'by Id', which reinforces the schema but does not add format, syntax, or interaction details beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific verb-resource pair ('Get a single payment from QuickBooks Online by Id') and clarifies the resource type ('accounting record of a customer payment already received'). This clearly distinguishes it from search-style siblings and other get_* tools like get_bill_payment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'By Id' and 'single payment' indicate this is for retrieving one known record rather than listing/searching. The explicit warning that payment processing is not supported and ProcessPayment must be omitted/false gives a clear when-not condition, though it does not name an alternative tool for processing payments.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_payment_methodget payment methodA
Read-only
Inspect

Get a single payment method from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the PaymentMethod.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds that the tool reads from QuickBooks Online and requires an Id, but it does not disclose behavior such as errors, missing records, or response shape. This is acceptable but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence that states the action, target, and key constraint. Every word serves a purpose and the core information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only get tool with one required parameter, full schema coverage on all parameters, and safety annotations provided, the description is sufficient. An agent has everything needed to call it correctly, and nothing important appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema. The description repeats that lookup is 'by Id' but adds no meaning beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Get a single payment method from QuickBooks Online by Id.' It clearly distinguishes this from sibling tools like create_payment_method, update_payment_method, and search_payment_methods because it is framed as a single-record lookup by identifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by Id' implies the intended use case: retrieve one known payment method. However, the description does not explicitly mention when to prefer this over search_payment_methods or what to do when no Id is known, so the guidance remains implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_physical_inventory_worksheetget physical inventory worksheetA
Read-only
Inspect

Generate the QuickBooks physical inventory worksheet report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); item.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the readOnlyHint and destructiveHint annotations, disclosing detailed output semantics: row types (section/data/total), depth, value alignment, id mapping, omission of all-zero rows, hidden_zero_rows counting, filter confirmation behavior, and the no_data flag. This gives the agent a thorough understanding of the tool's exact behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single complex sentence but is information-dense and front-loaded with the primary purpose. While it could be restructured into clearer bullets, every clause provides necessary behavioral details. It is not overly verbose for the richness it conveys, though the long run-on structure slightly detracts.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining return values, and it does so comprehensively: columns, column_types, rows, id alignment, hidden_zero_rows, filters, and no_data. It also clarifies how ignored parameters appear. For a report-read tool with read-only annotations, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters (params and company), so the baseline is 3. The description does not add any meaning to the parameters themselves; it only references filters, which is output-related. There is no compensation for parameter detail beyond the schema, so the score remains at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Generate' and the specific resource 'QuickBooks physical inventory worksheet report', making the purpose unambiguous. However, it does not explicitly distinguish itself from sibling report tools like get_inventory_valuation_detail or get_balance_sheet, so it lacks the explicit sibling differentiation required for a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It offers no context, prerequisites, or exclusions, leaving the agent to infer usage solely from the tool name. No mention of alternatives is made, so this dimension scores poorly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_preferencesget preferencesB
Read-only
Inspect

Get the company's QuickBooks preferences.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond that: it does not describe what 'preferences' covers, whether it returns all preference groups, or anything about the response. There is no contradiction, but no added transparency either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler or redundancy. It earns its place by adding 'company's QuickBooks' context beyond the tool title, and there is no unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with one optional parameter fully documented in the schema and read-only annotations, the description is largely sufficient. The main gap is that it does not hint at the scope or shape of the returned preferences, but this is minor given the low tool complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single optional 'company' parameter is already fully explained in the schema. The description adds nothing about the parameter, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb-resource pair: 'Get the company's QuickBooks preferences.' It is specific enough to distinguish the resource from other getters and from update_preferences. It does not list what the preferences include, which keeps it from a 5, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The retrieval intent is implied by the verb 'Get', so an agent can infer this is the read tool for preferences. However, the description gives no explicit guidance about when to use this versus update_preferences for modifications or other company-level getters, and it states no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_profit_and_lossget profit and lossA
Read-only
Inspect

Generate the QuickBooks profit and loss report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); adjusted_gain_loss (true|false).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and non-destructive behavior. The description adds significant behavioral context beyond annotations: explains row types, how Money cells are represented, id assignment, omission of zero rows, the meaning of filters and no_data. This fully discloses behavior without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized, front-loading the core purpose and then explaining response details. It is longer than minimal but every sentence carries functional information, so it earns a 4 for structure without excessive fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description covers return structure, edge cases, and filter semantics comprehensively. It does not, however, contrast with get_profit_and_loss_detail or mention pagination/limits, which would make it fully complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are documented. The description adds valuable semantic detail beyond the schema, such as how filters behave when ignored and the significance of hidden_zero_rows. This enriches parameter understanding without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates the QuickBooks profit and loss report and specifies the return structure. It distinguishes itself implicitly from get_profit_and_loss_detail by name, but does not explicitly mention the sibling or the difference, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the related get_profit_and_loss_detail or other report tools. The description focuses on how to use parameters and interpret output, but lacks any explicit when-to-use or when-not-to-use instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_profit_and_loss_detailget profit and loss detailA
Read-only
Inspect

Generate the QuickBooks profit and loss detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); account, account_type, employee, payment_method; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses important runtime behaviors: all-zero rows are omitted and counted in hidden_zero_rows, filters lists only the filters QuickBooks confirmed applying (ignored parameters are absent), and no_data indicates an empty period. This gives an agent reliable expectations for parsing and interpreting results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but each sentence carries meaningful detail about return structure, row semantics, hidden rows, filter confirmation, and no_data. It is not a short one-liner, but the complexity of the output deserves this level of specificity; the key purpose is front-loaded and the details are logically ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description fully compensates by documenting the complete return shape: columns, column_types, rows with type/depth/values, id alignment, hidden_zero_rows, filters, and no_data. For a report-generation tool with rich output and many sibling report tools, nothing essential is missing for correct invocation and interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameters are already documented in the schema. The description adds value by explaining that the filters parameter may not appear in the output if QuickBooks ignored it, and it clarifies how values and ids map differently between summary and detail reports. This goes beyond the raw schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Generate the QuickBooks profit and loss detail report') and consistently uses 'detail' to distinguish it from the sibling summary tool get_profit_and_loss. The report structure is also named explicitly, so an agent can tell exactly what this tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for detailed P&L reporting, and the mention of 'summary reports' hints that a summary variant exists, but it never explicitly says when to choose this tool over get_profit_and_loss or other report tools. There is no direct exclusion or alternative routing; the agent must infer usage from the 'detail' naming.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_purchaseget purchaseB
Read-only
Inspect

Get a single purchase from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Purchase.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true and destructiveHint=false annotations already establish that this is a safe read, and the description aligns by saying 'Get'. It adds little beyond 'single ... by Id' (which is also in the schema) and QuickBooks Online as the source; no extra behavior such as auth, not-found handling, or company selection is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that puts the action and object first and drops filler. It is as concise as the definition can be without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id with fully documented parameters and safety annotations, the definition is mostly sufficient. The absence of an output schema means no return shape is specified, but the verb 'Get' and object 'purchase' reasonably imply the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema fully explains id and company. The description's 'by Id' merely echoes the required parameter and adds no new semantic context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States exactly the operation (get), resource (purchase), system (QuickBooks Online), and lookup key (Id), making the tool's function immediately clear. It doesn't explicitly distinguish from siblings such as get_purchase_order or search_purchases, but the resource and single-by-Id semantics prevent major ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or alternative guidance is provided. An agent is not told to use search_purchases when lacking an Id or to use get_purchase_order for purchase orders, so decision support relies entirely on the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_purchase_orderget purchase orderA
Read-only
Inspect

Get a single purchase order from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the PurchaseOrder.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds no notable behavioral context beyond what is already in the annotations and schema, but it does not contradict them either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no filler. It front-loads the resource and retrieval method, making it immediately actionable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read operation with two well-documented parameters and no nested objects. The description plus annotations cover the essential behavioral context; the only minor gap is not describing the return shape, but that is strongly implied by 'Get a single purchase order.'

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters are already documented clearly in the schema. The description adds no additional parameter semantics, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('a single purchase order') and explicitly states the retrieval key ('by Id'). This clearly distinguishes it from list/search tools like search_purchase_orders and from create/update/delete variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly communicates the intended context: use this tool when you have a QuickBooks Id and need exactly one purchase order. It does not explicitly name alternatives or exclusions, but the 'single ... by Id' wording makes the appropriate usage unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_recurring_transactionget recurring transactionA
Read-only
Inspect

Get a single recurring transaction from QuickBooks Online by Id. A template QuickBooks uses to create a transaction on a schedule, or on request.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the RecurringTransaction.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate readOnlyHint=true and destructiveHint=false, so the description does not need to restate that this is a safe read operation. It adds useful domain context by explaining that a recurring transaction is a template used to create transactions on a schedule or on request, but it does not disclose error behavior, response shape, or any additional operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence states exactly what the tool does, and the second briefly explains the domain concept in a way that helps the agent understand the resource. It is front-loaded and appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-resource lookup with read-only annotations and fully documented parameters, the description is largely complete for selecting and invoking the tool. It does not explain return values, but that is not necessary for a straightforward get-by-Id operation, and no output schema exists to contradict that expectation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% parameter description coverage, so the schema already documents id and company clearly. The description only confirms lookup 'by Id' and does not add further meaning beyond what the schema provides, which fits the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Get a single recurring transaction'), names the resource ('from QuickBooks Online'), and identifies the lookup key ('by Id'). It does not explicitly distinguish itself from the sibling search_recurring_transactions, but the singular 'by Id' phrasing makes the intended use apparent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied: use this when you already have the QuickBooks Id and need exactly one recurring transaction. However, there is no explicit guidance about when to choose this tool over search_recurring_transactions, nor any mention of the optional company parameter requirement when multiple companies are connected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_refund_receiptget refund receiptA
Read-only
Inspect

Get a single refund receipt from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the RefundReceipt.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, which fully covers the safety profile. The description adds no behavioral context beyond a simple 'Get'—it does not mention what happens if the Id is not found, whether the company parameter is required, or any other nuances. With annotations present, the description offers minimal added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence that immediately states the action and the key qualifier. There is no redundancy or filler—every word earns its place. It is perfectly sized for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward get-by-id operation with only two parameters (both documented in the schema) and annotations covering read-only/non-destructive behavior, the description is adequate. It does not mention the company parameter or error handling, but those are either covered by the schema or low-risk for a read operation. The only gap is not pointing to search_refund_receipts as an alternative when the Id is unknown, which is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'id' and 'company' already explained. The description's mention of 'by Id' aligns with the id parameter but adds no new semantic detail. Baseline 3 is appropriate because the schema carries the full parameter documentation; the description does not enhance it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get', the resource 'a single refund receipt', and the method 'by Id'. It is unambiguous and distinguishes from search_refund_receipts (which implies searching without an ID) and from create/update/delete operations. The purpose is immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description specifies the key condition: you need an Id to retrieve a single receipt. This clearly indicates when to use this tool over search_refund_receipts, which would be used when an Id is not known. However, it does not explicitly mention alternatives or exclusions, so it misses a small opportunity to guide the agent toward the search tool when an Id is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sales_receiptget sales receiptA
Read-only
Inspect

Get a single sales receipt from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the SalesReceipt.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds a bit of context ('from QuickBooks Online', 'single') but discloses no additional behavioral traits such as not-found handling or idempotency, which is acceptable for a simple read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler or redundant qualifiers. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-by-id tool, the annotations cover the safety profile and the schema covers both parameters. The description supplies enough context for correct invocation, though the absence of an output schema means return-value behavior is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents both the 'id' and 'company' parameters. The description only restates the Id-based lookup and adds no extra meaning beyond what the schema provides, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a clear resource ('single sales receipt'), and the retrieval key ('by Id'). It is well differentiated from search_sales_receipts by making clear this is a single-record retrieval, though it does not explicitly name the sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by Id' implies the tool should be used when the agent already has a specific QuickBooks receipt ID and wants that one record. However, it gives no explicit guidance about when not to use it or how it compares to search_sales_receipts, leaving the usage context implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tax_agencyget tax agencyA
Read-only
Inspect

Get a single tax agency from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the TaxAgency.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds that it returns a single entity, clarifying cardinality beyond the schema, but provides no details on error handling or company context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the verb and resource, with no redundant words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-ID tool, the core operation is covered. However, with no output schema, return value and error behavior remain implicit, and the optional company context is only in the schema, not the description. Acceptable but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both id and company documented. The description adds no extra parameter semantics beyond mentioning 'by Id', which is already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('get'), resource ('tax agency'), and method ('by Id'). Clearly distinguishes from sibling tools like search_tax_agencies or create_tax_agency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as search_tax_agencies. It does not mention that a known Id is needed or that search should be used to find it, leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tax_codeget tax codeA
Read-only
Inspect

Get a single tax code from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the TaxCode.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the retrieval scope ('single ... by Id') but no additional behavioral context such as auth requirements, rate limits, or return caveats. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One crisp, front-loaded sentence with no filler. It states the operation, the resource, the source system, and the key selector (Id) in minimal space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only single-entity fetch with fully documented parameters, the description provides the essential context. Since there is no output schema, a brief note about the returned TaxCode payload would make it fully complete, but this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both 'id' and 'company' are clearly documented in the input schema. The description adds no parameter details, but the schema carries that burden effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource — 'Get a single tax code from QuickBooks Online by Id.' The 'single' and 'by Id' wording clearly distinguishes it from sibling list/search tools like search_tax_codes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the right use case: fetch one tax code when you know its QuickBooks Id. It does not explicitly name alternatives such as search_tax_codes or state when not to use this tool, so usage context is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tax_paymentget tax paymentA
Read-only
Inspect

Get a single tax payment from QuickBooks Online by Id. Sales tax (GST/HST, QST, VAT) payments and refunds recorded against filed returns; AU, CA and UK companies only.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the TaxPayment.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal a safe read operation, and the description adds meaningful behavior beyond that: it specifies the operation is for a single record, includes both payments and refunds, and is restricted to AU, CA, and UK companies. This helps the agent understand eligibility and scope without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. The core action and lookup method are front-loaded, followed only by the essential scope and eligibility details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple two-parameter schema, read-only annotations, and the clear country constraint, the description provides what an agent needs to select and invoke the tool. An explicit pointer to search_tax_payments for discovering IDs would be a minor enhancement but is not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameters 'id' and 'company' are already well documented. The description adds little beyond confirming the ID is the lookup key and the optional company context, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get a single tax payment from QuickBooks Online by Id.' It differentiates this from sibling search tools like search_tax_payments by emphasizing 'single' and 'by Id,' and clarifies the scope as sales tax payments/refunds for AU, CA, and UK companies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied: use this when you need one tax payment by its QuickBooks Id. It does not explicitly state when to prefer alternatives such as search_tax_payments, nor does it name exclusions or prerequisites beyond the country limitation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tax_rateget tax rateB
Read-only
Inspect

Get a single tax rate from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the TaxRate.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the behavioral trait that this returns a single tax rate (not a list), which goes slightly beyond the schema. However, it does not describe error behavior, authentication needs, or response format, but with annotations present this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, fluff-free sentence. The resource and operation are front-loaded, and it is appropriately sized for a simple getter tool. Every word adds some value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id tool with readOnly annotations and fully described parameters, this is almost sufficient. However, it does not mention the optional company parameter's role in multi-company setups nor what the return value looks like (though no output schema exists). An agent can likely call it correctly, but some expectations (e.g., non-existence behavior) are unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (both id and company are described), so the baseline is 3. The description only repeats the id's role ('by Id') and adds no meaning for the optional company parameter or any format details. It does not compensate for anything missing in the schema, but nothing is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a specific resource ('tax rate'), and a scope ('single... by Id'), making it clear this is a single-record retrieval. It distinguishes implicitly from the sibling search_tax_rates (search vs. single by Id) and from other tax-related getters (get_tax_code, get_tax_agency), though it does not name an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives like search_tax_rates. It only implies that the caller needs an Id, but does not state that this tool is appropriate when an Id is already known and a search tool should be used otherwise.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tax_summaryget tax summaryA
Read-only
Inspect

Generate the QuickBooks tax summary (GST/HST, QST, VAT filing lines) for a period. Without agency_id, returns { agencies: [{ agency_id, agency_name, report }] } with one report per tax agency that had activity; with agency_id, returns that agency's report directly. Each report is { columns, column_types, rows } as for the other reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro; accounting_method (Cash|Accrual); agency_id (a TaxAgency Id; omit to get one report per tax agency).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds valuable behavioral context: the two return shapes (aggregated agencies vs. single agency) and the report format ('columns, column_types, rows'). It does not cover auth or rate limits, but those are secondary given the read-only flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero filler. The core action is front-loaded, and the conditional behavior and output format are stated efficiently in the following sentences. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must explain return values, and it does so precisely for both agency_id branches. The report shape is given explicitly and tied to a known convention ('as for the other reports'), making the tool self-sufficient for an agent to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of parameters and already describes agency_id's behavior (omit to get one report per agency). The description reinforces this and adds the output structure, but that is more about return values than parameter semantics. The marginal addition over the schema is modest, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate'), a precise resource ('QuickBooks tax summary'), and the tax types (GST/HST, QST, VAT filing lines). This clearly distinguishes it from sibling tax tools like get_tax_agency, get_tax_payment, or search_tax_payments, which deal with individual entities rather than a filing-line summary report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives usage context (for a period) and explains the agency_id conditional, but does not mention when to prefer this tool over sibling tax tools or provide any exclusions. There is no explicit alternative named, so an agent must infer the selection criteria from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_termget termA
Read-only
Inspect

Get a single term from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Term.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds no further behavioral context such as return format, not-found behavior, or company-selection nuance, but for a simple read operation this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler or repetition. It front-loads the operation, resource, source system, and lookup key, making every word useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only get-by-ID tool, the description is mostly complete: annotations cover safety and the schema covers both parameters. The only minor gaps are the lack of an explicit link to search_terms when the ID is unknown and no mention of the returned object shape, but these are not critical for a single-Term fetch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both id and company already documented clearly in the input schema. The description adds no additional parameter meaning beyond restating the id lookup, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies the verb 'Get', the resource 'a single term from QuickBooks Online', and the lookup mechanism 'by Id'. This clearly distinguishes it from search_terms, which searches rather than fetches by ID, and from create/update_term variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when the agent already has a QuickBooks term Id, but it does not explicitly state when to use this tool versus search_terms. No alternatives, exclusions, or when-not guidance are provided, so the agent must infer context from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_time_activityget time activityA
Read-only
Inspect

Get a single time activity from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the TimeActivity.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description's 'Get' aligns with those signals. The description adds the source system and ID-based lookup as context, but does not disclose return behavior or error handling. With annotations covering the safety profile, this is adequate but not enriched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. Every word contributes to identifying what the tool does and how it is invoked.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple point-lookup tool with full schema coverage and read-only annotations, so the description is essentially complete for invocation. It does not describe the return object or not-found behavior, but for a straightforward get-by-ID operation with supporting annotations, the gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'id' and 'company' already documented. The description only reinforces the 'by Id' retrieval concept already present in the schema, so it adds no new parameter-level meaning beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get'), names the resource ('single time activity'), identifies the system ('QuickBooks Online'), and specifies the lookup key ('by Id'). It clearly distinguishes itself from search_time_activities by indicating point retrieval by identifier rather than searching or filtering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by Id' gives clear context for when to use this tool: when the caller already has the QuickBooks TimeActivity ID. It does not explicitly name an alternative for searching, but the intended usage is clear and no misleading guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transaction_detail_by_accountget transaction detail by accountA
Read-only
Inspect

Generate the QuickBooks transaction detail by account report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several); account, account_type, transaction_type; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint/destructiveHint annotations by explaining output row types, id alignment differences between summary and detail reports, Money-cell representation, omission of all-zero rows via hidden_zero_rows, confirmed filters behavior, and the no_data flag. This is rich behavioral context that an agent cannot infer from the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first clause, and every subsequent clause adds necessary behavioral detail about the return format. Despite being dense, there is no filler; each part of the description earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully covers the return contract: columns, rows, types, ids, hidden_zero_rows, filters, and no_data. Combined with the 100%-covered input schema and safe read annotations, an agent has everything needed to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all parameters. The description adds no per-parameter guidance; its mention of filters is about output behavior rather than how to supply parameter values. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks transaction detail by account report.' It also differentiates itself from sibling transaction-list tools by detailing the report-style return shape (section/data/total rows, depth, id label cells), making the tool's identity unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the purpose statement: an agent can infer this is the tool for a QuickBooks transaction-detail-by-account report. However, there is no explicit 'use this when' guidance, no exclusions, and no mention of alternatives such as get_transaction_list, get_general_ledger, or the other transaction-list siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transaction_listget transaction listA
Read-only
Inspect

Generate the QuickBooks transaction list report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); start_duedate, end_duedate (YYYY-MM-DD); customer, vendor, department, name (customer, vendor or employee Ids); transaction_type (Invoice, ReceivePayment, Bill, BillPaymentCheck, Check, CreditCardCharge, Deposit, JournalEntry, Transfer, SalesReceipt, CreditMemo, Estimate, PurchaseOrder, VendorCredit, ...), source_account_type (Bank, AccountsReceivable, AccountsPayable, Income, Expense, CreditCard, ...), docnum, memo, cleared (Cleared|Uncleared|Reconciled), arpaid/appaid (All|Paid|Unpaid), bothamount (an exact amount), payment_method, term, group_by (Name|Account|Transaction Type|Customer|Vendor|Employee|Location|Payment Method|Day|Week|Month|Quarter|Year|None), start_createdate/end_createdate, start_moddate/end_moddate; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds substantial behavioral detail: the exact output structure (columns, column_types, rows), row types, id alignment differences between summary and detail reports, handling of all-zero rows via hidden_zero_rows, the meaning of the filters field, and the no_data flag. This goes well beyond the annotation baseline and helps the agent predict edge cases. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it starts with the purpose, then details the return structure, and covers special cases (hidden_zero_rows, filters, no_data). Every sentence adds value, and the structure is logical. It is longer than some descriptions, but the complexity of the report output justifies the length. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, the description fully specifies the return format, row types, cell semantics, id resolution, and edge-case behavior. Parameters are fully documented in the schema. The tool is a read-only report generator, so the description covers everything an agent needs to call it correctly and interpret the result. It is complete for its purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the params field lists every relevant QuickBooks filter, grouping, and sort option with examples. The description itself adds no additional parameter-specific guidance, so it stays at the baseline of 3. It does not hinder usage, but the description does not elevate parameter understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates the QuickBooks transaction list report and specifies the exact return structure. It is distinct from siblings like get_transaction_list_by_customer or get_transaction_list_by_vendor by its general scope, but it does not explicitly contrast itself with these related tools. Still, the verb-resource pairing is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose this tool over its siblings (e.g., get_transaction_list_by_customer, get_transaction_list_by_vendor, get_transaction_detail_by_account). It does not state filters that would route an agent to a more specific report or mention any exclusions. This is a clear gap for an agent selecting among many get_* report tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transaction_list_by_customerget transaction list by customerA
Read-only
Inspect

Generate the QuickBooks transaction list by customer report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); start_duedate, end_duedate (YYYY-MM-DD); customer, department; transaction_type (Invoice, ReceivePayment, Bill, BillPaymentCheck, Check, CreditCardCharge, Deposit, JournalEntry, Transfer, SalesReceipt, CreditMemo, Estimate, PurchaseOrder, VendorCredit, ...), source_account_type (Bank, AccountsReceivable, AccountsPayable, Income, Expense, CreditCard, ...), docnum, memo, cleared (Cleared|Uncleared|Reconciled), arpaid/appaid (All|Paid|Unpaid), bothamount (an exact amount), payment_method, term, group_by (Name|Account|Transaction Type|Customer|Vendor|Employee|Location|Payment Method|Day|Week|Month|Quarter|Year|None), start_createdate/end_createdate, start_moddate/end_moddate; sort_by (tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_order (ascend|descend); this report does not accept a columns parameter.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral detail beyond annotations: it explains the return format, row types, hidden_zero_rows, filters, and no_data semantics. It also clarifies that ignored parameters are absent from filters. With readOnlyHint and destructiveHint already set, the description enriches the behavioral profile without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense paragraph that front-loads the purpose and then details output structure and edge cases. Every sentence contributes necessary information without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a report tool with complex output and many parameters, the description is complete. It explains the return object, row types, ids, filtering behavior, and no_data flag. Since there is no output schema, the description carries the full burden of explaining the return value, and it does so thoroughly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% description coverage for all parameters, so the baseline is 3. The description does not add parameter-specific meaning beyond what the schema already explains; it only notes that a columns parameter is not accepted, which is also in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate the QuickBooks transaction list by customer report.' It specifies the verb and resource, and details the output structure. It distinguishes itself from sibling tools like get_transaction_list and get_transaction_list_by_vendor by focusing on the customer dimension.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and description ('transaction list by customer'), but there is no explicit guidance on when to use this tool versus alternatives such as get_transaction_list_by_vendor or get_transaction_list_with_splits. No exclusions or alternative conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transaction_list_by_vendorget transaction list by vendorA
Read-only
Inspect

Generate the QuickBooks transaction list by vendor report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); start_duedate, end_duedate (YYYY-MM-DD); vendor, department; transaction_type (Invoice, ReceivePayment, Bill, BillPaymentCheck, Check, CreditCardCharge, Deposit, JournalEntry, Transfer, SalesReceipt, CreditMemo, Estimate, PurchaseOrder, VendorCredit, ...), source_account_type (Bank, AccountsReceivable, AccountsPayable, Income, Expense, CreditCard, ...), docnum, memo, cleared (Cleared|Uncleared|Reconciled), arpaid/appaid (All|Paid|Unpaid), bothamount (an exact amount), payment_method, term, group_by (Name|Account|Transaction Type|Customer|Vendor|Employee|Location|Payment Method|Day|Week|Month|Quarter|Year|None), start_createdate/end_createdate, start_moddate/end_moddate; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the readOnlyHint annotation by thoroughly detailing output behavior: row types, depth, value alignment, Money cell representation, id semantics for summary vs detail reports, omitted all-zero rows with hidden_zero_rows, confirmed filters, and the no_data flag. This gives an agent a clear mental model of the tool's behavior and edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: the opening sentence states the action, and the rest efficiently describes return structure and edge cases. It is front-loaded, uses compact notation, and avoids repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema and the tool's complex return shape, the description does an excellent job explaining return values, row layout, id alignment, and special cases like hidden zero rows and no_data. It is slightly less complete in providing usage context or selecting among sibling reporting tools, but the core call-and-response semantics are well covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes all parameters with 100% coverage, so the baseline is 3. The description adds some useful context by explaining how filters in the response reflect applied vs ignored parameters, but it does not elaborate further on parameter meanings or combinations beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Generate the QuickBooks transaction list by vendor report.' This is a specific verb and resource, and the 'by vendor' qualifier naturally distinguishes it from sibling report tools such as get_transaction_list_by_customer or get_transaction_detail_by_account.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied by the 'by vendor' phrasing and the report name, so an agent can infer this is for vendor-focused transaction reporting. However, the description gives no explicit guidance about when to choose this tool instead of the many closely related sibling list/report tools, nor does it state any exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transaction_list_with_splitsget transaction list with splitsA
Read-only
Inspect

Generate the QuickBooks transaction list with splits report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); transaction_type, source_account_type, docnum, name, payment_method, group_by; columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only and non-destructive. The description adds valuable behavioral context: row types, depth semantics, id alignment, hidden_zero_rows counting, filter confirmation behavior, and the no_data flag. These details go well beyond the annotations and help an agent interpret the response correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loaded with the purpose, followed by return structure and edge cases. Every clause serves a purpose, though the single-paragraph format with many semicolon-separated details could be easier to scan if structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, the description thoroughly compensates by explaining the response shape, filtering behavior, and no-data handling. It is nearly complete for a read-only report tool, though it omits explicit guidance on when to use this versus related transaction-list reports.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already fully documents the params object and company field. The description focuses on return semantics rather than parameter behavior; with full schema coverage, baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it generates the QuickBooks transaction list with splits report, with a specific verb and resource. It does not explicitly differentiate from siblings like get_transaction_list or get_transaction_list_by_customer, though the 'with splits' qualifier provides some distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to choose this tool over sibling transaction-list tools, nor any exclusions or alternative routing. The intended use is only implied by the report name and the detailed return semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transferget transferA
Read-only
Inspect

Get a single transfer from QuickBooks Online by Id. An accounting entry recording a transfer between accounts. It does not move money between bank accounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Transfer.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds meaningful domain context beyond annotations by clarifying that a transfer is an accounting entry and 'does not move money between bank accounts,' which helps agents avoid conceptual misuse even though it is a read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core operation and identification method are front-loaded, and the additional accounting clarification is brief and relevant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with two parameters, one required, and annotations covering safety, the description provides enough context to select and invoke it correctly. It could mention return value shape or explicitly reference search_transfers, but those are minor gaps given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both id and company are already documented in the schema. The description does not add parameter-level detail, but with full schema coverage the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb and resource: 'Get a single transfer from QuickBooks Online by Id.' It distinguishes itself by targeting a single transfer rather than a list, and the added domain context 'accounting entry recording a transfer between accounts' clarifies exactly what entity is being fetched.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a specific transfer Id is already known, but it does not explicitly contrast with search_transfers or mention when to use an alternative. The 'by Id' phrasing offers some guidance, but there is no clear when-to-use versus when-not-to-use statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_trial_balanceget trial balanceA
Read-only
Inspect

Generate the QuickBooks trial balance report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description goes well beyond that by disclosing subtle behaviors: all-zero rows are omitted and counted in hidden_zero_rows, filters lists only the filters QuickBooks confirmed it applied (ignored parameters are absent), and no_data: true signals an empty period. It also explains the id alignment difference between summary and detail reports. This is rich behavioral context that an agent cannot infer from the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads the core purpose and return structure before diving into edge-case behaviors. Every sentence carries meaningful information about return format, row semantics, or filter behavior. It is longer than strictly necessary, but the density justifies the length; a slight structural improvement would be splitting the return-format details from the edge-case behaviors.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only report tool with 100% schema coverage and no output schema, the description is remarkably complete. It explains the return structure, row types, value alignment, id semantics, hidden zero rows, filter confirmation behavior, and the no_data flag. An agent has everything needed to invoke the tool and interpret its response correctly. The only minor gap is not explicitly listing the report's default date range when no date parameters are provided, but the schema's optional params make that a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (params and company) in detail, including the date_macro enum list and accounting_method options. The description adds context about how filters are reported back (filters lists only confirmed applied filters), which indirectly clarifies parameter behavior, but it doesn't add new meaning to the parameters themselves beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the QuickBooks trial balance report.' It then details the exact return shape (columns, column_types, rows, row types, depth, values, id semantics), which distinguishes it from sibling report tools like get_balance_sheet or get_general_ledger. The level of specificity leaves no ambiguity about what this tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this is the tool to call when a trial balance report is needed, and the detailed return format helps an agent understand what to expect. However, it does not explicitly state when to prefer this over related financial report siblings (e.g., get_general_ledger, get_balance_sheet) or mention any exclusions. The context is clear but the when-not-to-use guidance is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vendorget vendorA
Read-only
Inspect

Get a single vendor from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Vendor.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide the safety profile (readOnlyHint=true, destructiveHint=false). The description adds useful scope information — that it returns exactly one vendor selected by Id — which is consistent with the annotations. However, it does not disclose additional behavioral traits such as not-found behavior, rate limits, or return format, though for a simple read operation this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It conveys the action, target system, and identifier in a compact way that is easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only operation with only two parameters, one required, the description plus schema and annotations are sufficient for an agent to invoke the tool correctly. The only minor gap is that no output schema exists and the description does not hint at the structure of the returned vendor object, but this is a small omission for such a simple retrieve-by-id tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, as both 'id' and 'company' have clear textual descriptions in the input schema. The description only echoes 'by Id' and adds no extra meaning beyond what the schema already provides, so the baseline score applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses the specific verb 'Get', names the resource 'a single vendor from QuickBooks Online', and the access method 'by Id'. This clearly distinguishes it from search/list tools like search_vendors and from vendor report tools such as get_vendor_expenses, so an agent can tell exactly what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'by Id' phrasing implies that this tool should be used when a specific vendor Id is known, and the required 'id' parameter confirms this. However, the description does not explicitly mention alternatives like search_vendors when no Id is available, nor does it state exclusion criteria. Usage is inferred rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vendor_balanceget vendor balanceA
Read-only
Inspect

Generate the QuickBooks vendor balance report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); vendor, department, appaid (All|Paid|Unpaid).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by explaining the exact return structure, row types, id alignment rules, hidden_zero_rows behavior, filter confirmation semantics, and no_data flag. It adds substantial behavioral context without contradicting the readOnlyHint and destructiveHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core verb and resource, and every sentence contributes to call correctness or result interpretation. Despite its length, there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by documenting exactly what the tool returns: columns, column_types, rows, row types, depth, id alignment, hidden_zero_rows, filters, and no_data. It also addresses edge cases like ignored parameters and empty periods, making the tool safe to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful parameter behavior: 'filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent)' clarifies how ignored vendor/department parameters surface in the result. It also explains no_data in relation to the report period, going beyond the schema's value lists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate the QuickBooks vendor balance report.' It clearly identifies the tool's function, though it does not explicitly differentiate it from close siblings like get_vendor_balance_detail or get_aged_payables, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_vendor_balance_detail or get_aged_payables. The description implies this is a summary-level report through row types, but it never states the selection criteria or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vendor_balance_detailget vendor balance detailA
Read-only
Inspect

Generate the QuickBooks vendor balance detail report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: report_date (YYYY-MM-DD, the as-of date; defaults to today) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); start_duedate, end_duedate (YYYY-MM-DD); vendor, department, term, appaid (All|Paid|Unpaid); columns (comma-separated, e.g. tx_date,txn_type,doc_num,name,memo,account_name,other_account,due_date,pmt_mthd,is_cleared,dept_name), sort_by (one of those columns), sort_order (ascend|descend).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint/destructiveHint annotations by disclosing the exact return shape, row types, Money cell representation, id alignment behavior, omission of all-zero rows with hidden_zero_rows count, filter confirmation semantics, and no_data flag. This is rich behavioral context that an agent needs to interpret the response correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and information-rich, front-loading the report generation purpose before diving into return-format details. Every sentence adds value, though the long run-on sentence about return structure could be slightly better organized. It is appropriately sized for the complexity of the output format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only report tool with no output schema, the description is remarkably complete: it explains the return structure, row types, value alignment, id behavior, zero-row handling, filter confirmation, and no-data signaling. The parameter schema covers all inputs. Nothing an agent needs to call this tool and interpret its response is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds context about how filters appear in the response (filters lists confirmed filters, ignored parameters are absent), which helps understand parameter behavior, but it does not add syntax or format details beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and resource ('QuickBooks vendor balance detail report'), and the detailed return-format explanation distinguishes it from sibling report tools like get_vendor_balance, get_aged_payable_detail, and get_vendor_expenses. It clearly identifies what the tool produces and how the output is structured.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by naming the report type and listing filter parameters, but it does not explicitly state when to prefer this over sibling report tools (e.g., get_vendor_balance vs get_vendor_balance_detail). The parameter schema provides the filter options, but there is no explicit when/when-not guidance or alternative tool names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vendor_creditget vendor creditA
Read-only
Inspect

Get a single vendor credit from QuickBooks Online by Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the VendorCredit.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the single-record scope and Id-based lookup semantics, but offers no additional behavioral context such as not-found handling, return shape, or connection requirements. Consistent with annotations and adds modest value beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler; every word carries information. The verb and resource are front-loaded, and the scoping qualifier ('single', 'by Id') immediately follows. Nothing could be cut without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter, read-only point lookup with no nested objects and full schema coverage, the description covers the essentials. The only notable gap is the absence of routing guidance among the large set of vendor/get/search siblings; return-value details are self-evident for a 'Get' tool and no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both id and company fully described in the input schema. The tool description reinforces that lookup is by Id but adds no new parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), resource ('a single vendor credit'), system ('QuickBooks Online'), and lookup method ('by Id'). The 'single ... by Id' phrasing functionally distinguishes it from search_vendor_credits (criteria-based list lookup) and other vendor-related getters like get_vendor_balance or get_vendor_expenses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is only implied: the phrase 'by Id' signals that this tool is for point lookups when the vendor credit Id is already known, as opposed to searching. However, it never names alternatives (e.g., search_vendor_credits for finding by criteria) or states any exclusions, leaving the agent to infer the boundary itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vendor_expensesget vendor expensesA
Read-only
Inspect

Generate the QuickBooks vendor expenses report. Returns { columns, column_types, rows }: each row has a type (section | data | total), a depth, and values aligned to columns (Money cells as numbers, first cell is the label on summary reports); id is the label cell's QuickBooks record id (e.g. the account); on detail reports ids aligns to values instead. All-zero rows are omitted and counted in hidden_zero_rows. filters lists the filters QuickBooks confirms it applied (a parameter it ignored is absent), and no_data: true means the period has nothing to report.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsNoQuickBooks report query parameters: start_date, end_date (YYYY-MM-DD) or date_macro (Today, Yesterday, This Week, Last Week, This Month, Last Month, This Month-to-date, This Fiscal Quarter, Last Fiscal Quarter, This Fiscal Year, Last Fiscal Year, This Fiscal Year-to-date, Last Fiscal Year-to-date, ...); accounting_method (Cash|Accrual); summarize_column_by (Total|Month|Quarter|Year|Week|Days|Classes|Departments|Customers|Vendors|Employees|ProductsAndServices); the filters customer, vendor, item, class, department (QuickBooks Ids, comma-separated for several).
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint and destructiveHint annotations by detailing return shape, row types, id alignment behavior, hidden_zero_rows counting, filter application semantics, and no_data handling. It also discloses edge cases like ignored parameters being absent from filters, which is valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence, but every clause carries useful information about report output structure and edge cases. It is front-loaded with the tool's purpose and then covers return semantics efficiently. It could be split into shorter sentences for readability, but there is little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description comprehensively explains the return structure, row types, id alignment, zero-row omission, filter reporting, and no_data state. The input schema already covers parameters, so an agent has everything needed to correctly invoke and interpret this read-only report tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining how filter parameters are reflected in the output: filters lists what QuickBooks confirmed applied, and an ignored parameter is absent. This helps the agent interpret parameter effects and diagnose failed filtering.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action and resource: 'Generate the QuickBooks vendor expenses report.' This distinguishes it from vendor balance and transaction-list tools by name. However, it does not explicitly differentiate itself from sibling reporting tools such as get_customer_income or get_profit_and_loss.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose this tool over alternatives, nor any exclusions or prerequisites. The intended use is only implied by the report name 'vendor expenses.' There is no mention of when a different report like get_vendor_balance_detail would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_connected_companiesList connected companiesA
Read-only
Inspect

List the QuickBooks companies connected to this Caribooks account, with their access level.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, establishing the safety profile. The description adds the access-level detail, which is useful but minor. It does not mention pagination, return format, or authentication requirements, but the annotations lower the burden for this simple read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the verb and resource, with the access-level qualifier adding necessary output information. Every word earns its place; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters, no output schema, and read-only annotations, the description is fully adequate. It states what the tool lists and the key attribute of the results. An agent has everything it needs to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description has no parameter burden. Per the baseline for 0-parameter tools, this is a 4. The description does not attempt to document nonexistent parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List'), identifies the resource ('QuickBooks companies connected to this Caribooks account'), and adds a distinguishing detail ('with their access level'). This clearly separates it from every other list_* sibling and leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are no sibling tools that list connected companies, so the unique resource makes the usage context clear. The description does not explicitly state exclusions or alternatives, but none are needed; an agent can infer it should be used whenever it needs to enumerate connected QuickBooks companies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_loopsList the loopsA
Read-only
Inspect

The loops this account has for a company, or for every company when none is named: each one's id, whether it is active or paused, the sentence it came from, what it does read back in the user's words, and when it last ran and runs next. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected company. Every company when left out.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly says 'Read-only', which aligns with the annotations (readOnlyHint: true) and adds clarity. It also mentions the specific fields returned (id, active/paused status, sentence origin, user-words read-back, last/next run times), which is helpful context beyond the annotations. No contradictions found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that packs all necessary information without any fluff. It front-loads the primary function (listing loops) and then provides scope and output details efficiently. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple nature of the tool (one optional parameter, no output schema needed because the description itemizes the return fields), the description is complete enough for an agent to call it correctly. It covers scope, filtering, and output composition, leaving little ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides complete descriptions for the single parameter 'company' (coverage 100%), so the description adds little additional meaning. However, the description's phrase 'or for every company when none is named' reinforces the optionality and default behavior, which is a slight enhancement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that this tool lists loops for a company or for all companies when none is named, which distinguishes it from related sibling tools like list_proposals and list_rules. While it doesn't explicitly name a sibling, the scope (loops vs. proposals vs. rules) is clear and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: to retrieve loops for a specific company or across all companies. It doesn't explicitly state when not to use it or mention alternatives, but the context is clear enough for an agent to select it over siblings like search_* tools when a simple list is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_proposalsList what the loops proposeA
Read-only
Inspect

List a company's proposed or completed bookkeeping changes, newest first. Returns each proposal's id, action type, approval mode, status, summary and supporting evidence. Pending proposals can be approved or rejected by the user in the Caribooks portal Review tab or through the daily digest link. This tool is read-only and cannot approve or reject proposals.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoKeep only the proposals waiting for the user, the ones carried out, or the ones that failed. All of them when left out.
companyNoWhich connected company. Every company when left out.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and destructiveHint, and the description reinforces this with 'This tool is read-only and cannot approve or reject proposals.' It adds beyond annotations by specifying returned fields, ordering ('newest first'), and where approvals happen, providing useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, front-loaded with the core action, then return fields, then behavioral constraints and user guidance. Every sentence contributes value with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with optional parameters and no output schema, the description covers the purpose, returned fields, ordering, read-only nature, and user approval path. It is complete enough for an agent to invoke it correctly; minor details like pagination are not essential here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add meaningful detail about the status or company parameters beyond what the schema already provides; it only loosely refers to 'a company's' proposals and 'proposed or completed' statuses.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'List a company's proposed or completed bookkeeping changes, newest first.' It also enumerates returned fields. However, it does not explicitly distinguish itself from siblings like list_loops, so it lacks full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context and an explicit exclusion: it is read-only and cannot approve or reject proposals, and it points users to the Caribooks portal Review tab or daily digest link for those actions. It does not, however, explain when to choose this tool over related tools such as list_loops.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_receipt_inboxList the receipt inboxA
Read-only
Inspect

List receipts, invoices and documents uploaded through the Caribooks portal or emailed to the account's receiving address. Returns each file's upload_id, company, kind, filing status and content preview. Kinds include transaction evidence, monthly statements and reference documents. Read-only; does not create transactions or attachments. include_filed includes filed statements and references. upload_id selects a document for a full-text read; text_offset pages through long documents, and cursor pages through older files. search, from, to, min_total and max_total find documents by the words, date and total read off them, filed receipts included; each line shows the vendor, date, number, total and summary read.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoOnly documents dated on or before this day, YYYY-MM-DD.
fromNoOnly documents dated on or after this day, YYYY-MM-DD, by the date written on them.
limitNoFiles per page, newest first. Defaults to 10.
cursorNoThe next cursor from a listing. Keep the same company and include_filed when continuing.
searchNoWords every matching document contains, in its vendor, number, summary, lines bought, text, file name, sender or subject. Case and accents do not matter. Searches filed documents too.
companyNoOnly the files sorted into this connected company. Every file when left out, each line saying which company it belongs to.
max_totalNoOnly documents whose total is at most this amount.
min_totalNoOnly documents whose total is at least this amount.
upload_idNoRead this file instead of listing previews. PDFs return up to 12,000 characters with a continuation until the end.
text_offsetNoContinue reading upload_id at the text_offset returned by the previous call. Defaults to 0.
include_filedNoAlso list filed statements and references. A named file includes these automatically; filed receipts live in QuickBooks.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and destructiveHint, and the description reinforces this with 'Read-only; does not create transactions or attachments.' Beyond annotations, it discloses cursor pagination, full-text continuation via upload_id/text_offset, and search behavior over filed receipts. No contradiction with the annotations is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient: purpose and return values are front-loaded, and the parameter-related sentences earn their place. The semicolon-heavy style packs a lot into seven sentences, and while it could be reorganized into clearer sections, there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 11 parameters, the description carries a heavy burden, and it covers the key ground: document sources, return fields, file kinds, read-only behavior, filtering, search, and pagination. It does not spell out edge cases like missing upload_id behavior, but the essential information an agent needs to invoke the tool correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by grouping search/from/to/min_total/max_total as a coherent filtering behavior and explaining the output line format ('vendor, date, number, total and summary read'). It also clarifies the relationship among upload_id, text_offset, and cursor, which is not obvious from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List receipts, invoices and documents uploaded through the Caribooks portal or emailed to the account's receiving address.' It also names what is returned (upload_id, company, kind, filing status, content preview), which makes the tool's purpose unmistakable. The inbox scope distinguishes it from the many search/get/create siblings in the surrounding tool list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: it is read-only, does not create transactions or attachments, and notes that filed receipts 'live in QuickBooks' when include_filed is used. It does not explicitly name alternative tools or state when not to use this tool, so it stops short of a 5, but the usage context is unambiguous enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_rulesList a company's rulesA
Read-only
Inspect

List saved bookkeeping rules for a connected company, including active rules and proposals awaiting user approval. Returns each rule's id, source and evidence. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected company. Optional when only one is connected.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly states 'Read-only', which aligns with the annotations (readOnlyHint=true, destructiveHint=false). It adds value by clarifying it includes proposals and the type of data returned, which is beyond the annotations. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, two sentences, with the most critical information (what it lists and read-only) front-loaded. No unnecessary words. Every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read-only list with one optional parameter. The description covers the purpose, scope (active and proposals), and return fields. It lacks explicit examples or related siblings, but given the simplicity and annotations, it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage for the 'company' parameter, explaining it is optional and which company. The description does not add additional parameter semantics beyond the schema, but the schema is sufficient. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists saved bookkeeping rules for a connected company, specifying the scope ('including active rules and proposals awaiting user approval') and the return fields ('id, source and evidence'). It distinguishes from siblings like list_proposals and list_loops by focusing on rules and their approval status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you need to view saved rules or pending proposals. However, it does not explicitly state when not to use it or mention alternatives like list_proposals or search functions. The context of 'connected company' is provided but not elaborately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remember_ruleRemember a rule about how this company's books are keptA
Destructive
Inspect

Save a user's bookkeeping preference for a connected company: a categorization, vendor fact, preference or procedure. For example, a categorization can associate a payee with an expense account. Every rule saved here waits for the user's approval on the Caribooks Rules page (https://caribooks.com/portal/rules) and applies only once approved there; tell the user so. stated_by_user records whether the user said it or it was inferred from a pattern. Rules are company-specific business data; they never change access permissions or confirmation requirements.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
textYesThe rule in the user's words, one sentence.
companyNoWhich connected company the rule belongs to. Optional when only one is connected.
matcherNoWhat the rule matches, as flat pairs: {"memo_contains": "UBERTRIP"}, {"payee": "Bambu Lab"}, {"amount_over": "500"}.
outcomeNoWhat to do when it matches, as flat pairs: {"account": "5245 Travel"}, {"class": "Formation A"}.
evidenceNoThe record that prompted the rule, e.g. Purchase 2979.
stated_by_userYestrue only when the user said it in this conversation; false when you inferred it.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses non-obvious behavior: saved rules wait for user approval and only apply after approval, and it adds the safety-relevant detail that rules never change access permissions or confirmation requirements. The destructiveHint annotation is not explained in terms of specific side effects, but the description still adds meaningful behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with each sentence earning its place: the main purpose, an example, the approval workflow, the stated_by_user distinction, and the safety reassurance about permissions. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with good schema coverage and no output schema, the description plus schema is nearly sufficient: it covers what the tool does, examples of supported rule kinds, and the crucial approval-gated behavior. It lacks explicit guidance on what the agent should do after the save completes, but that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high (86%), so the schema does most of the parameter documentation work. The description adds useful semantic color, such as the categorization example linking a payee to an expense account and the clarification of stated_by_user. It does not deeply explain the nested matcher/outcome objects beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action "Save a user's bookkeeping preference for a connected company" and enumerates the kinds of rules it covers (categorization, vendor fact, preference, procedure). It clearly identifies the resource (rules) and gives a concrete example. It does not explicitly differentiate it from sibling tools like list_rules or retire_rule, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys the relevant context: use this tool when a user states a bookkeeping preference, and it instructs the agent to mention the approval requirement on the Caribooks Rules page. It does not explicitly state when not to use this tool or point to alternative tools, so it lacks the exclusion guidance needed for a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retire_ruleRetire a ruleA
Destructive
Inspect

Retire a saved bookkeeping rule or decline a proposed rule by id, following the user's request. The rule remains in history but is no longer active.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected company. Optional when only one is connected.
rule_idYesThe rule's id, as shown by list_rules.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true, but the description adds valuable context: the rule is not permanently deleted but becomes inactive while remaining in history. It also notes that the action follows the user's request, which is a behavioral cue. This goes beyond the annotation's binary safety flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that packs the core action, the two target kinds of rules, and the post-condition. No filler; the most important information (what it does and what happens) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with full schema coverage and no output schema, the description covers the essential behavior. It states what is retired and the result. It could mention permission requirements or reversibility, but those are not critical for this operation and the description is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% – both parameters (company and rule_id) are already described in the schema. The description adds no extra parameter semantics, so the baseline of 3 applies. It correctly references 'rule_id' but doesn't elaborate beyond the schema's explanation that it's shown by list_rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (retire) on a specific resource (saved bookkeeping rule or proposed rule) and adds the effect ('remains in history but is no longer active'). It distinguishes from other rule tools by specifying 'retire or decline' vs create/remember/list. The purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (when the user wants to retire or decline a rule) but does not explicitly name alternatives or state when NOT to use it. It provides clear context but lacks explicit exclusions or comparison to sibling tools like remember_rule or list_rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_accountssearch accountsA
Read-only
Inspect

Search account records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all account records.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The definition goes well beyond that via the criteria documentation: page limits (1000 records, truncation), AND-only filters with no OR, case-sensitive field values, unavailable data (government IDs, birth dates, card details), and inactive-record exclusion. This is rich behavioral context consistent with the annotations, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with zero wasted words and the core purpose front-loaded. The minimalism is a deliberate trade-off: behavioral detail is delegated to the schema, which is a reasonable structure even if the description alone feels thin.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with two well-documented parameters and no output schema, the definition covers criteria shapes, operators, restrictions, pagination, and the company disambiguation rule thoroughly. The main gap is the absence of any description of the returned account record structure, which an agent would have to infer or discover.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the criteria parameter's prose is exceptionally detailed on accepted shapes, operators, field limitations, and pagination semantics. The description itself adds nothing about parameters, so the baseline 3 applies—the schema carries the load as expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Search account records in QuickBooks Online' names a specific verb, resource, and platform, so an agent knows immediately what the tool does. However, it makes no attempt to differentiate from siblings such as get_account, get_account_list, or the other search_* tools, which all share the same basic 'search entity records' framing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description itself gives no guidance on when to use this tool versus alternatives like get_account or get_account_list. The schema's criteria prose implies usage contexts (filtering, pagination with limit/offset, Active filtering), which is helpful, but there is no explicit when/when-not logic or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_attachablessearch attachablesA
Read-only
Inspect

Search attachment records (files and notes pinned to QuickBooks records). Each result names the file and the records it is attached to. Search results do not include a TempDownloadUri download link.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all attachable records.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly and non-destructive behavior, so the bar is lower. The description adds valuable behavioral context: it discloses that results only name the file and associated records and explicitly excludes a download link. The criteria parameter description further documents pagination behavior (page size 1000, truncation) which supplements the safety profile from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is two tightly written sentences that immediately state the purpose, the content of results, and a key limitation. No unnecessary words, and the important behavioral note about the missing download link is included without clutter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain what is returned—it does so by stating that each result names the file and its attached records, and that no download link is included. Combined with the detailed schema for the criteria parameter (filters, sorting, pagination), an agent has enough context to call the tool correctly, though it stops short of fully describing the result structure beyond that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the criteria parameter has a rich description covering formats, operators, limitations, and pagination. The main tool description does not add any parameter-specific semantics beyond what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('search') and a clear resource ('attachment records (files and notes pinned to QuickBooks records)'). It states exactly what the tool does and what results contain, and explicitly notes that results lack a TempDownloadUri link, distinguishing it from other attachment tools like get_attachable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for searching attachment records but does not explicitly name alternatives or state when not to use it. The mention that results lack a download link hints that this tool is not for downloading, but it doesn't direct the agent to get_attachable or other tools. Guidance is present but only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_bill_paymentssearch bill paymentsA
Read-only
Inspect

Search bill payment records in QuickBooks Online. An accounting record of a payment to a vendor. It does not initiate a bank or card payment.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all bill_payment records.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds behavioral context: it clarifies no bank/card initiation, and the criteria description details limitations (no OR, no Line fields, pagination, truncation, case sensitivity, data unavailability). This goes beyond annotations, so a high score is warranted, though not perfect because it doesn't mention response format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: it opens with the purpose and non-initiation caveat. The criteria parameter documentation is long but necessary given the tool's complexity. It does not waste words, though the length is justified by the utility. Slight deduction for not summarizing key limitations more tightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with extensive schema coverage, the description is highly complete. It covers purpose, non-side-effects, filtering syntax, limitations, and pagination. Missing output schema is fine since it's a search tool. An agent has enough to call it correctly without external documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides very high coverage (100%) for the criteria parameter, but the description adds substantial value by explaining operators, AND semantics, field limitations, pagination, and case-sensitivity. The company parameter is simple and schema-covered. This exceeds schema-only meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it searches bill payment records in QuickBooks Online, specifying it is an accounting record of a payment to a vendor and explicitly notes it does not initiate a bank or card payment. This distinguishes it from payment initiation tools like 'send_payment' or 'create_bill_payment'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for searching/filtering historical payment records, and the extensive criteria documentation covers how to construct filters. It does not explicitly mention when not to use it or alternative search tools (e.g., search_payments), but the criteria details provide strong usage guidance for filtering.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_billssearch billsA
Read-only
Inspect

Search bill records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all bill records.

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The schema's criteria description discloses several non-obvious behaviors: filters are ANDed, QuickBooks lacks OR, cannot filter on Line fields, only top-level fields are filterable, field names are case-sensitive, sensitive data is excluded, and each call returns at most 1000 records with truncation. This goes well beyond the readOnlyHint and destructiveHint annotations, providing crucial context for correct usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately conveys the tool's purpose, with no filler. It's appropriately concise and doesn't repeat information found in the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich criteria description in the schema, the tool definition covers filtering options, pagination, and data exclusions, making it sufficiently complete for an agent to use correctly. The only minor gap is the explicit return format, but that's implied and not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters have descriptive text in the schema, covering the accepted formats, operators, and constraints for criteria, and the optional company parameter. Since the schema provides 100% coverage, the description adds no additional parameter semantics beyond what's already documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the exact verb 'search' and the resource 'bill records' in QuickBooks Online, making the tool's purpose unambiguous. However, it doesn't explicitly contrast with related tools like get_bill or other search_* tools, so it's clear but not fully distinguishing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose this tool over alternatives such as get_bill or other search tools. While the criteria description gives advice on using filters and pagination, it doesn't address tool selection, leaving the agent to infer when search is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_budgetssearch budgetsA
Read-only
Inspect

Search budget records in QuickBooks Online. A profit-and-loss budget with one amount per account per period.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Each budget comes back with every BudgetDetail line. Omit to return the first page of all budget records.

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate readOnlyHint=true and destructiveHint=false, and the description adds substantial behavioral context: pagination behavior ('one page of at most 1000 records'), truncation to fit the answer, filtering limitations ('QuickBooks has no OR, cannot filter on Line fields'), and sensitive-data exclusions ('Government identifiers, birth dates and card-processing details are not available'). These go well beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, with two sentences that both earn their place: one states the action and subject, the other clarifies the budget type. There is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Together with the rich schema and annotations, the description covers key operational constraints: pagination, truncation, filter limitations, and excluded sensitive fields. It also notes that each budget returns every BudgetDetail line. It lacks explicit return-shape detail beyond that and does not mention sibling tool routing, but for a read-only search tool with strong annotations, this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the criteria parameter's schema description is very detailed, covering operators, filter fields, and examples. The tool description itself does not add new parameter meaning beyond what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: 'Search budget records in QuickBooks Online.' It also adds a defining detail ('A profit-and-loss budget with one amount per account per period') that helps distinguish budgets from other entity types. However, it does not explicitly contrast with sibling tools like get_budget or create_budget, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance about when to use search_budgets versus alternatives such as get_budget or get_budget_vs_actuals. The intended use is implied by the name and the phrase 'Search budget records', but the description does not state when to choose this tool or what it is not for. The criteria parameter includes operational constraints, but not alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_classessearch classesA
Read-only
Inspect

Search class records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all class records.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The criteria parameter discloses important behavioral traits beyond the readOnly/destructive annotations: one-call returns at most 1000 records, large pages are truncated, filters are ANDed with no OR support, only top-level fields are filterable, field matching is case-sensitive, inactive records are excluded unless Active is explicitly filtered, and sensitive data is unavailable. This is rich, useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused sentence containing only necessary information: the operation, the entity, and the platform. It is front-loaded and free of filler. The long criteria schema text is justified by the complexity of the query language and does not bloat the narrative description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema, the definition conveys most of what an agent needs: platform scope, company selection, filter syntax, operator semantics, pagination behavior, and inactive-record handling. It stops just short of fully complete because it does not describe class-specific searchable fields or the shape of returned class records, but the filtering documentation largely compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The narrative description adds no parameter-level meaning beyond what the schema already provides. The criteria parameter's schema description is excellent, but it lives in the schema, not in the tool description, so the description itself does not augment it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search') and resource ('class records') with the platform scope 'QuickBooks Online', so an agent knows exactly what action and entity are involved. It distinguishes from non-class tools by naming the resource, though it does not explicitly differentiate from siblings like get_class or search_accounts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The one-line description gives no guidance on when to use search_classes versus get_class for a single record or versus other search_* tools. The criteria schema provides substantial how-to guidance (pagination, filtering limitations, inactive records), but it does not address alternatives or exclusions, leaving usage selection largely implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_company_currenciessearch company currenciesB
Read-only
Inspect

Search company currency records in QuickBooks Online. The currencies the company transacts in; only meaningful with multicurrency enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all company_currency records.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the no-write/no-destructive nature is covered. The description adds the multicurrency context, which is helpful, but it does not disclose operational behaviors like pagination, page-size limits, or filter constraints; those live in the criteria schema, not the description. No contradiction with annotations, but the description adds only moderate behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded with the primary action, and contains no filler. The second sentence about multicurrency adds meaningful context without bloating the definition. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich criteria schema and read-only annotations, the overall definition is sufficient for an agent to call the tool correctly. It lacks an explicit mention of sibling alternatives and does not describe return format, but those gaps are partially mitigated by schema coverage and the simple nature of a search tool. Slightly more routing guidance would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the criteria parameter has a very detailed schema description covering operators, filtering constraints, and pagination behavior. The tool description itself doesn't explain the parameters, so it adds little beyond the schema. Per the rubric, the baseline of 3 applies because the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action and resource: 'Search company currency records in QuickBooks Online.' It adds useful domain context by noting these are the currencies the company transacts in and that the tool is only meaningful with multicurrency enabled. It doesn't explicitly differentiate from siblings like search_exchange_rates or get_company_currency, but the resource name is specific enough to avoid major ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to prefer this tool over alternatives such as search_exchange_rates or get_company_currency. The only usage signal is the multicurrency prerequisite, which is a condition rather than a routing rule. The description does not mention exclusions or alternative tools, leaving the agent to infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_company_infossearch company infosA
Read-only
Inspect

Search company info records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. There is one company info record; get_company_info reads it directly. Omit to return the first page of all company_info records.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe read operation, and the description adds valuable behavioral detail beyond that: QuickBooks has no OR, only top-level fields can be filtered, field names are case-sensitive, sensitive data is unavailable, and results are limited/truncated to one page. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is one efficient sentence. The criteria parameter text is dense but every clause carries a distinct operational constraint or clarification, with no filler or repetition. It is appropriately detailed for a search tool with complex filtering semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition covers search syntax, operator semantics, QuickBooks-specific limitations, pagination and truncation behavior, sensitive-data restrictions, and the relevant sibling alternative. Despite no output schema, the return behavior is described well enough for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema description coverage is 100%, the criteria description substantially enriches the parameter meaning with accepted forms, operators, wildcard syntax, sort/limit/offset options, filtering restrictions, and field examples. The company parameter is also clearly explained with connection disambiguation and optionality.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Search'), the resource ('company info records'), and the system ('QuickBooks Online'). It does not explicitly differentiate itself from the sibling get_company_info, though the schema text does point that out; the main description alone is clear but not distinguishing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The criteria description gives explicit guidance: it notes there is only one company info record, names get_company_info as the direct read, and explains how to page or filter when using this search tool. This effectively tells the agent when the search tool is appropriate and when to prefer a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_credit_card_payment_txnssearch credit card payment txnsA
Read-only
Inspect

Search credit card payment txn records in QuickBooks Online. An accounting record of a bank payment toward a credit card balance. It does not move money.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all credit_card_payment_txn records.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. The description adds valuable behavioral context: it does not move money, it returns one page of at most 1000 records, large pages are truncated, and it cannot filter on Line fields or use OR. This goes beyond the annotations and helps the agent set expectations about the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded with the core purpose. The first sentence states what the tool does, and the second clarifies the domain meaning. The criteria parameter description is long but dense with necessary operational details. It earns its place because it covers filtering syntax, limitations, and pagination in a compact form. Slightly verbose in the criteria section, but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema, the description covers the essential operational details: what records are searched, how to filter, pagination limits, and known limitations. It does not describe the return format or fields returned, but for a search tool with a generic criteria parameter, the description is sufficiently complete for an agent to invoke it correctly. The lack of an output schema is partially mitigated by the detailed criteria documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds meaningful context about the criteria parameter: it explains the accepted formats (array of {field, value, operator}, simple object, or {criteria, asc|desc, limit, offset, count}), the operators, the ANDed filtering, the lack of OR, the case-sensitivity, and the list of filterable fields. This is substantial added value beyond the schema's generic description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search'), a specific resource ('credit card payment txn records in QuickBooks Online'), and clarifies the domain meaning ('An accounting record of a bank payment toward a credit card balance'). It also distinguishes itself from related tools by noting it does not move money, which helps an agent differentiate it from create/update/delete credit card payment tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: searching credit card payment transaction records. It also gives practical guidance on filtering, pagination, and limitations (no OR, no Line fields, case-sensitive fields, one page of at most 1000 records). It does not explicitly name sibling alternatives like get_credit_card_payment_txn or create_credit_card_payment_txn, but the context is strong enough for an agent to infer when to use this search tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_credit_memossearch credit memosA
Read-only
Inspect

Search credit memo records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all credit_memo records.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. The description adds critical behavioral context: it discloses that QuickBooks has no OR support, that filters are ANDed, that field names/values are case-sensitive, that government identifiers/birth dates/card-processing details are unavailable, and that one call returns at most 1000 records with truncation. This is rich behavioral disclosure beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads the core purpose ('Search credit memo records in QuickBooks Online') before diving into criteria details. Every sentence adds value, but the criteria explanation is quite long and could benefit from slight restructuring (e.g., bullet points or separation of pagination from filter syntax). Still, it is appropriately sized for the complexity of the tool and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema, the description covers everything an agent needs: what it searches, how to filter, what limitations exist, pagination behavior, and what happens when criteria is omitted. The annotations cover the safety profile (read-only, non-destructive). The only minor gap is that it doesn't describe the exact return shape, but for a search tool with no output schema, the description's coverage of behavior and constraints is comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds significant meaning beyond the schema: it explains the criteria parameter's flexible formats (array of {field, value, operator}, simple object, or {criteria, asc|desc, limit, offset, count}), lists valid operators, gives example filterable fields, and explains pagination semantics. The company parameter is also clarified as optional when only one company is connected. This goes well beyond the baseline 3 for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Search credit memo records in QuickBooks Online.' It uses a specific verb ('Search') and resource ('credit memo records'), and the title reinforces this. It distinguishes itself from sibling tools like get_credit_memo (which retrieves a single record) and create_credit_memo (which creates), so an agent can easily tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides extensive usage guidance: it explains the criteria parameter formats, operators, filtering limitations (no OR, no Line fields, only top-level fields), case sensitivity, and pagination behavior. It explicitly states what to do when needing more than one page ('narrow with filters or walk it with limit and offset') and what happens if criteria is omitted. This is far beyond a simple 'when to use' statement and gives the agent clear decision-making rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_customerssearch customersB
Read-only
Inspect

Search customer records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all customer records.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint=false, and destructiveHint=false, so safety is covered. The tool description itself adds no behavioral context about pagination, truncation, or filtering limits; those details live in the input schema rather than the description, so this is adequate but not a strong description-level disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The tool description is a single efficient sentence, and the criteria description is dense but information-bearing. It is slightly monolithic as one long paragraph, but every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema's criteria description plus the annotations cover query syntax, operator limitations, field restrictions, case sensitivity, paging and truncation, inactive-record defaults, and sensitive-field omissions. No output schema exists, but the resource is clear from the tool name and description, so the agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline applies. The criteria parameter description is genuinely rich, covering operators, AND semantics, paging, and inactive-record behavior, but that value is in the schema, not the tool description, so the description does not add parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: Search customer records in QuickBooks Online. It is clear enough to distinguish from other search_* tools by entity, though it does not contrast itself with the single-record get_customer sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool-level description gives no guidance on when to prefer this over get_customer, get_customer_balance, get_customer_sales, or other search tools. Usage is only implied by the name and the detailed criteria schema, not stated with any when-to-use or when-not-to-use advice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_customer_typessearch customer typesC
Read-only
Inspect

Search customer type records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all customer_type records.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description merely says 'Search', adding no behavioral nuance beyond what the annotations provide. It does not mention pagination, truncation, or limitations, even though those details exist in the schema's criteria description rather than the tool description itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is concise but does not go beyond the minimal statement of purpose, which prevents a 5 for structural completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema's criteria description carries substantial operational detail (pagination, filtering, case sensitivity), and annotations cover the safety profile. However, there is no output schema and the tool description itself does not describe the return shape or the fields of a customer type record, leaving some inference required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both the company and criteria parameters are already well documented in the input schema. The tool description adds no parameter-specific meaning, which aligns with the baseline 3 for fully covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Search') and the resource ('customer type records in QuickBooks Online'), making the tool's purpose immediately identifiable. It does not explicitly distinguish itself from sibling tools like get_customer_type or search_customers, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus get_customer_type, search_customers, or other search_* siblings. There are no stated conditions, exclusions, or alternative tool references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_departmentssearch departmentsC
Read-only
Inspect

Search department records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all department records.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The top-level description adds no behavioral traits beyond what the annotations already provide: it merely says the tool searches records. Important behaviors like no-OR filtering, one-page-at-most-1000 truncation, case sensitivity, and inactive-record exclusion are present only in the criteria parameter description, not in the tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded, waste-free sentence: 'Search department records in QuickBooks Online.' It is concise and easy to parse, though it is brief enough that it omits useful behavioral context that would make it genuinely helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The detailed criteria schema makes the definition minimally viable, but with no output schema and no tool-level mention of pagination, default inactive filtering, or the get_department alternative, an agent must infer important invocation context. It is adequate but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the criteria parameter already explaining operators, ANDed filters, pagination, and inactive-record behavior. The tool description adds no parameter-level meaning, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Search department records in QuickBooks Online.' It is clear about what the tool does, but it does not differentiate this tool from the sibling get_department or other search_* tools, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as get_department. It does not state exclusions, prerequisites, or recommended use cases; the only substantive usage detail lives inside the criteria parameter schema, not in the tool description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_depositssearch depositsB
Read-only
Inspect

Search deposit records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all deposit records.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the read-only safety profile is covered. The description adds no behavioral context beyond the word 'search' — pagination, truncation, and QuickBooks data limitations appear only in the input schema, and nothing contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler. It is appropriately concise, though it adds little substance beyond what the title already conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The rich input schema carries most of the operational weight, documenting criteria syntax, filter restrictions, and one-page truncation. However, with no output schema and no sibling routing in the description, the agent gets no return-shape guidance and only implied selection cues, leaving a moderate completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters have detailed documentation, especially the criteria parameter with operators, AND semantics, field limitations, case sensitivity, and pagination. The tool description adds zero parameter meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the action and object ('Search deposit records') and scopes it to QuickBooks Online. It is unambiguous about being a search tool, but it does not contrast with get_deposit or mention that it returns a page of matching records, so it doesn't fully distinguish from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus get_deposit or other search_* tools, and no alternatives or exclusions are named. The only usage hints live inside the criteria parameter schema, not in the tool description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_employeessearch employeesB
Read-only
Inspect

Search employee records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all employee records.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses no behavioral traits beyond the annotations. Annotations already provide readOnlyHint=true and destructiveHint=false, so safety is covered, but the description adds no context about pagination, truncation, case-sensitivity, inactive-record handling, or unavailable fields. Those details are in the schema's criteria description, not the tool description. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence with no filler. It is front-loaded and states the core function without unnecessary words. It is appropriately sized for a tool whose complexity is largely delegated to the input schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is thin, but the input schema's criteria description provides extensive contextual detail: filtering operators, pagination, truncation, case-sensitivity, inactive-record behavior, and unavailable fields. Annotations cover safety. The only notable gap is the unspecified return shape of employee records, which is minor for a search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool description itself adds no parameter meaning beyond the schema. Schema coverage is 100%, and the criteria parameter has an exceptionally detailed description covering operators, pagination, and limitations. Baseline 3 is appropriate because the schema carries the semantic load; the description offers no additional clarification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('search') and resource ('employee records') and identifies the platform (QuickBooks Online). It is clear but does not explicitly differentiate from sibling tools like get_employee or search_customers; the resource is evident from the name and description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this tool to search employee records. It does not explicitly state when to choose this over get_employee (single record) or other search_* tools, nor does it mention alternative routing. The schema's criteria description provides detailed usage guidance, but the description itself lacks when-to-use versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_estimatessearch estimatesB
Read-only
Inspect

Search estimate records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all estimate records.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The top-level description adds no behavioral detail beyond the annotations; it only restates the operation. It does not mention filtering semantics, pagination, truncation, or unavailable sensitive fields, which are documented only in the schema. No contradiction with readOnlyHint=true/destructiveHint=false exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no filler, and the core purpose is stated immediately. The brevity is appropriate given the heavy lifting done by the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Together with the rich criteria parameter documentation, the tool definition covers filters, AND semantics, unsupported fields, case sensitivity, pagination, and default behavior. The absence of an output schema and the terse top description leave the exact return shape implied rather than stated, so it is not a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents both parameters thoroughly (100% coverage), including the full criteria grammar and operators. The description itself adds no parameter-level detail, so it sits at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete operation and object: 'Search estimate records in QuickBooks Online.' It is clear enough to distinguish from create/update/delete_estimate and from the singular get_estimate, although it does not explicitly contrast with sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to choose search_estimates versus get_estimate or other search tools. The only behavioral hint, 'Omit to return the first page of all estimate records,' lives in the schema's criteria parameter, not in the description, and there is no explicit 'use this when...' statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_exchange_ratessearch exchange ratesB
Read-only
Inspect

Search exchange rate records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Filter on SourceCurrencyCode and AsOfDate (YYYY-MM-DD) to read one rate. Omit to return the first page of all exchange_rate records.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description only restates the read/search behavior without adding useful context such as pagination, result truncation, filtering constraints, or return behavior. The detailed behavioral notes live inside the schema, not the top-level description, so the description itself adds little beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler: it names the action, resource, and platform. It is concise, though it is closer to under-specification than to rich but efficient guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read-only search with no required parameters, and the criteria schema provides substantial operational guidance about filtering, pagination, and truncation. However, there is no output schema and no explicit mention of the return shape, so the description alone is not fully complete for an agent trying to predict the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters have complete schema descriptions (100% coverage), including the detailed criteria description covering operators, AND semantics, field limitations, pagination, and truncation. The top-level description adds no parameter-level meaning, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Search'), a specific resource ('exchange rate records'), and the platform ('QuickBooks Online'), so an agent can tell this is a query/list operation. It does not explicitly distinguish itself from the sibling get_exchange_rate, but the verb and noun make the search intent reasonably obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The top-level description offers no explicit guidance about when to prefer this tool over get_exchange_rate or other search tools. Some usage direction is embedded in the criteria parameter description, such as filtering on SourceCurrencyCode and AsOfDate to read one rate, but this is implied usage rather than explicit selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_invoicessearch invoicesA
Read-only
Inspect

Search invoice records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all invoice records.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The criteria parameter description adds meaningful behavioral context beyond that: one call returns one page of at most 1000 records, large pages are truncated, filters are AND-only, field values are case-sensitive, and certain data (government identifiers, birth dates, card-processing details) is unavailable. The main description itself is terse, but the definition as a whole discloses the tool's operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is a tightly scoped single sentence with no filler. The criteria text is a dense block of necessary operational detail, front-loaded with accepted formats before caveats; it could benefit from structural formatting, but every clause carries functional value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, zero-required-parameter read-only search tool with no output schema, the definition covers company selection, criteria construction, pagination, truncation, and data-availability limits. The notable omissions are the response shape and routing to related siblings like get_open_invoices or get_invoice, but neither is required for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the criteria text goes far beyond the bare union type (object/array/string/null). It specifies the three accepted shapes, operators (=, <, >, <=, >=, LIKE with % wildcards, IN), AND semantics, the whitelist of filterable fields, case-sensitivity, and pagination options (limit, offset, count, asc/desc) — everything needed to construct a valid call. The company parameter's optional-when-single rule is also documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search invoice records in QuickBooks Online.' This clearly separates it from resource-level siblings like search_bills or search_customers. However, it does not differentiate from get_invoice or get_open_invoices; the agent must infer that 'search' implies flexible criteria-based lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when/when-not guidance or alternative tools are named in the description; an agent must infer from the name when to pick this over get_invoice or get_open_invoices. The criteria parameter text offers strong how-to guidance (narrow with filters, walk pages with limit/offset), but that is usage-within-tool, not tool-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_itemssearch itemsB
Read-only
Inspect

Search item records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all item records.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The top-level description itself adds no behavioral detail such as pagination, inactive-record handling, or unsupported OR filters — those live in the criteria parameter description rather than the tool description. This is not contradictory, but it is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence naming verb, resource, and system, with no filler. Every word earns its place, and the longer criteria description is justified by the genuinely complex query behavior it documents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Together with the very detailed criteria schema, the definition gives an agent enough to construct queries, paginate, and avoid unsupported OR or Line-field filters. It falls short of 5 because there is no explicit statement of the return shape and no guidance on when to prefer get_item for a single known record.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the criteria parameter description already documents operators, sort/limit/offset, AND semantics, and QuickBooks limitations. The top-level description adds no parameter meaning, so per the high-coverage baseline this is adequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the action (search) and resource (item records) in QuickBooks Online, which immediately distinguishes it from the many create/update/get siblings. It stops short of 5 because it does not explicitly say it returns a list or page, or how it differs from get_item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use search_items versus get_item or the other search_* tools, and no mention of when-not-to-use or alternatives. The name implies a search use case, but the description leaves selection among siblings to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_journal_entriessearch journal entriesB
Read-only
Inspect

Search journal entry records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all journal_entry records.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description only restates the operation and adds no behavioral context beyond what annotations already provide. It does not mention pagination, result size limits, filtering limitations, or any other runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence with no filler and conveys the essential purpose immediately. It is appropriately concise, though it is terse enough to omit useful usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool, the readOnlyHint annotation and the detailed criteria schema description provide most of what an agent needs to invoke the tool correctly. The one-line description alone would be insufficient, but the schema fills the key operational gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline applies even though the tool description itself says nothing about parameters. The criteria parameter's own schema description is rich and covers operators, AND semantics, and pagination.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action and resource: searching journal entry records in QuickBooks Online. It is unambiguous, but it does not distinguish this tool from sibling tools like get_journal_entry or the many other search_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as get_journal_entry, get_journal_report, or search tools for other record types. No exclusions, prerequisites, or preferred use cases are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_payment_methodssearch payment methodsC
Read-only
Inspect

Search payment method records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all payment_method records.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds no behavioral context such as pagination limits, inactive record exclusion, or field-filter restrictions. The schema's criteria description covers these behaviors, but the description itself contributes zero additional transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose with no wasted words. While it is under-specified, that is a completeness issue rather than a conciseness or structure problem.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a complex, flexible criteria parameter and no output schema, yet the description offers only a bare statement of purpose. The extensive schema description partially compensates, but the description itself fails to summarize search behavior or result characteristics, leaving the agent to infer capability from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'company' and 'criteria' fully documented in the input schema. The description adds no extra parameter meaning, but at this high coverage level the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'search' and resource 'payment method records' within QuickBooks Online, clearly separating it from create/update/delete siblings. It does not explicitly contrast with get_payment_method, but the plural 'records' signals a list-oriented search, making the purpose fairly clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like get_payment_method or other search_* tools. The only usage hints live in the criteria parameter's schema description (e.g., filtering, pagination), not in the tool description itself, leaving the agent to discover them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_paymentssearch paymentsA
Read-only
Inspect

Search payment records in QuickBooks Online. An accounting record of a customer payment already received. Payment processing is not supported; ProcessPayment must be omitted or false.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all payment records.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the agent knows this is a read-only operation. The description goes beyond that by disclosing that ProcessPayment must be omitted or false, a specific behavior not captured by the annotation. The schema's criteria description adds further behavioral details (no OR, no Line-field filtering, pagination/truncation), though these come from the schema rather than the narrative description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. The core action is front-loaded, and the second sentence adds a crucial domain clarification and a processing constraint. No repetition of schema data or annotation content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the read-only annotation, two parameters (0 required), and a comprehensive schema description covering filter syntax, limitations, and pagination, the tool is well-documented overall. The narrative description covers core semantics and the processing constraint. A minor gap is that the response shape is not described, but no output schema exists and this is probably acceptable for a search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents both parameters thoroughly. The criteria parameter's description is especially rich, covering operators, filtering limitations, and pagination behavior. The narrative description adds no parameter-specific meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Search payment records in QuickBooks Online', a clear verb+resource statement. It then defines 'payment' specifically as 'an accounting record of a customer payment already received', which differentiates it from sibling tools like search_bill_payments (vendor payments) and search_credit_card_payment_txns. The additional constraint 'Payment processing is not supported' further distinguishes it from payment-processing operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by clarifying that this covers customer payments already received, and warns that payment processing is not supported. However, it names no alternative tools and gives no explicit when-to-use/when-not-to-use guidance. An agent must infer the boundary with search_bill_payments or get_payment from the definition of 'payment' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_purchase_orderssearch purchase ordersB
Read-only
Inspect

Search purchase order records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all purchase_order records.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds no behavioral information beyond the word 'Search'. It does not mention pagination, truncation, filtering limitations, or the fact that one call returns one page—those details live in the parameter schema rather than the tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It states the verb, resource, and system context efficiently, earning its place without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The overall tool definition is quite complete because the criteria parameter description provides detailed behavior on operators, case-sensitivity, privacy constraints, pagination, and truncation. The main description is thin, but the schema richness compensates; only alternative-tool guidance is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully documents both the company and criteria parameters. The tool description itself adds no parameter-level meaning, but per the baseline for high schema coverage, a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Search') and resource ('purchase order records in QuickBooks Online'), so the tool's purpose is immediately clear. It does not explicitly distinguish itself from sibling search tools, but the resource is specific enough that an agent can infer it targets purchase orders rather than purchases or bills.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as get_purchase_order or search_purchases. There are no exclusions, prerequisites, or selection criteria; usage is only implied by the tool's name and resource.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_purchasessearch purchasesC
Read-only
Inspect

Search purchase records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all purchase records.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds no behavioral information such as pagination limits, filter constraints, or data availability restrictions. Those details appear only in the criteria parameter description, not in the tool description itself. Thus, the description contributes little beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, a single sentence with no filler. However, it is under-specified to the point of being almost tautological, providing only a verb and resource. It is not verbose, but it also does not earn its place by adding any informative value beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The overall tool definition is rich because the criteria parameter description thoroughly covers operators, limitations, pagination, and excluded data. The description itself is minimal, but it does not need to repeat those details since the schema handles them. Still, the description lacks any note about the return format or relation to sibling tools, making it only marginally adequate for a complex search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed descriptions for both the company and criteria parameters. The tool description itself adds no parameter information; it merely restates the tool's purpose. Since the schema already carries the full burden of parameter semantics, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the verb 'Search' and the resource 'purchase records', making the purpose unambiguous. However, it does not differentiate itself from the many sibling search tools (e.g., search_bills, search_invoices) beyond the resource name, and it offers no unique identifier or distinguishing capability. Still, it is specific enough to be understood.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. There is no mention of using get_purchase for single records, no filter strategies, and no exclusions. It is a single sentence that gives no context for selection among the many search_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_recurring_transactionssearch recurring transactionsB
Read-only
Inspect

Search recurring transaction records in QuickBooks Online. A template QuickBooks uses to create a transaction on a schedule, or on request.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Each result wraps the template's transaction under its type (Invoice, Bill, ...) with RecurringInfo. Omit to return the first page of all recurring_transaction records.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool readOnly and non-destructive, and the description adds no behavioral details beyond defining what a recurring transaction is. It does not disclose pagination, OR limitations, or filtering constraints; those live only in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the operation front-loaded. The second sentence adds domain context but is not strictly required; still, nothing is wordy or misplaced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The rich criteria schema compensates for the sparse description, covering pagination, filters, operators, and result wrapping. The description itself is thin, but the schema fills the practical gaps for an agent making a call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema carries the parameter semantics. The description adds nothing about the company or criteria parameters beyond what the schema already documents, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Search recurring transaction records in QuickBooks Online') and clarifies the entity as a template for scheduled/on-request transactions. It is clear, though it does not explicitly contrast with the sibling get_recurring_transaction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No description-level guidance explains when to use search vs. get_recurring_transaction or any other alternative. The intended usage must be inferred from the tool name and generic search semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_refund_receiptssearch refund receiptsB
Read-only
Inspect

Search refund receipt records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all refund_receipt records.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the read-only nature. The description itself adds no additional behavioral traits such as pagination, truncation, lack of OR filtering, or case sensitivity—these are only in the schema. The description fails to disclose these behaviors beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, direct sentence with no fluff. It is appropriately concise and front-loads the core purpose without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the schema provides extensive details on criteria and pagination, the description alone is minimal and doesn't hint at the tool's complexity or differentiate it from get_refund_receipt. It doesn't mention that it returns a page of records or how to narrow results, though the schema compensates. For a search tool with such a complex criteria parameter, a bit more context in the description would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for both parameters (company and criteria), with the criteria description being highly detailed. The tool description adds no extra meaning to the parameters, but since the schema already covers them, the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search) and the target resource (refund receipt records in QuickBooks Online). It distinguishes from sibling get_refund_receipt (which fetches a single record) and the create/update/delete tools by implying a listing/search operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like get_refund_receipt or other search_* tools. It doesn't mention that this is for multiple records or how to filter, though the schema does. There is no explicit when-to-use or when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_reimburse_chargessearch reimburse chargesC
Read-only
Inspect

Search reimburse charge records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all reimburse_charge records.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds no behavioral detail beyond 'search', such as pagination, truncation, filtering limitations, or the fact that it returns one page at a time. Since it contributes no extra transparency, a low score is warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one concise sentence with no filler, making it easy to scan. However, its brevity is partly a result of underspecification rather than deliberate economy, so it does not earn a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and only a one-sentence description, the agent is not told what the search returns or what a reimburse charge record looks like. The schema's criteria notes are comprehensive, but the description itself adds almost nothing, leaving the definition incomplete for an unfamiliar agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the criteria parameter's description is extensive (operators, ANDed filters, field restrictions, paging). The tool description itself mentions no parameters; if there is meaning here, it comes from the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action and resource: it searches reimburse charge records in QuickBooks Online. It is understandable on its own, though it adds little beyond the tool's name and does not explicitly contrast with sibling search tools, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool instead of other search tools, such as search_bills or search_purchases, nor about prerequisites like connected companies. The only usage context is embedded in the schema's criteria description, not in the tool description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_sales_receiptssearch sales receiptsB
Read-only
Inspect

Search sales receipt records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all sales_receipt records.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and destructiveHint=false, so the one-line search description does not need to repeat a safety warning. It adds no behavioral detail beyond 'Search' – pagination, truncation, lack of OR, and field restrictions appear in the criteria parameter's schema text rather than in the tool description, leaving a modest gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler or redundant phrasing. It is efficient, though its brevity means it carries no supporting structure such as scope examples or exclusions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Together with the rich criteria schema and read-only annotations, the definition gives an agent enough to invoke the tool correctly, including filter behavior and page-size limits. It stops short of 5 because there is no output schema and the tool description does not describe result shape or default sorting beyond what the criteria text says.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the criteria parameter is extensively documented with operators, AND semantics, limitations, and paging. The tool description contributes no additional parameter meaning, so the high-coverage baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource, 'Search sales receipt records in QuickBooks Online,' which is clear and not a tautology. It does not explicitly distinguish itself from sibling search_* tools or from get_sales_receipt, so it misses the top-level differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool description gives no guidance on when to choose this over alternatives such as get_sales_receipt or search_refund_receipts, and states no exclusions or prerequisites. The criteria parameter documents filtering mechanics but does not address tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_tax_agenciessearch tax agenciesC
Read-only
Inspect

Search tax agency records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all tax_agency records.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is essentially a restatement of the title and adds no behavioral detail beyond what the readOnlyHint=true and destructiveHint=false annotations already convey. Pagination, truncation, filtering limitations, and field-level restrictions appear only in the parameter schema, not in the tool description itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler, and it names the verb and resource immediately. It loses a point because it closely echoes the title without adding much substantive information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The parameter schema is rich enough to support correct invocation, covering operators, pagination, and unsupported fields. However, the definition as a whole omits selection guidance versus get_tax_agency and does not describe the shape of returned records, though the absence of an output schema makes this less critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with thorough documentation of both the 'company' parameter and the highly flexible 'criteria' parameter. The one-line tool description adds no parameter-level meaning, but the schema already carries the full burden, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Search') and names a concrete resource ('tax agency records in QuickBooks Online'), so an agent can tell this is a read-only search over tax agencies. It does not explicitly differentiate from sibling tools like get_tax_agency, search_tax_codes, or search_tax_rates, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool instead of an alternative such as get_tax_agency for a single record or search_tax_rates for tax rates. The tool name implies a search use case, but no when-to-use condition, exclusion, or alternative is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_tax_codessearch tax codesA
Read-only
Inspect

Search tax code records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all tax_code records.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The criteria parameter description adds substantial behavioral detail beyond the annotations: filters are ANDed with no OR support, only top-level QuickBooks fields can be filtered, field names and values are case-sensitive, sensitive identifier fields are unavailable, one page returns at most 1000 records, and large pages are truncated. These runtime constraints go far beyond what readOnlyHint and destructiveHint provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The top-level description is one short, filler-free sentence. The long criteria text is dense but purposeful, front-loading accepted query shapes and then operational constraints. Each clause adds decision-relevant information with no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool, the definition is nearly complete: it explains company selection, query syntax, operators, unsupported filters, privacy limitations, pagination/truncation behavior, and default results. It does not list tax-code-specific filterable fields or describe the return shape, but no output schema exists and the generic field list makes invocation safe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the input schema. The tool description itself adds no parameter-level meaning beyond that, and the rich criteria syntax lives in the schema rather than in the description. The baseline 3 applies because the schema already does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Search') and resource ('tax code records') in QuickBooks Online, so an agent knows what the tool operates on. It does not explicitly distinguish itself from sibling tools like search_tax_agencies or get_tax_code, so it falls just short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use statement, no mention of alternatives, and no exclusions. The verb 'Search' plus 'tax code records' implies the intended use case, but the agent must infer it from tool naming rather than from stated guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_tax_paymentssearch tax paymentsA
Read-only
Inspect

Search tax payment records in QuickBooks Online. Sales tax (GST/HST, QST, VAT) payments and refunds recorded against filed returns; AU, CA and UK companies only.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all tax_payment records.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds the geographic scope and the fact that it includes refunds, but does not disclose pagination or filtering limitations beyond what the schema already covers. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler, front-loaded with the core purpose and then the scope constraint. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite being brief, the description, combined with the highly detailed schema for the criteria parameter and the annotations, gives an agent sufficient context to call the tool correctly. The output format is not described, but there is no output schema and the search behavior is adequately covered by the criteria documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema thoroughly documents both 'company' and 'criteria' (including operators, filtering rules, pagination, and truncation). The description adds the geographic restriction but does not add extra meaning beyond what the schema already provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Search') and resource ('tax payment records in QuickBooks Online'), and adds distinguishing context: sales tax payments/refunds and the AU/CA/UK geographic scope. This clearly separates it from other search tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for sales tax payment records but does not explicitly compare to alternatives like get_tax_payment or search_tax_agencies. The geographic restriction is a useful usage constraint, but no when-to-use or when-not-to-use guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_tax_ratessearch tax ratesB
Read-only
Inspect

Search tax rate records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all tax_rate records.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is clear. The description itself adds no behavioral context such as pagination, truncation, case sensitivity, or record limits; these details reside in the schema's criteria parameter description, which is structured data and not credited here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that states the tool's purpose without redundancy. It front-loads the action and resource, leaving detailed parameter semantics to the schema. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema documentation for the criteria parameter and the presence of annotations, the minimal description is adequate but not complete. It does not explicitly state that the tool returns a pageable list of tax rate records, though this is implied by the name and elaborated in the schema. With no output schema, some return-format clarity from the description would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'company' and 'criteria' are already well documented in the schema. The description adds no additional parameter meaning. Per the baseline rule for high schema coverage, a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Search') and resource ('tax rate records') within the QuickBooks Online context. It distinguishes itself from siblings like search_tax_agencies and search_tax_codes by naming the specific record type, though it does not explicitly contrast with get_tax_rate or other search tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives. It does not mention when to prefer search_tax_rates over get_tax_rate, nor does it note any exclusions or prerequisites. The detailed criteria description is operational, not tool-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_termssearch termsB
Read-only
Inspect

Search term records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all term records.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile, so a terse description is acceptable. Rich behavioral facts (1000-record page cap, page truncation, inactive records excluded by default, no OR support) are disclosed in the criteria parameter schema rather than in the description itself, which adds transparency but not from the description's own text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; every word in 'Search term records in QuickBooks Online' earns its place. No redundancy with title or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the tool is operationally well documented because the criteria parameter covers filtering, pagination, truncation, inactive-record defaults, and unavailable data types. The notable gap is tool-selection guidance (search_terms vs get_term), but an agent can invoke this tool correctly from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the criteria description is unusually thorough (operator syntax, field restrictions, case sensitivity, pagination, truncation). The tool description itself adds no parameter meaning, so it correctly sits at the baseline with the schema doing the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('search') and resource ('term records') within QuickBooks Online, so an agent can tell what entity it operates on. Among siblings there are both search_terms and get_term, and the description does not explicitly separate them, but 'Search term records' is unambiguous about scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to choose this over get_term or the many other search_* tools. The criteria parameter documents query mechanics, but nothing tells the agent when search is the appropriate tool versus fetching a single term.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_time_activitiessearch time activitiesB
Read-only
Inspect

Search time activity records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all time_activity records.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and destructiveHint=false, so the safety profile is covered without the description restating it. The criteria parameter documentation goes beyond the annotation by disclosing one-page/1000-record behavior, truncation, and lack of OR support, but the one-line description itself adds no behavioral nuance. Since annotations carry the safety burden and the schema adds needed behavior details, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. It lets the detailed criteria parameter description carry complexity instead of duplicating it, making the definition appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the read-only annotation and the very detailed criteria schema, the definition is complete enough for an agent to invoke it correctly: criteria behavior, page size, truncation, and filter limitations are all documented. The only meaningful gap is the absence of explicit routing guidance between this and related get/search tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the 'criteria' parameter extensively documented: accepted shapes, operators, field restrictions, case sensitivity, and pagination semantics. The description text adds no additional parameter meaning beyond what the schema already provides, so it sits at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Search') and resource ('time activity records') scoped to QuickBooks Online, so an agent can tell what the tool does. It does not explicitly contrast itself with get_time_activity or other search tools, though the 'search' verb implies a listing/filtering behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over get_time_activity, when not to use it, or which tool should be chosen in alternative situations. The schema's criteria documentation implies operational usage and pagination strategy, but the description itself provides no when-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_transferssearch transfersA
Read-only
Inspect

Search transfer records in QuickBooks Online. An accounting entry recording a transfer between accounts. It does not move money between bank accounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all transfer records.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this tool as read-only and non-destructive. The description adds useful behavioral context: transfer records are accounting entries and this tool does not move money between bank accounts, preventing a common misinterpretation. It doesn't cover auth, rate limits, or side effects, but those are less critical for a read-only search.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the main verb and resource front-loaded. The second sentence earns its place by clarifying the domain meaning of 'transfer', and there is no redundant restatement of the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The full definition provides company selection, flexible criteria, filtering constraints, pagination behavior, and read-only annotations. The only gap is the absence of an explicit output schema or return-field description, but for a search-over-records tool the description and schema are otherwise sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema's criteria description is rich—operators, ANDed filters, case sensitivity, pagination, and unsupported fields. The tool description itself contributes no parameter-level meaning, so it sits at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names the action and resource explicitly ('Search transfer records in QuickBooks Online'), and the second sentence clarifies what a transfer is in this domain. It is clear, but it doesn't contrast with get_transfer or the other search_* siblings, so differentiation relies mostly on the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the search use case but never says when to choose search_transfers over get_transfer or create_transfer. The schema's criteria text explains filtering behavior, but selection guidance among sibling tools is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_vendor_creditssearch vendor creditsA
Read-only
Inspect

Search vendor credit records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Omit to return the first page of all vendor_credit records.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish this as read-only and non-destructive, and the top-level description adds no behavioral trait beyond 'search'. It is consistent with the annotations, but pagination, truncation, filtering limits, and unavailable fields live in the schema rather than the description, so the description itself is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The tool description is a single, front-loaded sentence with no wasted words. The long criteria explanation is correctly placed in the schema instead of cluttering the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Taken as a full definition, the clear one-liner plus the detailed criteria spec and read-only annotations is nearly complete for a search tool. The only notable gap is the absence of an output schema or explicit return-field description, so the agent must infer the shape of returned vendor credit records.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents both the company and criteria parameters in exhaustive detail. The one-line tool description adds no parameter-level meaning, so it stays at the schema-covered baseline rather than being elevated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Search') and an unambiguous resource ('vendor credit records') scoped to QuickBooks Online. This lets an agent distinguish it from the singular get_vendor_credit and from search_vendors, which target different records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to prefer this tool over get_vendor_credit or other search_* siblings; the only cue is the verb and resource in the description. The criteria parameter documentation does provide operational advice on narrowing filters and paging, but tool-selection exclusions and alternatives are left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_vendorssearch vendorsB
Read-only
Inspect

Search vendor records in QuickBooks Online.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
criteriaNoCriteria accepts: an array of {field, value, operator} objects (operators: =, <, >, <=, >=, LIKE with % wildcards, IN), a simple {Field: value} object, or {criteria: [...], asc|desc: "Field", limit, offset, count}. Filters are ANDed; QuickBooks has no OR, cannot filter on Line fields, and only filters top-level fields (TxnDate, DocNumber, CustomerRef, VendorRef, Balance, TotalAmt, DueDate, MetaData.LastUpdatedTime, ...). Field names and values are case-sensitive as QuickBooks stores them. Government identifiers, birth dates and card-processing details are not available through these tools. One call returns one page of at most 1000 records, and a large page is truncated to fit the answer: narrow with filters or walk it with limit and offset rather than asking for everything. Inactive records are left out unless Active is filtered explicitly (Active: false, or Active IN (true, false) for both). Omit to return the first page of all vendor records.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and destructiveHint=false, and 'Search' is consistent with that. However, the description itself adds no behavioral context beyond the annotations: it does not mention pagination, truncation, inactive-vendor exclusion, or query limitations. Those behaviors are present only inside the parameter schema, not in the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, no repetition of the tool name, and the action plus resource are front-loaded. It is appropriately compact for a read-only search tool whose parameter complexities are handled by the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The overall definition is usable because the schema documents criteria grammar, pagination guidance, and company disambiguation, while annotations confirm a safe read operation. However, there is no output schema and the description never states the response shape or that output is a page of matching vendor records. This leaves an agent to infer what the tool returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both company and criteria parameters documented; the criteria description is especially detailed about operators, AND semantics, and QBO limitations. The top-level description adds no parameter-level meaning, so the baseline score of 3 for high schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search vendor records in QuickBooks Online.' This clearly tells an agent what the tool does and distinguishes it from sibling tools that search other QBO entities such as search_customers or search_invoices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance about when to use search_vendors versus get_vendor, get_vendor_balance, search_vendor_credits, or other related tools. It does not state conditions, exclusions, or mention that criteria-based searching is the differentiator. Useful usage context lives only in the schema's criteria description, not in the tool description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_credit_memosend credit memoA
Destructive
Inspect

Email a credit memo to the customer as a PDF from QuickBooks Online, and mark it sent. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the CreditMemo to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=true), the description reveals critical behaviors: sending is asynchronous, delivery requires user approval in Caribooks, the memo cannot be unsent, and the record must still exist at delivery time. This is exactly the kind of behavioral context an agent needs and goes well beyond the structured annotation hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the primary action, the second explains asynchronous delivery and record-liveness constraints, and the third discloses irreversibility and the required user approval step. Critical warnings are front-loaded and the language is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a send action with no output schema, the description covers everything an agent needs to know: what happens, when it happens, what can prevent it, and what the user must do. Combined with a fully described schema, this is a complete operational picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all four parameters, so the schema already carries the semantic weight. The description does not add parameter-level detail, but it does reinforce the need for explicit approval, which complements the 'confirm' parameter's 'must be true' requirement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Email'), a specific resource ('a credit memo ... as a PDF from QuickBooks Online'), and a clear outcome ('mark it sent'). This clearly distinguishes it from sibling send_* tools that target other document types like invoices, estimates, or sales receipts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear that this tool is for emailing credit memos and that it is subject to an approval workflow, which is essential context for deciding when to invoke it. It does not explicitly name sibling alternatives or state conditions to avoid using it, but the resource-specific purpose is sufficiently clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_estimatesend estimateA
Destructive
Inspect

Email an estimate to the customer as a PDF from QuickBooks Online, and mark it sent. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Estimate to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false, openWorldHint=true, and destructiveHint=true. The description adds crucial behavior beyond these: asynchronous delivery, 'mark it sent', record-must-exist caveat, deletion/voiding risk, and approval delay. It thoroughly explains what happens when the tool is invoked.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries essential information: the core action, async behavior, delivery prerequisites, and the user-approval requirement. The structure front-loads the main purpose and then layers critical caveats without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating send operation with no output schema, the description covers the essential lifecycle, approval flow, and failure conditions. It does not describe return values, but the most important operational facts are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents id, email, company, and confirm. The description does not add parameter-specific detail beyond the schema, but no compensation is needed because the schema is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Email an estimate to the customer as a PDF from QuickBooks Online, and mark it sent.' This clearly differentiates it from sibling send_* tools by naming the exact record type (estimate) and the system (QuickBooks Online).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: sending is async, requires user approval in Caribooks, and cannot be undone. It tells the agent to inform the user of the approval step. It does not explicitly name alternative tools like send_invoice, but the resource is unambiguous enough for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_invoicesend invoiceA
Destructive
Inspect

Email an invoice to the customer as a PDF from QuickBooks Online, and mark it sent. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Invoice to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructive and non-read-only behavior, so the bar for adding context is lower, but the description goes well beyond annotations by disclosing asynchronous sending, irreversibility, approval dependency, and failure conditions such as deletion or voiding preventing delivery. This is highly transparent about real-world side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then gives essential caveats in a sensible order. Every sentence provides necessary behavioral information, including the exact approval URL, without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating send operation with no output schema, the description covers all critical invocation-time knowledge: the action, async delivery, irreversibility, user approval, and deletion/voiding risks. An agent has enough context to invoke this correctly and inform the user appropriately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all four parameters, so the schema carries the parameter documentation burden. The description adds general workflow context but does not add new meaning to individual parameters beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: email an invoice as a PDF from QuickBooks Online and mark it sent. This clearly distinguishes it from sibling send_estimate, send_purchase_order, and other send_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: the email is sent asynchronously, requires user approval in Caribooks, cannot be unsent, and depends on the record still existing at delivery time. It does not explicitly name alternatives, but the scope is clear enough for an agent to know when to invoke this tool versus other send_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_paymentsend paymentA
Destructive
Inspect

Email the customer a receipt for a payment received, as a PDF from QuickBooks Online. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Payment to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by revealing asynchronous delivery, the need for the record to still exist at delivery, deletion/voiding as failure causes, the fact the email cannot be unsent, and the Caribooks approval requirement. This gives the agent crucial behavioral expectations that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficient and every sentence contributes: the action, the asynchronous timing, the record-existence caveat, and the irreversible approval step with a direct URL. The most important information is front-loaded at the start.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for an irreversible, side-effectful, externally-delivering tool. It covers what happens after calling, why approval is required, what can prevent delivery, and what the agent should tell the user. With no output schema, no additional return-value explanation is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents id, email, company, and confirm. The description adds useful context around delivery and approval but does not meaningfully elaborate on the individual parameters beyond what the schema says, matching the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and resource: 'Email the customer a receipt for a payment received, as a PDF from QuickBooks Online.' This clearly distinguishes it from sibling send_* tools by focusing on payment receipts, not invoices, credit memos, or sales receipts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The first sentence implies the tool's intended use case, and the approval/confirmation context explains when the email actually goes out. However, it does not explicitly compare against alternatives or state when this tool should or should not be chosen over sibling send_* tools, so the guidance is mostly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_purchase_ordersend purchase orderA
Destructive
Inspect

Email a purchase order to the vendor as a PDF from QuickBooks Online. The email parameter is required: QuickBooks answers 'Email Address is required to send email' even when the purchase order carries a POEmail. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the PurchaseOrder to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The destructiveHint annotation is already present, but the description goes well beyond it by disclosing that the email cannot be unsent, delivery is asynchronous, the record must still exist, and user approval is required in Caribooks. It also explains a surprising QuickBooks quirk about the email parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense, purposeful sentences with the core action front-loaded. Every additional clause carries a necessary caveat or instruction, and there is no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating asynchronous operation with no output schema, this is remarkably complete: it covers required email behavior, approval dependency, delivery timing, irreversibility, and a failure mode. An agent knows what to do, what to warn the user about, and what can go wrong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds important meaning for the `email` parameter: it is effectively required even though the schema marks it optional with a default, and it explains why. This genuinely helps an agent avoid a known failure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific action ('Email a purchase order to the vendor as a PDF') and names the exact resource, making it immediately distinguishable from send_invoice, send_estimate, and other send_* siblings. The delivery mechanism and document type are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when this tool is appropriate: emailing a purchase order. It also tells the agent what to convey to the user about the approval flow. It does not explicitly name alternatives or exclusion conditions, but no other sibling serves this exact purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_refund_receiptsend refund receiptA
Destructive
Inspect

Email a refund receipt to the customer as a PDF from QuickBooks Online, and mark it sent. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the RefundReceipt to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavioral context beyond the annotations: asynchronous sending after approval, irreversible nature, dependency on record existence, and the explicit instruction to tell the user. These are critical for safe invocation and are not derivable from readOnlyHint/destructiveHint alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with the core action first, followed by important caveats in order of relevance. Some minor redundancy between 'Cannot be unsent' and the approval explanation, but overall every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the major behavioral constraints: approval requirement, asynchronous email, non-reversibility, and record existence. Includes a link to the approvals page. No output schema exists, but for a fire-and-forget email action, return value details are not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the description doesn't add much parameter-specific meaning. It mentions 'to the customer' but doesn't elaborate on the email or company selection parameters. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (email), resource (refund receipt), format (PDF), and system (QuickBooks Online). This clearly distinguishes it from sibling send_* tools, which target different document types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains that the email requires user approval in Caribooks, cannot be unsent, and that the record must exist at delivery. This gives clear preconditions and warns the agent to inform the user. It doesn't explicitly name alternatives, but the resource-specific name makes the intended use unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_sales_receiptsend sales receiptA
Destructive
Inspect

Email a sales receipt to the customer as a PDF from QuickBooks Online, and mark it sent. QuickBooks sends the email asynchronously a few seconds after it is approved. The record must still exist when delivery occurs; deletion or voiding can prevent delivery. Cannot be unsent, so the email waits for the user's approval in Caribooks (Approvals page, https://caribooks.com/portal/review) and goes out only once they approve it there; tell them so.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the SalesReceipt to email.
emailNoRecipient address. Defaults to the address on the QuickBooks record.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed sending this email.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond the annotations: async delivery, approval requirement in Caribooks, inability to unsend, and the need for the record to exist at delivery time. Annotations only indicate a destructive write (readOnlyHint=false, destructiveHint=true), so this additional detail is valuable and not redundant. It does not fully explain the exact status change of 'mark it sent', but the coverage is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately structured: the first sentence gives the core purpose, followed by necessary caveats in a logical order. It is slightly longer than strictly needed, but every sentence covers an important behavioral constraint (async, record existence, irrevocability, approval), so the length is justified. The inclusion of a URL for the approval page is a practical addition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and a simple parameter set, the description covers the essential behavioral context an agent needs: how to handle user confirmation, the async nature, potential delivery failure, and the requirement to inform the user about approval. It does not mention error responses or success indications, but the absence of an output schema and the simplicity of the action make the description sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add new meaning to the parameters themselves; it reinforces the confirm parameter indirectly by stating the email waits for approval, but this is not a parameter-level detail. It neither harms nor significantly enhances the schema's parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific action ('Email a sales receipt to the customer as a PDF from QuickBooks Online, and mark it sent') with a clear verb and resource. It also distinguishes itself from sibling 'send_' tools by explicitly naming the sales receipt, so an agent can differentiate it from send_invoice or send_estimate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through its clear action statement, but it does not explicitly mention alternatives or conditions for when not to use this tool (e.g., 'use send_invoice instead for invoices'). The sibling tools are similarly named, so the agent must infer the appropriate one based on the document type, which is only implicitly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_accountupdate accountA
Destructive
Inspect

Update an existing account in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Account first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat Name and AccountType, unchanged, or QuickBooks rejects it as missing. QuickBooks cannot delete a account; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag this as destructive, so the description doesn't need to repeat that. It adds a small amount of context by emphasizing the account must already exist and that Id/SyncToken are required, but it does not disclose side effects such as sparse vs full update behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It conveys the essential action and the most important requirement immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the description itself is brief, the input schema's entity description is very rich and covers sparse updates, full updates, restrictions, and the Active:false retirement pattern. Together with annotations, the definition provides enough information for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema carries the burden of explaining entity and company parameters. The description only repeats the Id/SyncToken requirement already documented in the schema, adding no new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Update'), a specific resource ('existing account in QuickBooks Online'), and a key precondition ('requires Id and SyncToken'). This clearly distinguishes it from create_account and the many other update_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies an update workflow by requiring Id and SyncToken, but it does not explicitly say when to use this tool versus create/delete/search alternatives. The richer guidance lives in the input schema rather than in the description text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_attachableupdate attachableA
Destructive
Inspect

Update an existing attachable in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Attachable first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and readOnlyHint=false, so the mutation/destructive nature is covered. The description adds the SyncToken requirement, which hints at stale-token/optimistic-locking behavior, but it does not disclose the sparse-update semantics or full-update clearing behavior itself; those live in the parameter schema rather than the tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler: subject, action, resource, and the essential precondition all appear in order. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Together with the rich entity schema, the tool gives an agent enough to invoke it correctly: prerequisite fields, sparse behavior, line-replacement caveat, sensitive-field restrictions, and company disambiguation. The only notable omission is an explicit statement of the response/return value, but with no output schema and a straightforward update operation that gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter description already details Id/SyncToken requirements, sparse updates, Line-array replacement, and forbidden sensitive fields. The top-level description's 'requires Id and SyncToken' restates that schema content, adding no new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Update an existing attachable in QuickBooks Online') and adds a key prerequisite, so an agent knows what the tool acts on. It implies the attachable must already exist, which loosely separates it from create_attachable and delete_attachable, though it does not name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear precondition: the attachable must already exist and the call requires Id and SyncToken. However, it never states when to prefer create_attachable/delete_attachable, what happens with a stale SyncToken, or any exclusion, so an agent must infer the boundary between update and create/delete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_billupdate billA
Destructive
Inspect

Update an existing bill in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Bill first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat VendorRef, unchanged, or QuickBooks rejects it as missing.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

While annotations already mark destructiveHint=true, the description goes well beyond that: it explains sparse updates by default, that a Line array replaces all lines, that sparse:false clears omitted writable fields, that VendorRef must be repeated, and that sensitive fields must not be sent. This is strong behavioral disclosure and does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is a single, front-loaded sentence stating the action, target, and prerequisite with no filler. Detailed behavioral nuances are delegated to the structured parameter schema, keeping the description appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The combination of annotations and the rich entity description covers preconditions, sparse/full update behavior, destructive semantics, and prohibited fields. A minor gap is that the main description does not mention the expected response or explicitly direct users to create_bill for new bills, but this is not essential given the sibling context and open object schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both entity and company parameters. The tool description only restates that Id and SyncToken are required, which adds no meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Update'), an explicit resource ('an existing bill in QuickBooks Online'), and the key precondition (requires Id and SyncToken). This clearly distinguishes it from create_bill, delete_bill, and get_bill among the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing bill' and the requirement to supply the current SyncToken (fetch the Bill first) make it clear this tool is for updating an already-created bill. It does not explicitly name create_bill as the alternative for new bills, but the context leaves little ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_bill_paymentupdate bill paymentA
Destructive
Inspect

Update an existing bill payment in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the BillPayment first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds that this is an 'update' of an 'existing' record, narrowing the destructive action to modification rather than deletion. However, it does not disclose the sparse vs. full update semantics or that a full update clears omitted fields; those critical behaviors are only in the schema, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the verb, resource, and prerequisite without any fluff. It is appropriately sized for the tool's complexity and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The overall definition is quite complete: annotations cover the destructive nature, the schema covers sparse/full update behavior, sensitive data restrictions, and company selection. The main description is lean but sufficient given the rich schema. The only notable gap is that no output/return behavior is mentioned, and there is no output schema, so the agent is left to infer what the tool returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters have detailed descriptions in the schema. The main description only mentions 'requires Id and SyncToken', which is already stated in the entity parameter description. Thus the description adds no meaning beyond what the input schema already provides, keeping it at the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('update'), the resource ('bill payment'), and the context ('QuickBooks Online'), and the requirement of Id/SyncToken makes it unambiguous that this targets an existing record. It is easily distinguished from sibling tools like create_bill_payment, delete_bill_payment, and void_bill_payment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description and schema provide clear context: it updates an existing bill payment and requires Id and SyncToken, with the schema explicitly instructing to fetch the BillPayment first. However, it does not explicitly name alternatives or state when-not to use this tool (e.g., for new payments use create_bill_payment, for voiding use void_bill_payment), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_budgetupdate budgetA
Destructive
Inspect

Update an existing budget in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Budget first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No contradiction with annotations: readOnlyHint=false and destructiveHint=true align with an update operation. The description adds the prerequisite context that a current SyncToken is required, but it doesn't describe overwrite or clear semantics; the schema compensates and the destructive annotation already signals mutation. With annotations covering the core safety profile, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tightly scoped sentence that front-loads the action and resource and includes the essential precondition. No filler, with details correctly left to the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Together with the input schema's detailed entity semantics and the destructive annotation, the description provides a clear, callable picture: what the tool does, what must be passed, and the risk profile. No output schema exists, but return-value documentation is not required here. It could name the full workflow explicitly, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema's entity description already documents Id, SyncToken, sparse update behavior, and sensitive-field exclusions. The tool description only echoes the Id/SyncToken requirement and adds no new parameter meaning. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Update'), the resource ('existing budget'), and the platform ('QuickBooks Online'), while the requirement for Id and SyncToken clarifies scope. This clearly distinguishes it from sibling create_budget and delete_budget operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing budget' implies the tool is for modifying an already-created record, not for creating or deleting one. The prerequisite of Id and SyncToken and the schema's instruction to fetch the Budget first give workflow guidance, but the description does not explicitly name alternatives or when-not-to-use conditions. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_classupdate classA
Destructive
Inspect

Update an existing class in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Class first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a class; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false. The description's 'Update' is redundant with those annotations and adds no behavioral context such as sparse update behavior, inability to delete, or effects on omitted fields. The schema covers those, but the description itself provides minimal added value beyond the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the core action and prerequisite without fluff or redundancy. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is terse, but the overall tool definition is rich: annotations cover destructive behavior and the schema explains all parameter nuances. The agent can correctly invoke the tool by inspecting the schema. The only gap is that the sparse/full update distinction is not in the description, but it is documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed parameter semantics including required Id/SyncToken, prohibited fields, sparse updates, and retirement via Active:false. The tool description only restates the Id/SyncToken requirement, adding no new meaning beyond the schema. Baseline 3 for high schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the action ('Update'), the resource ('class'), and the system ('QuickBooks Online'), and adds the key prerequisite of Id and SyncToken. This distinguishes it from create_class and other update_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for existing classes ('Update an existing class') but does not explicitly mention when to use it versus create_class, nor that the class must be fetched first (that detail lives in the schema). Usage guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_company_currencyupdate company currencyA
Destructive
Inspect

Update an existing company currency in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the CompanyCurrency first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, covering the mutating nature. The description adds the prerequisite of Id and SyncToken, which hints at optimistic concurrency, but it doesn't elaborate on consequences (e.g., permanent changes, overwriting behavior). Since annotations carry the safety profile, the description adds minimal behavioral context beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the core purpose and includes the critical precondition (Id and SyncToken) in a concise parenthetical. Every word earns its place, and the structure is optimal for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema that thoroughly documents the entity parameter (sparse updates, prohibited fields, etc.), the tool description does not need to repeat those details. It covers the essential prerequisite (fetching the currency to get Id/SyncToken) and implicitly signals the update workflow. A note about using create/delete for different operations would improve completeness, but the schema mitigates the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage with detailed descriptions for both 'entity' and 'company'. The tool description mentions 'requires Id and SyncToken', which is already included in the entity parameter's description. No new parameter-level information is added beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Update'), the resource ('existing company currency'), and the target system ('QuickBooks Online'). It distinguishes from sibling tools like create_company_currency or delete_company_currency by explicitly saying 'existing'. The requirement of Id and SyncToken adds specificity, making the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for updating existing records, but it does not explicitly mention when to use it over alternatives (e.g., create_company_currency for new currencies, delete_company_currency for removal). It also doesn't state that fetching the currency first is required to obtain SyncToken, though the schema implies it. No direct guidance on alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_company_infoupdate company infoA
Destructive
Inspect

Update an existing company info in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the CompanyInfo first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, but the description adds substantial behavior beyond that: updates are sparse by default, sending a Line array replaces all lines, sparse:false clears every omitted writable field, and sensitive fields like SSNs and birth dates must be managed directly in QuickBooks. This is exactly the kind of destructive/behavioral detail an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single tight sentence that fronts the action, resource, system, and the key precondition, with zero filler. The more detailed behavioral and parameter information is appropriately delegated to the schema rather than bloating the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive update with a nested object parameter and no output schema, the definition is nearly complete: it covers preconditions, sparse vs. full update behavior, line replacement, sensitive-field exclusions, and optional company selection. The remaining gap is that it never names the companion read tool for fetching CompanyInfo or states what the update call returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The entity parameter is richly documented in the schema, including Id/SyncToken requirements and sparse update semantics, and the company parameter is also explained. The tool description only restates the Id/SyncToken requirement already present in the schema, adding no new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update'), a specific resource ('existing company info'), and the target system ('QuickBooks Online'). The word 'existing' plus the Id/SyncToken requirement clearly distinguishes update from creation, but it does not explicitly differentiate from sibling read tools like get_company_info or search_company_infos.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description and schema make the prerequisite clear: fetch the CompanyInfo first and supply Id plus current SyncToken. It also tells the agent which fields must not be sent. However, it never names alternatives or states when-not-to-use conditions, so routing among the many company-related sibling tools is inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_credit_card_payment_txnupdate credit card payment txnA
Destructive
Inspect

Update an existing credit card payment txn in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the CreditCardPaymentTxn first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat Amount, BankAccountRef and CreditCardAccountRef, unchanged, or QuickBooks rejects it as missing.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this is a mutating and potentially destructive operation. The description adds that it targets an 'existing' txn and requires Id/SyncToken, but does not detail the sparse/full update behavior or sensitive-field restrictions (those live in the schema description, not the tool description).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that immediately states the action, target, and a critical prerequisite. No redundant wording, and it avoids repeating information already present in the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema description and the annotations (readOnlyHint=false, destructiveHint=true), the brief description suffices as a high-level summary. It tells the agent what the tool does and the key requirement, while the schema covers all invocation details; no output schema exists, so return-value expectations are not required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the entity parameter thoroughly documented (required Id/SyncToken, sparse update semantics, prohibited fields). The tool description itself adds no parameter detail beyond the name and the Id/SyncToken mention, so the schema carries the semantic load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Update'), the specific resource ('existing credit card payment txn'), and a key prerequisite ('requires Id and SyncToken'), which distinguishes it from create_credit_card_payment_txn. It does not explicitly name sibling alternatives, but the verb and resource are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for modifying an already-existing txn by requiring Id and SyncToken, but it does not explicitly state when to choose this over create/delete/search variants. The schema adds prerequisites (fetch first), but the description itself offers no when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_credit_memoupdate credit memoA
Destructive
Inspect

Update an existing credit memo in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the CreditMemo first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry readOnlyHint=false and destructiveHint=true, so the description does not need to restate that this mutates data. The description adds only the Id/SyncToken prerequisite and no side-effect detail such as reversibility or what a full update clears; those details live in the parameter schema rather than the tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence front-loads the verb and resource and tucks the essential prerequisite into a parenthetical. There is no filler, redundancy, or unnecessary background.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Between the rich entity schema, the company-field description, and the destructive/readOnly annotations, an agent has enough information to invoke this tool correctly even without an output schema. The main missing piece is explicit sibling routing, which is already a usage-guideline gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity description thoroughly explains required fields, sparse vs full updates, Line-array replacement, and sensitive-data exclusions. The tool description only restates the Id/SyncToken requirement, adding no new parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Update'), a specific resource ('credit memo'), and the domain ('QuickBooks Online'). The qualifier 'existing' plus the Id/SyncToken requirement clearly scopes the tool to modifying an already-created memo rather than creating, deleting, or sending one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage context: use this tool when you need to change an existing credit memo, and the entity schema reinforces fetching the CreditMemo first to get the current SyncToken. It does not explicitly name alternative tools or give when-not-to-use conditions, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_customerupdate customerA
Destructive
Inspect

Update an existing customer in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Customer first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a customer; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructive/write, and the description goes well beyond them: sparse updates by default, full-update clearing behavior, Line array replacement semantics, forbidden sensitive fields, and the retirement-via-Active:false workaround. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is one efficient sentence with the purpose front-loaded. The parameter descriptions are long but necessary to capture QuickBooks-specific semantics; there is no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition covers all behavioral prerequisites and edge cases (sparse/full updates, lineup handling, no deletion, forbidden fields, company selection). The only gap is absence of an output schema and no mention of the return value, which an agent might need; still, the core information is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both entity and company have detailed descriptions, so the structured schema already carries the load. The tool's main description only restates the Id/SyncToken requirement already in the entity parameter description; it adds no additional parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'Update an existing customer in QuickBooks Online', which uses a specific verb (update), a specific resource (customer), and a platform. It clearly distinguishes from create_customer and other entity-specific update tools by limiting to existing customers and noting the Id/SyncToken requirement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives implicit vs explicit guidance: 'existing customer' implies this is not for creation; the entity parameter instructs to fetch the Customer first to get the SyncToken, and notes the workaround for deleting (Active: false, since QuickBooks cannot delete). It doesn't explicitly name sibling alternatives like create_customer or delete_customer, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_departmentupdate departmentA
Destructive
Inspect

Update an existing department in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Department first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a department; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true and readOnlyHint=false, and the description adds substantial behavioral detail: sparse updates, full updates clearing omitted fields, inability to delete departments, retirement via Active:false, and prohibited sensitive-field updates. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The top-level description is a single precise sentence, and the entity parameter description is dense but focused. Each sentence earns its place by conveying a critical operational rule without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description thoroughly covers the main call requirements, update semantics, and the deletion workaround. It does not explicitly describe what the response returns, which would be useful given no output schema, but the input-side guidance is strong enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description enriches the entity parameter with operational meaning: it must include Id and SyncToken, sparse behavior, line replacement, full-update semantics, sensitive-data restrictions, and the retirement workaround. The company parameter is also clearly explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb and resource: 'Update an existing department in QuickBooks Online.' Requiring Id and SyncToken clarifies that this tool targets existing records, distinguishing it from create_department, get_department, and search_departments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage conditions: requires Id and current SyncToken, instructs to fetch the Department first, explains the sparse-update default versus full update with sparse:false, and gives an alternative to deletion by using Active:false. This clearly guides when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_depositupdate depositA
Destructive
Inspect

Update an existing deposit in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Deposit first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat DepositToAccountRef, unchanged, or QuickBooks rejects it as missing.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint: true, and the description adds no behavioral detail beyond that. The sparse-update behavior, Line array replacement, and DepositToAccountRef requirement are documented in the parameter schema, not the tool description. The description offers no additional context about what happens during the update.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the purpose and key requirement. It is efficient with zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the entity parameter (sparse updates, forbidden fields, line replacement), the tool description alone is insufficient to invoke correctly without reading the schema. However, the schema is comprehensive, and the description covers the most critical precondition (Id and SyncToken). It is minimally complete but leaves key behaviors to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The tool description redundantly mentions Id and SyncToken, which is already in the entity parameter description, adding no new meaning. The rich behavioral details live in the schema, not the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update'), resource ('existing deposit in QuickBooks Online'), and a critical precondition (requires Id and SyncToken). It clearly distinguishes from create_deposit and delete_deposit by specifying 'existing'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies usage for modifying an existing deposit and mentions the need for Id/SyncToken, but it does not explicitly exclude alternatives like create_deposit or mention that you should fetch the deposit first via get_deposit. No when-not-to-use guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_employeeupdate employeeA
Destructive
Inspect

Update an existing employee in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Employee first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a employee; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the description is not required to restate mutation semantics. It adds the requirement of Id and SyncToken, which is useful operational context. However, it does not mention the sparse-update behavior, the inability to delete employees (use Active: false), or the restriction on sending sensitive fields — those details live in the schema, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that leads with the action and resource, then immediately gives the key prerequisite. There is no padding, and every word contributes to the agent's understanding of when and how to invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a nested entity parameter and a high-complexity schema, and the annotation indicates destructive behavior. The description is minimal but does highlight the critical Id and SyncToken requirement. However, it omits the nuanced update semantics (sparse/full, restrictions on sensitive data, the Active: false retirement pattern) that are only found in the schema. While the schema enriches the definition, the description alone is not fully self-sufficient for this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both 'entity' and 'company' parameters have descriptive text. The tool description itself adds no parameter-specific semantics beyond the mention of Id and SyncToken, which is redundant with the schema. Per the baseline rule, a score of 3 is appropriate when the schema carries the descriptive weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Update') and the resource ('an existing employee in QuickBooks Online'), which immediately distinguishes it from create/delete/search operations on employees. It also adds the prerequisite requirement (Id and SyncToken), which further clarifies the scope of the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating the action and prerequisite, but it does not explicitly compare against sibling tools like create_employee, get_employee, or search_employees. An agent would benefit from a direct when-to-use/alternative statement, and the brief description leaves that to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_estimateupdate estimateA
Destructive
Inspect

Update an existing estimate in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Estimate first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal a non-read-only, destructive operation, so the description does not need to establish that. It adds the useful context that updates target an existing estimate and require Id/SyncToken, but does not disclose the sparse-update or line-replacement behavior; that detail lives in the schema rather than the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler words. It states the action, the target, and the key prerequisite clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The nested entity schema carries the substantive update semantics, and annotations cover the safety profile. The description is complete enough for an agent that reads the schema, though it would be marginally stronger if it called out the sparse-update behavior or explicitly routed to create_estimate for new estimates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity property description already explains Id/SyncToken requirements, sensitive field restrictions, sparse versus full updates, and line replacement behavior. The description's mention of Id and SyncToken is accurate but redundant with the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Update') and resource ('an existing estimate in QuickBooks Online'), and the word 'existing' plus 'requires Id and SyncToken' distinguishes it from create_estimate and delete_estimate. An agent can tell this targets a previously created estimate without opening sibling schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this tool is for modifying an existing estimate and establishes the Id/SyncToken prerequisite, implying the estimate should be fetched first. It does not explicitly state 'use create_estimate for new estimates,' but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_exchange_rateupdate exchange rateA
Destructive
Inspect

Set the exchange rate QuickBooks uses for a currency on a given date (multicurrency companies only).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesThe rate to set: { SourceCurrencyCode: 'USD', AsOfDate: 'YYYY-MM-DD', Rate: 1.37 }. An exchange rate has no Id or SyncToken. Only companies with multicurrency enabled have exchange rates.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false, so the mutation nature is disclosed. The description adds useful context: it is date-scoped and limited to multicurrency companies. However, it does not explain what happens if a rate already exists for that date (overwrite), or any side effects. Given the annotations carry the safety profile, the description adds moderate behavioral context but not rich detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the core action and includes the critical constraint. Every word earns its place with no filler. It is appropriately brief for a tool with a well-covered schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a nested required parameter and no output schema. The schema provides the structure, and the description gives the core intent. However, it lacks any information about return values, error conditions, or behavior on existing rates. For a mutation tool with no output schema, an agent might need to know what to expect after invocation. Given the simplicity, a 3 is fair, but it could improve by mentioning overwrite behavior or confirmation of success.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters. The entity property has an explicit description showing the exact structure and noting that exchange rates have no Id or SyncToken. The description itself adds no additional parameter meaning beyond the schema. Since the schema is thorough, a baseline of 3 is appropriate; the description does not need to supplement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Set the exchange rate'), the resource (exchange rate), and the exact scope (for a currency on a given date). It also includes a critical constraint (multicurrency companies only), which distinguishes it from currency-related tools like update_company_currency. An agent can immediately understand what this tool does and when it applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives such as get_exchange_rate or search_exchange_rates. There is no mention of prerequisites (e.g., checking existence) or exclusions. It only states a constraint (multicurrency companies only) but does not clarify that this tool is for setting/updating rather than reading. An agent would need to infer usage from the name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_inventory_adjustmentupdate inventory adjustmentA
Destructive
Inspect

Update an existing inventory adjustment in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the InventoryAdjustment first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal destructiveHint=true and readOnlyHint=false, so the mutation risk is covered. The description adds the useful prerequisite that Id and SyncToken are required, but it does not disclose the sparse-update semantics or line-replacement behavior—those appear only in the schema. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence with no filler. It front-loads the action and resource, then states the critical prerequisite.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the description is terse, the entity schema is unusually rich and carries the update semantics, sparse behavior, line replacement, and field restrictions. Combined with the destructiveHint annotation, the overall definition gives an agent enough to invoke the tool correctly. The lack of output schema is a minor gap for an update operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter's schema already explains Id/SyncToken requirements, sparse updates, line replacement, and PII restrictions. The tool description itself adds no parameter meaning beyond echoing the required Id and SyncToken, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Update') and resource ('inventory adjustment in QuickBooks Online'), and the phrase 'existing' plus the prerequisite 'requires Id and SyncToken' clearly distinguishes it from create_inventory_adjustment, delete_inventory_adjustment, and get_inventory_adjustment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the usage context clear: updating an already-existing adjustment rather than creating or deleting one. It also states the key prerequisite (Id and SyncToken). It does not explicitly name alternative tools or list exclusion conditions, but the context is sufficient for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_invoiceupdate invoiceA
Destructive
Inspect

Update an existing invoice in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Invoice first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, indicating a write operation. The description and entity schema go well beyond this by explaining sparse updates by default, that a Line array replaces all lines, the requirement for a current SyncToken, and prohibited fields (SSN, birth dates, card details). This is rich behavioral disclosure that adds context not present in the annotations, with no contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately conveys the verb and resource, and includes the critical requirement (Id and SyncToken). There is no filler or redundancy; it is as concise as possible while still being informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with rich schema descriptions covering the entity behavior and the company parameter. The description, while minimal, is sufficient given the schema and annotations. It does not mention response format or side effects, but with no output schema and a clear update action, this is acceptable. The entity schema provides the needed operational details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage with detailed descriptions for both parameters, especially the entity object, which explains required fields, sparse behavior, and prohibited data. The tool description only restates 'requires Id and SyncToken', which is already in the schema, so it adds minimal value beyond the structured definition. The baseline of 3 applies given high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (update) and the resource (invoice) with a specific context (QuickBooks Online). It also adds a critical requirement (Id and SyncToken) that differentiates this from create, delete, or void operations. This is a precise, unambiguous purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for existing invoices by requiring Id and SyncToken, and the entity schema instructs to fetch the invoice first, which gives clear contextual guidance. It also explains sparse vs. full update behavior in the schema, aiding selection of appropriate update semantics. However, it does not explicitly contrast with sibling create/delete/void tools, though the verb 'update' and the existing-entity requirement make the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_itemupdate itemA
Destructive
Inspect

Update an existing item in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Item first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a item; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal a destructive write (destructiveHint=true, readOnlyHint=false), and the description's 'Update' is consistent with that. The description adds the Id/SyncToken precondition but does not surface the sparse-by-default, full-update, or retire-via-Active:false behaviors in the tool description itself; those are documented in the entity parameter schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that names the action, target, and prerequisite without filler or repetition of the title. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The terse tool description is supported by a rich entity parameter schema and clear annotations, giving the agent the key invocation details: required fields, sparse/full-update behavior, sensitive-field restrictions, and company selection. The main gap is the lack of any response-format guidance, but there is no output schema to lean on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity property already details Id/SyncToken requirements, sensitive-data exclusions, sparse updates, line-array replacement, and Active:false retirement. The tool description's parenthetical only restates the Id/SyncToken requirement and adds no new parameter-level meaning, so the baseline score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete action ('Update') on a specific resource ('an existing item') in QuickBooks Online, with a clear prerequisite requiring Id and SyncToken. This makes the intent unambiguous and distinguishes it from creating a new item, though it does not explicitly contrast with sibling update_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing item' and the Id/SyncToken requirement imply this is for modifying an already-created item rather than creating a new one. However, it does not explicitly name alternatives like create_item, nor does it state when not to use the tool or that the item should be fetched first.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_journal_entryupdate journal entryA
Destructive
Inspect

Update an existing journal entry in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the JournalEntry first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation risk is known. The description adds valuable behavioral context: sparse updates by default, a Line array sent replaces all lines, and sparse:false clears every writable field left out. It also warns against sending sensitive fields. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action and requirement. The additional sentences about sparse updates and sensitive fields are dense with necessary information. It could be slightly more structured, but every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the critical preconditions (Id, SyncToken, fetch first), the update semantics (sparse vs full), and the sensitive-field restriction. It does not describe the return value, but that is a minor gap given the rich behavioral detail already provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the sparse-update behavior, the line-replacement semantics, and the sensitive-field restriction. The company parameter is already well described in the schema, so the description's focus on entity behavior is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update'), a specific resource ('existing journal entry'), and the system ('QuickBooks Online'), and it names the two required identifiers (Id and SyncToken). This clearly distinguishes it from create_journal_entry and delete_journal_entry among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: it requires Id and SyncToken, and instructs the agent to fetch the JournalEntry first to obtain the current SyncToken. It does not explicitly name alternatives or when-not-to-use, but the fetch-first instruction and the sparse-update semantics provide adequate usage guidance for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_paymentupdate paymentA
Destructive
Inspect

Update an existing payment in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Payment first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already flag the operation as destructive and non-read-only, and the description adds the important precondition that a current SyncToken is required to update. The richer behavior—sparse updates, Line replacement, sensitive-field restrictions—lives in the schema rather than the description, but the description still surfaces the key versioning/precondition trait beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence states the action, object, context, and the key precondition with no filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema fully documents sparse/full update semantics, line replacement, sensitive-field prohibitions, and company selection, while annotations cover the destructive nature. For this complexity level, an agent has everything needed to invoke the tool correctly, even without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity schema already explains the required Id/SyncToken, sparse vs. full updates, Line-array replacement, restricted fields, and the optional company parameter. The tool description repeats only the Id/SyncToken requirement, adding no new parameter meaning beyond what the schema already provides, so it sits at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Update'), a specific resource ('an existing payment'), and the target system ('QuickBooks Online'). The words 'existing' and 'requires Id and SyncToken' clearly distinguish this from create_payment, get_payment, and lifecycle tools like void_payment or delete_payment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys clear usage context: use it for modifying an already-created payment, not creating or reading one, and you must first have the Id and current SyncToken. It does not explicitly name alternatives or say 'use create_payment for new payments', so it stops short of a full when-not statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_payment_methodupdate payment methodA
Destructive
Inspect

Update an existing payment method in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the PaymentMethod first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a payment method; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint:true and readOnlyHint:false, and the description aligns without contradiction. It adds valuable context beyond annotations: states that QuickBooks cannot delete a payment method, suggests using Active:false to retire, warns against sending sensitive fields, and explains sparse update semantics. This is meaningful behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise given the complexity of update semantics. It front-loads the purpose and prerequisites, then covers important caveats. Every sentence adds meaningful information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (sparse updates, full update behavior, deletion constraint, protected fields) and the absence of an output schema, the description covers the essential operational details an agent needs. It doesn't address return values, but that's acceptable without an output schema. The company parameter is simple and self-explanatory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter's description in the schema already covers the required fields, sparse behavior, and constraints. The tool description adds no new parameter-level information beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (update), resource (existing payment method), and product (QuickBooks Online). It specifies required fields (Id and SyncToken), distinguishing it from create_payment_method and search_payment_methods by the 'existing' qualifier. No ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit prerequisites: 'fetch the PaymentMethod first' and 'requires Id and SyncToken'. Explains the sparse vs. full update behavior and when to use sparse:false. Does not explicitly name alternative tools, but the context is clear that this is for modifying existing payment methods, not creating or searching.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_preferencesupdate preferencesA
Destructive
Inspect

Update the company's QuickBooks preferences. Requires a sparse payload with Id, SyncToken, sparse: true and the preference block being changed. Read-only preference fields are rejected by QuickBooks.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Preferences first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal destructiveHint=true, but the description adds meaningful behavior beyond that: sparse updates avoid clobbering, full updates clear omitted writable fields, read-only preference fields are rejected, and a current SyncToken is required. This is exactly the behavioral context an agent needs before mutating company preferences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first front-loads the purpose and the critical payload requirements, the second adds the key field-level constraint. Detailed rules are correctly deferred to the schema rather than bloating the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive update tool with rich schema coverage and annotations, this definition covers the full call cycle: obtain current Preferences for Id/SyncToken, send a sparse update, avoid read-only and sensitive fields, and understand full-update clearing behavior. The absence of an output schema is acceptable because the critical preconditions are thoroughly described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents entity and company, including sensitive-field prohibitions and sparse-vs-full semantics. The description adds explicit value by requiring sparse: true and warning that read-only fields are rejected, which helps the agent construct a valid payload. There is minor tension with the schema's 'sparse by default' phrasing, but the added guidance is still useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Update the company's QuickBooks preferences.' Among a large sibling set of update_* tools, it uniquely targets preferences and is clearly distinct from get_preferences. The sparse-payload requirement adds operational specificity without obscuring the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description and entity schema give a clear invocation context: fetch Preferences first to obtain the current SyncToken, send a sparse payload, and avoid read-only or sensitive fields. It does not explicitly name get_preferences as the read alternative or list when-not-to-use conditions, but the prerequisites and sparse/full-update behavior are well conveyed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_purchaseupdate purchaseA
Destructive
Inspect

Update an existing purchase in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Purchase first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat PaymentType and AccountRef, unchanged, or QuickBooks rejects it as missing.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations are generic (readOnly=false, destructiveHint=true), but the description/schema disclose substantial behavior: sparse updates by default, replacing lines with a Line array, sparse:false clearing omitted writable fields, replaying PaymentType and AccountRef on sparse updates, and prohibitions on sensitive identifiers. This goes well beyond what the annotations provide and there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The tool-level description is a single front-loaded sentence stating what the tool does and its key prerequisite. The more complex sparse-update details are correctly placed in the parameter schema where they belong, and every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an update tool with a nested object and destructive potential, the combined description and schema cover the prerequisites, mutation semantics, line-replacement behavior, field restrictions, and optional company selection. The absence of an explicit response shape is minor for invoking the tool, and the schema is rich enough that an agent can call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; the entity parameter already documents Id, SyncToken, sparse semantics, and forbidden fields, and the company parameter is described as optional. The tool description adds only the brief 'requires Id and SyncToken' note, so it does not materially expand on the schema — the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Update an existing purchase in QuickBooks Online.' The parenthetical 'requires Id and SyncToken' immediately distinguishes this from create/delete/search operations, and among the many sibling update_* tools the resource name and QuickBooks context make the target unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the right usage context by requiring an existing purchase with Id and SyncToken, and the schema says to fetch the Purchase first, which hints at get_purchase as a prerequisite. However, it never explicitly names alternatives such as create_purchase for new purchases or states when not to use this tool, so the routing guidance is implied rather than direct.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_purchase_orderupdate purchase orderA
Destructive
Inspect

Update an existing purchase order in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the PurchaseOrder first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The input schema adds substantial behavioral detail beyond the destructiveHint annotation: sparse updates are the default, a Line array replaces all lines, and sparse:false clears every writable field left out. It also warns against sending government IDs, birth dates, or card-processing details, which is valuable operational context for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The top-level description is a single front-loaded sentence with no wasted words. The longer entity description is dense but each sentence earns its place by explaining update semantics, destructive behavior, or sensitive-field restrictions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with nested parameters, the definition covers prerequisites, sparse vs. full update behavior, line replacement, and sensitive-field handling. It does not describe the return value, but with no output schema specified for this update operation, that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter description thoroughly documents requirements, sparse semantics, and sensitive-field restrictions. The tool-level description adds no additional parameter meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Update'), a resource ('purchase order'), and a scope ('existing ... in QuickBooks Online'), which clearly distinguishes it from create or send operations. It does not explicitly contrast sibling tools, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the core precondition ('requires Id and SyncToken'), and the entity schema reinforces it with 'fetch the PurchaseOrder first.' This gives clear context for when the tool is appropriate, though it does not explicitly name alternatives like create_purchase_order or get_purchase_order.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_refund_receiptupdate refund receiptA
Destructive
Inspect

Update an existing refund receipt in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the RefundReceipt first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the description's 'update' is consistent but adds little beyond the requirement for Id and SyncToken. It does not disclose sparse update behavior or line replacement, which are relegated to schema description. No contradiction; limited added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core action and essential requirement. It is appropriately concise with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal but adequate given the rich schema description for the entity parameter. It covers the core 'what' and the precondition, but does not mention important behavioral nuances like sparse vs full updates, which are only in the schema. Since no output schema exists, return value expectations are unaddressed, though not explicitly required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed descriptions for both entity and company parameters. The tool description adds no parameter meaning beyond restating the Id and SyncToken requirement already present in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the verb 'update', the resource 'refund receipt', and qualifies it as 'existing', which clearly distinguishes it from create or get tools. It also adds the requirement of Id and SyncToken, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context (existing refund receipt, requires Id and SyncToken) but does not mention alternatives or when not to use it. It implies update vs create but doesn't explicitly say 'use create_refund_receipt for new receipts', leaving selection partially to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_sales_receiptupdate sales receiptA
Destructive
Inspect

Update an existing sales receipt in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the SalesReceipt first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation risk is covered by structured data. The description adds only the Id/SyncToken prerequisite; richer behaviors like sparse updates, line replacement, and cleared fields live in the input schema rather than in the tool description itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the operation and its key prerequisite with no filler. It is appropriately sized for a tool whose input schema carries the detailed parameter semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema descriptions and annotations together make the tool callable with confidence, covering required fields, protected data, and destructive intent. The main gap is the lack of an output schema or return-value information, but that is minor for a straightforward update operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the entity parameter is richly documented: required Id/SyncToken, forbidden sensitive fields, sparse-by-default behavior, Line array replacement semantics, and the effect of sparse:false. The description itself merely echoes the Id/SyncToken requirement, so the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the verb ('Update'), the resource ('existing sales receipt'), and the system ('QuickBooks Online'). The phrase 'existing' and the Id/SyncToken requirement distinguish this from create, delete, void, and send operations on the same document type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when modifying an already-created sales receipt and when you have the current Id and SyncToken. However, it does not explicitly name alternatives or state when not to use it versus create_sales_receipt, void_sales_receipt, delete_sales_receipt, or search_sales_receipts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_termupdate termA
Destructive
Inspect

Update an existing term in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Term first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat Name, unchanged, or QuickBooks rejects it as missing. QuickBooks cannot delete a term; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation is known. The description adds valuable context: it clarifies the sparse vs full update semantics (only sent fields change, Line array replaces all), the requirement to repeat Name for sparse updates, and the impossibility of deletion. It also warns about sensitive fields (SSN, birth dates) that must not be sent. This significantly exceeds the annotation baseline, though it doesn't describe response format or error handling, hence 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense paragraph that front-loads the verb and resource, and then packs in critical operational details without redundancy. Every sentence carries weight: the SyncToken requirement, the sparse update behavior, the Name repetition rule, and the deletion workaround. It's concise (around 150 words) but no word is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and relatively simple parameters (2), the description covers the key pitfalls and behaviors. It explains the semantics of sparse vs full updates, the irreversibility of deletion, and the need to fetch the current SyncToken. It doesn't describe the exact response shape or error scenarios, but given that the schema covers parameter types and the description covers usage pitfalls, it's nearly complete. The only minor gap is lack of explicit mention of response structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents both parameters (entity and company). The description adds nuance about the entity object's behavior (sparse updates, Line replacement, sensitive field restrictions), which is helpful. However, it doesn't add meaning to the 'company' parameter beyond the schema's description. Since the schema already covers the basics, a baseline of 3 is appropriate, with the extra context justifying keeping it at 3 (not higher because no new parameter-specific details).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific action ('Update an existing term') and the resource ('term in QuickBooks Online'), with the key requirement 'requires Id and SyncToken'. It clearly distinguishes this from create_term and delete operations on terms, and the context of sibling tools (update_term vs create_term vs delete_*) makes the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: 'fetch the Term first' to get SyncToken, and essential operational details like sparse update behavior, the need to repeat Name, and how to retire a term (since deletion isn't possible). It also warns against sending sensitive fields, which is crucial for correct usage. This is effectively a mini-user manual for proper invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_time_activityupdate time activityA
Destructive
Inspect

Update an existing time activity in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the TimeActivity first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the description adds the requirement of Id and SyncToken, which is a behavioral prerequisite (must fetch first). However, it does not disclose the sparse update behavior, line replacement semantics, or restrictions on sensitive fields—those are only in the schema description, not the tool description. The description provides minimal behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the purpose and the critical prerequisite. Every word earns its place; there is no fluff or redundancy. It is appropriately sized for a tool with a rich schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested entity object, sparse update semantics, sensitive field restrictions), the tool description is minimal. It covers the core purpose and the Id/SyncToken requirement, but it does not mention the sparse behavior, line replacement, or the need to avoid sending sensitive fields. However, the schema description is comprehensive, so an agent reading the schema will have the necessary details. Still, the description itself is not fully complete for an agent that relies on it as the primary guide.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (entity and company) in detail, including the sparse update behavior, Id/SyncToken requirements, and field restrictions. The tool description itself adds no parameter-level meaning, so the baseline score of 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update'), a specific resource ('existing time activity in QuickBooks Online'), and the required fields (Id and SyncToken). This clearly distinguishes it from create_time_activity (creates), delete_time_activity (deletes), and search_time_activities (searches). No ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (updating an existing record) and provides a prerequisite (requires Id and SyncToken), but it does not explicitly state when to use this vs. alternatives, nor does it mention that the entity must be fetched first. It lacks explicit when-not-to-use guidance or naming of sibling tools, though the verb and resource make the primary purpose clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_transferupdate transferA
Destructive
Inspect

Update an existing transfer in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Transfer first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat FromAccountRef, ToAccountRef and Amount, unchanged, or QuickBooks rejects it as missing.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false, so the mutating nature is covered. The description adds the prerequisite of Id and SyncToken, which is useful, but it does not disclose behaviors like sparse updates or the fact that a Line array replaces all lines—those details live in the parameter schema, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the purpose and the key requirement. There is zero waste; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema provides extensive detail on the entity parameter, including sparse update semantics and required fields. However, the description does not mention what the tool returns (no output schema) or that the transfer must be fetched first to obtain the current SyncToken (though that is in the schema). For a mutating tool with no output schema, the description could be more complete, but the schema compensates significantly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both entity and company have detailed descriptions. The tool description itself adds nothing about parameters beyond what the schema already states. Since the schema carries the full weight, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool updates an existing transfer in QuickBooks Online, and explicitly notes the requirement for Id and SyncToken. This distinguishes it from create_transfer (new) and delete_transfer (removal) without needing to name siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for existing transfers via the word 'existing,' but it does not explicitly mention alternatives like create_transfer or delete_transfer, nor does it state when to prefer this tool over other update tools. The guidance is clear but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_vendorupdate vendorA
Destructive
Inspect

Update an existing vendor in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the Vendor first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. QuickBooks cannot delete a vendor; to retire one, update it with Active: false.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the existing destructiveHint annotation by explaining sparse updates, line-array replacement, full-update behavior, restricted field categories, and the Active:false retirement pattern. This gives the agent a strong understanding of the operation's side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is a single, focused sentence, and the parameter schema contains detailed but well-organized guidance. Minor redundancy exists between the main description's 'requires Id and SyncToken' and the schema's 'Must include Id and the current SyncToken', but overall the definition is efficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition covers prerequisites, restricted fields, sparse vs. full update semantics, line replacement, the vendor-deletion limitation, and company selection. Given the rich schema and annotations, the lack of an output schema does not represent a meaningful gap for calling this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already provides detailed guidance for both entity and company parameters. The main description only restates the Id and SyncToken requirement, which is already in the schema, so it adds no new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Update an existing vendor') and the target system ('QuickBooks Online'), with a concise prerequisite. This distinguishes it from create_vendor, search_vendors, and vendor-credit tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description and entity schema explain that Id and SyncToken are required, instruct the agent to fetch the Vendor first, and describe how to effectively retire a vendor since deletion is not possible. It provides clear context for when to use this tool, though it does not explicitly contrast it with create_vendor or other alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_vendor_creditupdate vendor creditA
Destructive
Inspect

Update an existing vendor credit in QuickBooks Online (requires Id and SyncToken).

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesFields to update. Must include Id and the current SyncToken (fetch the VendorCredit first). Do not send government identifiers (SSN or tax IDs), birth dates or card-processing details. Manage those fields directly in QuickBooks. Updates are sparse by default: only the fields sent change, and a Line array sent replaces all lines. Send sparse: false for a full update, which clears every writable field left out. A sparse update must also repeat VendorRef, unchanged, or QuickBooks rejects it as missing.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true, and the description implicitly confirms this by describing update semantics. The description adds key behavioral context like sparse vs full updates, which is not in annotations. However, it doesn't discuss authorization requirements or potential side effects beyond the update itself. Since annotations already cover destructiveness, the burden is lower, and the description adds some value but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise, one sentence, and gets the essential requirement (Id and SyncToken) early. It doesn't over-explain but also doesn't waste words. It's well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the update logic (sparse vs full, repeating VendorRef), the description and schema together cover most needed details. No output schema exists, so return format is not explained, but that's not described in the tool description. It might be helpful to note common failure scenarios, but overall it's reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage, so it already explains the entity and company parameters in detail. The description adds little beyond what the schema provides, leaning mostly on the schema's extensive property descriptions. Baseline 3 is appropriate since schema covers everything.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The tool description explicitly states it updates an existing vendor credit in QuickBooks Online and requires Id and SyncToken. Since there are many other update tools for different resources, the specific resource (vendor credit) is clear. However, it doesn't explicitly differentiate itself from other update tools, but given the resource name is part of the tool name, that is mostly self-evident.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear guidance on the required Id and SyncToken, mentions fetching the VendorCredit first, and warns about sensitive fields to avoid. It also explains the sparse update behavior and the need to repeat VendorRef for sparse updates. It doesn't explicitly say when to not use this tool or what alternative to use, but the guidance is specific and helpful for correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

void_bill_paymentvoid bill paymentA
Destructive
Inspect

Void a bill payment in QuickBooks Online, keeping the record and its number but zeroing its amounts. Cannot be undone; requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the BillPayment to void.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this void.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the destructive nature is known. The description adds valuable context beyond the annotations: the record and its number are preserved, amounts are zeroed, the operation cannot be undone, and explicit user confirmation is required. This is meaningful behavioral disclosure for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. The core action and its key behavioral consequence are front-loaded, followed by the irreversibility warning and confirmation requirement. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-resource operation with a fully documented schema and annotations covering the safety profile, the description is nearly complete. It covers what happens (void, zero amounts), what doesn't happen (record/number preserved), and the required precondition (explicit confirmation). It doesn't describe the return value, but there is no output schema and the operation's outcome is clear enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (id, company, confirm). The description adds the critical semantic that 'confirm' must be true and only set after explicit user confirmation, which reinforces the schema. However, it doesn't add much beyond that, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Void'), a specific resource ('a bill payment in QuickBooks Online'), and the precise behavioral scope: it keeps the record and its number but zeroes its amounts. This clearly distinguishes it from delete_bill_payment and update_bill_payment, which are the closest siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when the goal is to void rather than delete or update a bill payment. It does not explicitly name alternatives like delete_bill_payment or update_bill_payment, but the contrast is strongly implied by the phrase 'keeping the record and its number but zeroing its amounts.' The explicit confirmation requirement also gives clear usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

void_invoicevoid invoiceA
Destructive
Inspect

Void an invoice in QuickBooks Online, keeping the record and its number but zeroing its amounts. Cannot be undone; requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Invoice to void.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this void.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description adds critical behavioral facts: the action cannot be undone and requires explicit user confirmation. It also explains the exact effect on the invoice record, which is more specific than the generic destructive hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence front-loads the purpose and distinguishing effect, and the second sentence delivers the irreversibility and confirmation requirement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation with three well-documented parameters, destructive annotation, and no output schema, the description covers the essential facts an agent needs: what voiding does, that it is irreversible, and that user confirmation is mandatory. No critical operational context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all three parameters. The description reinforces that confirmation is required but does not add meaning beyond the schema's 'Must be true' note, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Void an invoice') and clarifies the operation's semantic scope: it keeps the record and number while zeroing amounts. This distinguishes void_invoice from delete_invoice and update_invoice without needing to inspect schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the intended use case clear: voiding when you want to preserve the record and number rather than delete it. It does not explicitly name alternatives like delete_invoice, but the 'keeping the record and its number but zeroing its amounts' phrasing gives strong contextual guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

void_paymentvoid paymentA
Destructive
Inspect

Void a payment in QuickBooks Online, keeping the record and its number but zeroing its amounts. Cannot be undone; requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the Payment to void.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this void.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the destructiveHint annotation: the record stays, the number stays, amounts become zero, the operation cannot be undone, and explicit user confirmation is required. These details tell the agent exactly what happens and what safeguards must precede invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences front-load the core action and effect, then add the two critical warnings. Every phrase earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-entity action with fully documented parameters and a destructiveHint annotation, the description captures the effect, irreversibility, and confirmation requirement. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents id, company, and confirm. The description reinforces the confirm requirement, but does not need to add new parameter-level semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Void a payment in QuickBooks Online') and clarifies the exact behavioral distinction: the record and its number are kept while amounts are zeroed. This implicitly differentiates void_payment from delete_payment and from create/update payment siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Describes clear operational context: the action is irreversible and requires explicit user confirmation, which tells the agent to gate the call on user approval. It does not explicitly name alternatives like delete_payment or void_invoice, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

void_sales_receiptvoid sales receiptA
Destructive
Inspect

Void a sales receipt in QuickBooks Online, keeping the record and its number but zeroing its amounts. Cannot be undone; requires explicit user confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesQuickBooks Id of the SalesReceipt to void.
companyNoWhich connected QuickBooks company to use (name, realm id, or connection id). Optional when only one company is connected.
confirmNoMust be true. Only set after the user has explicitly confirmed this void.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already mark the tool as destructive, the description adds important behavioral details: it is irreversible, requires explicit user confirmation, and preserves the record and number while zeroing amounts. This goes meaningfully beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one dense, front-loaded sentence that states the action, the effect, and the two most critical cautions without any filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an irreversible destructive mutation with no output schema, the description covers the essential context: what the tool does, what it preserves, that it cannot be undone, and that confirmation is required. It does not specify the response format, but that is less critical here given the clear side effects and existing annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description's mention of explicit user confirmation reinforces the confirm parameter, but that information is already in the schema, so the description adds no new parameter-level meaning beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact verb ('Void'), the resource ('sales receipt'), and the distinguishing effect ('keeping the record and its number but zeroing its amounts'). This clearly separates it from delete_sales_receipt, which would remove the record, and update_sales_receipt, which edits it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool by explaining that it zeroes amounts while preserving the record, which contrasts with deletion. However, it never explicitly names alternatives like delete_sales_receipt or states when not to use this tool, so the guidance is implied rather than direct.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedlist_receipt_inbox5 fields changed
      • addedInput schema / properties / from
        Added value: +{
        +  "description": "Only documents dated on or after this day, YYYY-MM-DD, by the date written on them.",
        +  "pattern": "^(\\d{4}-\\d{2}-\\d{2})?$",
        +  "type": "string"
        +}
      • addedInput schema / properties / max_total
        Added value: +{
        +  "description": "Only documents whose total is at most this amount.",
        +  "type": "number"
        +}
      • addedInput schema / properties / min_total
        Added value: +{
        +  "description": "Only documents whose total is at least this amount.",
        +  "type": "number"
        +}
      • addedInput schema / properties / search
        Added value: +{
        +  "description": "Words every matching document contains, in its vendor, number, summary, lines bought, text, file name, sender or subject. Case and accents do not matter. Searches filed documents too.",
        +  "maxLength": 200,
        +  "type": "string"
        +}
      • addedInput schema / properties / to
        Added value: +{
        +  "description": "Only documents dated on or before this day, YYYY-MM-DD.",
        +  "pattern": "^(\\d{4}-\\d{2}-\\d{2})?$",
        +  "type": "string"
        +}
  2. 220 tool updates
    • First observedactivate_loop
    • First observedattach_file
    • First observedattach_receipt
    • First observedcreate_account
    • First observedcreate_attachable
    • First observedcreate_bill
    • First observedcreate_bill_payment
    • First observedcreate_budget
    • First observedcreate_class
    • First observedcreate_company_currency
    • First observedcreate_credit_card_payment_txn
    • First observedcreate_credit_memo
    • First observedcreate_customer
    • First observedcreate_department
    • First observedcreate_deposit
    • First observedcreate_employee
    • First observedcreate_estimate
    • First observedcreate_inventory_adjustment
    • First observedcreate_invoice
    • First observedcreate_item
    • First observedcreate_journal_entry
    • First observedcreate_payment
    • First observedcreate_payment_method
    • First observedcreate_purchase
    • First observedcreate_purchase_order
    • First observedcreate_recurring_transaction
    • First observedcreate_refund_receipt
    • First observedcreate_sales_receipt
    • First observedcreate_task
    • First observedcreate_tax_agency
    • First observedcreate_tax_service
    • First observedcreate_term
    • First observedcreate_time_activity
    • First observedcreate_transfer
    • First observedcreate_vendor
    • First observedcreate_vendor_credit
    • First observeddelete_attachable
    • First observeddelete_bill
    • First observeddelete_bill_payment
    • First observeddelete_budget
    • First observeddelete_company_currency
    • First observeddelete_credit_card_payment_txn
    • First observeddelete_credit_memo
    • First observeddelete_deposit
    • First observeddelete_estimate
    • First observeddelete_inventory_adjustment
    • First observeddelete_invoice
    • First observeddelete_journal_entry
    • First observeddelete_payment
    • First observeddelete_purchase
    • First observeddelete_purchase_order
    • First observeddelete_recurring_transaction
    • First observeddelete_refund_receipt
    • First observeddelete_sales_receipt
    • First observeddelete_time_activity
    • First observeddelete_transfer
    • First observeddelete_vendor_credit
    • First observeddiscard_inbox_file
    • First observedfind_expenses_without_attachments
    • First observedget_account
    • First observedget_account_list
    • First observedget_aged_payable_detail
    • First observedget_aged_payables
    • First observedget_aged_receivable_detail
    • First observedget_aged_receivables
    • First observedget_attachable
    • First observedget_balance_sheet
    • First observedget_balance_sheet_detail
    • First observedget_balance_sheet_summary
    • First observedget_bill
    • First observedget_bill_payment
    • First observedget_budget
    • First observedget_budget_vs_actuals
    • First observedget_cash_flow
    • First observedget_changes
    • First observedget_class
    • First observedget_class_sales
    • First observedget_company_currency
    • First observedget_company_info
    • First observedget_credit_card_payment_txn
    • First observedget_credit_memo
    • First observedget_customer
    • First observedget_customer_balance
    • First observedget_customer_balance_detail
    • First observedget_customer_income
    • First observedget_customer_sales
    • First observedget_customer_type
    • First observedget_department
    • First observedget_department_sales
    • First observedget_deposit
    • First observedget_employee
    • First observedget_estimate
    • First observedget_exchange_rate
    • First observedget_general_ledger
    • First observedget_inventory_adjustment
    • First observedget_inventory_valuation_detail
    • First observedget_inventory_valuation_summary
    • First observedget_invoice
    • First observedget_item
    • First observedget_item_sales
    • First observedget_journal_entry
    • First observedget_journal_report
    • First observedget_open_invoices
    • First observedget_payment
    • First observedget_payment_method
    • First observedget_physical_inventory_worksheet
    • First observedget_preferences
    • First observedget_profit_and_loss
    • First observedget_profit_and_loss_detail
    • First observedget_purchase
    • First observedget_purchase_order
    • First observedget_recurring_transaction
    • First observedget_refund_receipt
    • First observedget_sales_receipt
    • First observedget_tax_agency
    • First observedget_tax_code
    • First observedget_tax_payment
    • First observedget_tax_rate
    • First observedget_tax_summary
    • First observedget_term
    • First observedget_time_activity
    • First observedget_transaction_detail_by_account
    • First observedget_transaction_list
    • First observedget_transaction_list_by_customer
    • First observedget_transaction_list_by_vendor
    • First observedget_transaction_list_with_splits
    • First observedget_transfer
    • First observedget_trial_balance
    • First observedget_vendor
    • First observedget_vendor_balance
    • First observedget_vendor_balance_detail
    • First observedget_vendor_credit
    • First observedget_vendor_expenses
    • First observedlist_connected_companies
    • First observedlist_loops
    • First observedlist_proposals
    • First observedlist_receipt_inbox
    • First observedlist_rules
    • First observedremember_rule
    • First observedretire_rule
    • First observedsearch_accounts
    • First observedsearch_attachables
    • First observedsearch_bill_payments
    • First observedsearch_bills
    • First observedsearch_budgets
    • First observedsearch_classes
    • First observedsearch_company_currencies
    • First observedsearch_company_infos
    • First observedsearch_credit_card_payment_txns
    • First observedsearch_credit_memos
    • First observedsearch_customer_types
    • First observedsearch_customers
    • First observedsearch_departments
    • First observedsearch_deposits
    • First observedsearch_employees
    • First observedsearch_estimates
    • First observedsearch_exchange_rates
    • First observedsearch_invoices
    • First observedsearch_items
    • First observedsearch_journal_entries
    • First observedsearch_payment_methods
    • First observedsearch_payments
    • First observedsearch_purchase_orders
    • First observedsearch_purchases
    • First observedsearch_recurring_transactions
    • First observedsearch_refund_receipts
    • First observedsearch_reimburse_charges
    • First observedsearch_sales_receipts
    • First observedsearch_tax_agencies
    • First observedsearch_tax_codes
    • First observedsearch_tax_payments
    • First observedsearch_tax_rates
    • First observedsearch_terms
    • First observedsearch_time_activities
    • First observedsearch_transfers
    • First observedsearch_vendor_credits
    • First observedsearch_vendors
    • First observedsend_credit_memo
    • First observedsend_estimate
    • First observedsend_invoice
    • First observedsend_payment
    • First observedsend_purchase_order
    • First observedsend_refund_receipt
    • First observedsend_sales_receipt
    • First observedupdate_account
    • First observedupdate_attachable
    • First observedupdate_bill
    • First observedupdate_bill_payment
    • First observedupdate_budget
    • First observedupdate_class
    • First observedupdate_company_currency
    • First observedupdate_company_info
    • First observedupdate_credit_card_payment_txn
    • First observedupdate_credit_memo
    • First observedupdate_customer
    • First observedupdate_department
    • First observedupdate_deposit
    • First observedupdate_employee
    • First observedupdate_estimate
    • First observedupdate_exchange_rate
    • First observedupdate_inventory_adjustment
    • First observedupdate_invoice
    • First observedupdate_item
    • First observedupdate_journal_entry
    • First observedupdate_payment
    • First observedupdate_payment_method
    • First observedupdate_preferences
    • First observedupdate_purchase
    • First observedupdate_purchase_order
    • First observedupdate_refund_receipt
    • First observedupdate_sales_receipt
    • First observedupdate_term
    • First observedupdate_time_activity
    • First observedupdate_transfer
    • First observedupdate_vendor
    • First observedupdate_vendor_credit
    • First observedvoid_bill_payment
    • First observedvoid_invoice
    • First observedvoid_payment
    • First observedvoid_sales_receipt

Publisher details

Operator
Peich Technologies, Inc. · Publisher source
Vendor relationship
Independent
Restrictions
Requires a QuickBooks Online company. 14-day free trial, then a paid plan (from $39/month). Canada and US.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    The AccountingQB-MCP server is a comprehensive QuickBooks Online integration for Claude, offering 108 tools to manage transactions, run reports, and prepare taxes for US and Canada through natural language. It supports sole proprietors and small businesses with features like GST/HST returns, 1099/T4A reporting, reconciliation, and cash flow forecasting.
    -
  • A
    license
    A
    quality
    D
    maintenance
    AI bookkeeper for small businesses that connects to QuickBooks Online. Enables users to query financial data like bank balances, P\&L reports, and invoices through natural language in Claude Desktop or Cursor.
    6
    40 npm
    MIT
  • A
    license
    C
    quality
    D
    maintenance
    Provides complete QuickBooks Online API integration for Claude Code and other MCP-compatible clients, enabling full CRUD operations on 29 entity types and 11 financial reports.
    100
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Gives Claude full read/write access to QuickBooks Online, including tools to create, delete, and query accounting data, plus reconciliation features that detect bank-feed gaps and duplicate transactions.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources