Skip to main content
Glama
Tobeworks

invoiceshelf-mcp

by Tobeworks

invoiceshelf-mcp

MCP server for InvoiceShelf (self-hosted invoicing). Clean-room rewrite — not a fork — built against the current @modelcontextprotocol/sdk (McpServer + registerTool), Zod 4, and native fetch. No axios, no runtime dependencies beyond the SDK and Zod.

Why this exists

An earlier fork of a third-party MCP server for InvoiceShelf had two bugs (missing required fields on create_invoice and send_invoice — see below) and no license, so a PR wasn't a clean option. Since the API's real behavior was already fully verified by then, this is a fresh implementation with the same 23-tool surface and both bugs fixed from the start.

Related MCP server: QuickBooks Online MCP Server

Setup

pnpm install
cp .env.example .env   # fill in your instance URL + API token
pnpm run build
pnpm run test           # smoke test: lists tools, calls test_connection

Add to your MCP client config:

{
  "mcpServers": {
    "invoiceshelf": {
      "command": "node",
      "args": ["/absolute/path/to/invoiceshelf-mcp/dist/index.js"],
      "env": {
        "INVOICE_SHELF_BASE_URL": "https://your-instance.example.com/api/v1",
        "INVOICE_SHELF_API_TOKEN": "1|xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
      }
    }
  }
}

InvoiceShelf API quirks handled here

InvoiceShelf's REST API requires several fields on write endpoints that aren't documented and produce a 422 if omitted:

  • create_invoice / create_estimate: invoice_number/estimate_number (fetched from /next-number), exchange_rate, discount_type, discount, discount_val, and computed sub_total/tax/total plus per-item totals. Handled in src/money.ts and the create_* tools — callers just pass euro amounts and line items.

  • send_invoice / send_estimate: the send endpoint needs from and to in the body. to comes from the customer's email on the invoice/estimate record, from from the company's mail_username setting. Both are looked up automatically.

Tools

23 tools across invoices, estimates, customers, and account info/dashboard — see src/tools/ for the full list and schemas.

M8ven Score

License

MIT — see LICENSE.

Available Tools

23 tools
convert_estimate_to_invoiceC

Convert an accepted estimate into an invoice.

ParametersJSON Schema
NameRequiredDescriptionDefault
estimateIdYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only indicate it is not idempotent and not destructive. The description adds no additional behavioral context—it does not state what happens to the estimate (e.g., marked as converted), whether a new invoice is created, or any side effects. With minimal annotation coverage, the description carries the burden but fails to disclose these details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no fluff. It is front-loaded with the core action. However, it is under-specified for a tool with behavioral nuance, but conciseness itself is not penalized—the lack of detail is captured in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description still leaves gaps: it does not describe the return value, any prerequisites beyond 'accepted', or what constitutes a successful conversion. Given the lack of an output schema, the description should provide at least some hint about the result, but it does not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero description coverage for the single parameter estimateId, and the description does not mention it at all. The agent gets no meaning beyond the type 'number' from the schema, and the description adds no clarification about what the ID represents or any constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('convert') and resource ('estimate' to 'invoice'), and the phrase 'accepted estimate' adds specificity that distinguishes it from generic create_invoice. However, it does not explicitly name alternatives or elaborate on the transformation, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus create_invoice or other siblings. The implication that an accepted estimate is the prerequisite is present, but no alternatives or exclusions are mentioned, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_customerC

Create a new customer.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
emailNo
phoneNo
currency_idNoDefaults to 1 (the company's base currency).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (idempotentHint: false, destructiveHint: false), the description adds no behavioral context. It does not disclose whether duplicate customers are allowed, whether permissions are required, what happens on creation, or what the response contains. For a mutation-style tool, this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with no filler words. However, it is under-specified rather than efficiently informative, providing a bare operational statement without the additional context needed for a 4-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, minimal annotations, and low parameter schema coverage, the description leaves too much unstated. The agent is left without guidance on required parameters, parameter semantics, side effects, or expected return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, with only currency_id documented. The description does not compensate by explaining the meaning or constraints of the required name parameter, nor email and phone. It adds essentially no parameter-level information beyond the property names already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('create') and a specific resource ('customer'), which clearly distinguishes it from sibling tools like get_customer, update_customer, and delete_customer, as well as create_invoice and create_estimate. It is unambiguous about what action this tool performs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as update_customer or get_customers. It provides no context about prerequisites, workflow positioning, or scenarios where another customer-related tool would be preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_estimateB

Create a new estimate. Handles InvoiceShelf's undocumented required fields internally.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
notesNo
customer_idYes
expiry_dateYesYYYY-MM-DD
estimate_dateYesYYYY-MM-DD
template_nameNoDefaults to "tobeworks".
reference_numberNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=false and destructiveHint=false, shifting some burden off the description. The description adds one useful behavioral fact: the tool internally handles InvoiceShelf's undocumented required fields. However, it does not disclose any side effects, such as whether the estimate is automatically sent, what status it gets, or what happens on validation failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the primary purpose and then adds a useful implementation detail, making every word earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create operation with 7 parameters, nested items, low schema coverage, and no output schema, this description is too sparse. It leaves out the return value, post-creation state, and any preconditions. The internal-fields note is helpful context but insufficient for an agent to fully understand the tool's complete behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%, so the description needs to compensate for undocumented parameters like customer_id, items, notes, and reference_number. It does not explain any of these parameters. The statement about undocumented required fields is generic and reassuring, but not parameter-specific.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Create a new estimate', which clearly specifies the verb and resource. This cleanly differentiates it from sibling tools like get_estimate, update_estimate, delete_estimate, and create_invoice. The additional context about handling undocumented fields does not obscure the core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as create_invoice or convert_estimate_to_invoice. It also does not mention prerequisites like the customer needing to exist before creating an estimate. Usage context is only implicitly inferred from the tool name and first sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_invoiceA

Create a new invoice. Handles InvoiceShelf's undocumented required fields internally.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
notesNo
due_dateYesYYYY-MM-DD
customer_idYes
tax_percentNoAdds the existing tax type with this percentage (e.g. 19) on the whole invoice. Omit for no tax.
invoice_dateYesYYYY-MM-DD
template_nameNoDefaults to "tobeworks".
invoice_numberNoAuto-generated via /next-number if omitted.
reference_numberNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide idempotentHint=false and destructiveHint=false, but no readOnlyHint. The description adds the behavioral detail that required fields are handled internally, which is helpful. Yet it does not mention side effects (e.g., invoice status, numbering, or notifications). Given the existing annotations, the description adds some value but stays minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The primary action is front-loaded, and the secondary sentence adds a meaningful detail about internal handling. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool without an output schema, the description is sparse. It addresses one potential confusion (undocumented required fields) but leaves out other contextual details that might be needed, such as whether the invoice is created as a draft, how taxes are applied, or any prerequisites like authentication. The schema covers some, but the tool could benefit from more comprehensive guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 56%, leaving 44% of parameters undocumented. The description provides zero parameter-level guidance, and does not compensate for the schema gaps. For example, 'reference_number' and 'notes' have no schema descriptions, and the tool description does not clarify them. The description adds no semantic value beyond what the schema already offers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create a new invoice' clearly states the action and resource, and the verb 'create' distinguishes it from siblings like update_invoice, delete_invoice, and create_estimate. It leaves no ambiguity about the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Handles InvoiceShelf's undocumented required fields internally' provides a usage hint that the caller doesn't need to supply certain fields. However, it does not explicitly contrast with alternatives (e.g., 'Use create_estimate for estimates') or state when to prefer this over other creation tools. Usage context is implied but not fully articulated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_customerB
DestructiveIdempotent

Delete a customer by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
customerIdYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, but the description adds no behavioral context beyond the word 'Delete'. It does not mention permanence, cascading effects, failure conditions, or any post-deletion impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. For a one-parameter tool, this is appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the tool is simple and annotations cover destructive/idempotent aspects, the description omits outcome details, edge conditions, and what happens after deletion. With no output schema, the agent is left guessing about side effects and expected behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description should compensate by explaining the customerId parameter. 'By ID' merely restates the property name and adds no information about format, source, or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Delete'), a clear resource ('customer'), and the selection method ('by ID'). This clearly distinguishes it from sibling tools like get_customer, update_customer, or delete_invoice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when deletion is appropriate, what prerequisites might exist, or which alternatives to prefer. The only implied context is in the word 'Delete' itself, which is insufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_estimateA
DestructiveIdempotent

Delete an estimate by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
estimateIdYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already disclose destructiveHint=true and idempotentHint=true, and the description adds no behavioral context beyond that. It does not mention side effects, soft-delete behavior, return value on success, or irreversibility, so the description itself provides minimal transparency beyond 'by ID'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no waste. It states the verb, resource, and target mechanism directly, which is ideal for a simple delete operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter delete tool with destructive and idempotent annotations, the description plus annotations are nearly sufficient. It lacks explicit statements about what the response looks like or whether deletion cascades, but the simplicity of the operation and the existing annotations cover the essential context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does convey that the operation is keyed by an estimate ID, but it does not clarify the format, origin, or meaning of estimateId beyond the property name already present in the schema. This is minimum viable for a single obvious parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Delete'), the resource ('estimate'), and the identifying mechanism ('by ID'). It distinguishes itself from siblings like update_estimate and convert_estimate_to_invoice without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than explicit: the name and description make it obvious you call this when you want to remove an estimate. It does not provide any exclusion criteria, alternatives, or prerequisites, so it does not fully guide an agent's decision among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_invoiceA
DestructiveIdempotent

Delete an invoice by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
invoiceIdYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true and idempotentHint=true, so the destructive nature is covered. The description adds no additional behavioral context beyond those annotations, but it does not contradict them either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is appropriately sized for a simple one-parameter delete tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter delete operation with annotations already describing destructive and idempotent behavior, the description is largely complete. It doesn't describe return values or side effects, but the lack of an output schema and the tool's simplicity make this acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage for 'invoiceId'. The phrase 'by ID' indicates the parameter's role, but it adds little beyond the property name and doesn't mention format or usage details beyond the schema's number type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb, resource, and selection key: 'Delete an invoice by ID.' It is clearly distinct from sibling tools like delete_estimate and delete_customer, and from get/update/send invoice operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the verb and resource, but no explicit when-to-use guidance or alternatives are provided. It doesn't say to prefer this over delete_estimate or delete_customer, though the resource is obvious from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customerA
Read-only

Get full details for a single customer by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
customerIdYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description is consistent with that. The phrase 'full details' adds some indication of return richness, but the description does not disclose other behavioral traits like response shape, permissions, or edge cases. That is acceptable given the annotation, but the description itself contributes little beyond it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource with no filler. Every word contributes meaning, and the structure makes the tool's purpose immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, read-only lookup tool, the description plus input schema provide enough to call it correctly. There is no output schema, and 'full details' is somewhat vague about what fields are returned, so a slightly more specific return description would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description partially compensates by saying 'by ID', which tells the agent customerId is the lookup key. However, it adds no detail about the ID format, uniqueness, or behavior when the customer is not found. With one obvious numeric parameter, this is minimally sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), resource ('customer'), and scope ('single ... by ID'), which clearly distinguishes it from sibling tools like get_customers (plural) and get_customer_invoices. An agent can identify it as the targeted single-customer lookup without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: retrieve one customer when you have an ID. However, it never explicitly says when to prefer this over alternatives or mentions any exclusion, such as 'use get_customers to list customers'. The guidance is minimal but not misleading.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customer_invoicesA
Read-only

Get all invoices for a specific customer, optionally filtered by status.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNo
customerIdYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already indicates a safe read operation, and the description aligns with that. However, it adds minimal behavioral context beyond the optional status filter, such as pagination, error behavior, or handling of missing customers. The description is not contradictory but provides little extra disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that is front-loaded with the core purpose and includes the optional filter. There is no redundant or extraneous wording, making it efficient and easily parsed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward list operation, the description covers the essential purpose but omits potentially useful details like pagination, ordering, or default behavior when no invoices exist. Given there is no output schema, the description could reasonably include more context, but it is not critically incomplete for basic usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description is the sole source for parameter meaning. It clarifies customerId as the customer identifier and status as an optional filter, but does not detail accepted status values, requiredness (customerId is required per schema), or behavior when omitted. This provides basic semantics but lacks depth.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Get), the resource (invoices), and the scoping (for a specific customer), with an optional status filter. This distinguishes it from sibling tools like get_invoices (which likely retrieves all invoices) and get_invoice (single invoice), making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for retrieving invoices tied to a specific customer, but it does not explicitly contrast it with alternatives like get_invoices or get_invoice. There is no explicit 'when to use' or 'when not to use' guidance, though the phrasing suggests the customer-scoped context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customersC
Read-only

List customers, with optional pagination and search.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
limitNo
searchNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already declares the operation is read-only, and the description's 'List' is consistent. However, the description adds no further behavioral context such as return format, pagination defaults, or error handling. With the annotation covering safety, the description contributes little beyond the schema's parameter names.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to stating the primary action and optional capabilities, making it exemplary in conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With three parameters, no output schema, and no required fields, the description is too sparse. It omits details about the response structure (e.g., whether it returns an array, pagination metadata) and does not explain search semantics. For a list tool, this is inadequate for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the parameters. It only mentions 'optional pagination and search' without defining what 'page', 'limit', or 'search' mean or how they interact. This is insufficient for an agent to construct valid calls with confidence.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('List') and resource ('customers'), and implies a plural collection operation, which distinguishes it from siblings like 'get_customer' (singular) and 'get_users' (different resource). It is unambiguous in its core function, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as 'get_customer' for a single record or 'get_customers' vs other list tools. The description only states what it does, leaving usage context entirely to the agent's inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dashboard_statsA
Read-only

Get dashboard summary stats (totals due, paid, overdue, etc.).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation readOnlyHint=true already signals a safe read operation, so the description does not need to restate that. The description adds a little context by naming the kind of stats returned, but it does not disclose aggregation scope, time range, or whether the stats are global or user-scoped. With the annotation covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the verb and resource, then gives concrete examples of the stats. Every word earns its place; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool with a readOnlyHint annotation, the description is nearly complete. The only gap is that it does not specify whether the stats are global or scoped to a user/workspace, which could matter for an agent deciding between this and get_invoices. Still, the tool is simple enough that the description suffices.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics burden on the description. The schema is trivially complete at 100% coverage. The description's mention of 'totals due, paid, overdue' gives the agent a sense of what the returned stats mean, which is useful context beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('dashboard summary stats') and lists example contents ('totals due, paid, overdue, etc.'). It is clear what the tool does, though it does not explicitly distinguish it from sibling tools like get_invoices or get_customer_invoices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is the tool for high-level dashboard summary numbers rather than detailed invoice lists, but it does not explicitly state when to use it versus alternatives. An agent can infer usage from the phrase 'dashboard summary stats', but there is no direct comparison or exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_estimateB
Read-only

Get full details for a single estimate by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
estimateIdYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint: true, which covers the read-only nature. The description adds no further behavioral context (e.g., what 'full details' includes, error behavior, or authorization requirements). It simply restates the action without enriching beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that leads with the action ('Get full details') and the key differentiator ('by ID'). It contains no unnecessary words or repetition, making it highly efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read operation with one parameter and a readOnlyHint annotation, the description is minimally adequate. It tells the agent what the tool does and the required identifier. However, it omits details about the return format (no output schema) and any error cases, which a more thorough description could provide.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the single parameter estimateId is only a number. The description merely says 'by ID' without explaining what the ID represents, its format, or any constraints beyond the schema's 'required' flag. It provides minimal added meaning, failing to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'Get' with a clear resource ('single estimate') and scope ('by ID'). It distinguishes itself from the sibling tool 'get_estimates' (plural) which likely lists estimates, making it easy for an agent to select the correct tool for a single-record fetch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (when you have an estimate ID) but does not explicitly contrast with get_estimates for listing or mention any preconditions. It offers no 'when-not-to-use' guidance, so it relies on the agent to infer the appropriate context from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_estimatesA
Read-only

List estimates, with optional pagination, search and customer filters.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
limitNo
searchNo
customer_idNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already declares this is a safe read operation, and the description's 'List' wording matches that. It adds useful context about optional pagination and filters, but does not disclose defaults, result ordering, pagination limits, or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no fluff. Every element—resource, operation, and optional modifiers—earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with four optional parameters, the description is adequate but not fully complete. It does not describe the return format, pagination defaults, or how the filters behave, and there is no output schema to fill that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the four undocumented parameters. It groups page/limit as 'pagination', search as 'search', and customer_id as 'customer filters', adding high-level meaning beyond raw types, but it omits details like page numbering, limit constraints, and search scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'List' with the resource 'estimates', making the operation unambiguous. The plural resource and optional pagination/filter phrasing clearly differentiate it from singular siblings like get_estimate and from other resources like get_invoices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for fetching a list of estimates with optional filters, but it never explicitly says when to prefer it over get_estimate or get_customer_invoices. Usage context is implied by 'List' and the filter parameters, not stated directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_invoiceA
Read-only

Get full details for a single invoice by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
invoiceIdYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already signals this is a safe read operation, so the description doesn't need to repeat that. The description adds 'full details' as return context but does not disclose behavior for missing IDs, permissions, or response structure. This is adequate for a simple read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to the tool's purpose and usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only lookup with one parameter and a readOnlyHint annotation, the description is largely complete. 'Full details' implies the response contains comprehensive invoice information. It could be slightly improved by noting not-found behavior, but nothing essential is missing for selecting and invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description bears the burden of explaining the parameter. The phrase 'by ID' clarifies that invoiceId is the invoice identifier, but it adds no detail about value format, requiredness, or error conditions. The single obvious parameter makes this minimally sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Get'), a specific resource ('full details for a single invoice'), and the key identifying scope ('by ID'). It is immediately distinguishable from plural list siblings like get_invoices and get_customer_invoices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies usage when a specific invoice ID is known and the agent needs one invoice's complete data. It doesn't explicitly name alternatives like get_invoices for listing, but the singular 'by ID' phrasing provides enough contextual direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_invoicesA
Read-only

List invoices, with optional pagination, search, status and customer filters.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
limitNo
searchNo
statusNo
customer_idNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already covers the safety profile, and the description adds 'optional pagination' as a behavioral trait. However, it does not describe pagination defaults, return shape, or any operational constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the core action, with every phrase earning its place by naming the available filters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list operation with five optional parameters, the description covers all parameter categories and the read-only behavior is covered by annotations. It lacks explicit return-format or pagination-default details, but these are not critical for a simple list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description compensates by mapping all five parameters to meaningful groups: pagination (page/limit), search, status, and customer. It adds purpose beyond the raw property types, though it does not define exactly what 'search' matches.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource, 'List invoices', and enumerates the key filter dimensions. It does not explicitly distinguish itself from get_invoice or get_customer_invoices, leaving the all-customer scope implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'List invoices' directly signals the general listing use case, and the optional filters tell an agent how to narrow results. It does not mention when not to use it or direct the agent to get_customer_invoices for a customer-scoped list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_userA
Read-only

Get full details for a single user by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIdYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint: true, so the read-only nature is known. The description adds minimal extra behavioral context—only 'full details,' which implies a comprehensive response. It does not disclose error behavior, field selection, or relationship to other resources, but with the annotation covering safety, the added value is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence, front-loaded with the verb and resource, and no filler. It communicates the essential information in under 10 words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-ID tool with one parameter and no output schema, the description is adequate. It states the action and scope. It doesn't list returned fields, but 'full details' implies completeness. The readOnlyHint is already provided by annotations. No critical missing information for invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must clarify the userId parameter. The phrase 'by ID' directly explains that userId is the unique identifier, which adds meaning beyond the bare schema (type: number). This is sufficient for a single-parameter tool, though it could be more explicit about what kind of ID (e.g., user ID in the system).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a resource ('full details for a single user'), and the key discriminator ('by ID'). This clearly distinguishes it from the plural sibling get_users, and matches the singular get_invoice/get_customer pattern in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this when you need full details for a single user identified by ID. It does not explicitly mention alternatives or when not to use it, but the distinction from get_users (plural) is implicit and obvious. Absence of exclusions is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_usersA
Read-only

List users on the InvoiceShelf instance.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
limitNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds the scope ('on the InvoiceShelf instance') but does not disclose pagination behavior, default limits, or whether the response is a paginated list. With annotations covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded with the verb and resource. No wasted words, and the scope is stated efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with readOnlyHint=true and two self-explanatory parameters, the description is mostly adequate. However, it does not mention pagination behavior or response format, and with no output schema, an agent might not know what to expect. The lack of usage guidance relative to get_user is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden for parameter meaning. The description does not explain 'page' or 'limit' beyond their names, which are fairly self-explanatory for a list operation. It adds no detail about defaults, maximum values, or pagination format, so it only partially compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('users on the InvoiceShelf instance'), clearly distinguishing it from sibling tools like get_user (singular) and get_customers (different resource). It lacks explicit differentiation from get_user, but the plural form and instance scope make the purpose clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a list operation with no filtering details, and the sibling set includes get_user for singular retrieval. However, it does not explicitly state when to use this tool versus get_user or how pagination should be handled. The context is clear enough for a basic list call but lacks explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_estimateA

Email an estimate to its customer. Looks up the sender address and the customer's email automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNo
subjectNo
estimateIdYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose idempotentHint=false and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context by stating that the sender address and customer's email are resolved automatically, which helps the agent understand that no such fields need to be provided. It does not mention failure modes or success confirmation, but the bar is lowered because of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly written sentence with no filler. The core action and the notable automatic-lookup behavior are both front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple send tool, the description covers the essential action and the main hidden behavior (automatic address lookup), and the annotations provide the safety context. Yet with no output schema, the agent gets no guidance on return values, and the optional body/subject parameters are not contextualized, leaving minor but real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description should compensate, but it does not explain body, subject, or estimateId in any meaningful way. The phrase 'Email an estimate to its customer' only hints that estimateId identifies the estimate; the optionality and purpose of body and subject are left entirely to the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Email'), a specific resource ('an estimate'), and the target recipient ('its customer'), making the action unambiguous. It does not explicitly distinguish itself from the sibling send_invoice, though the resource term 'estimate' implicitly separates them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when an estimate needs to be emailed to its customer. It does not, however, provide any exclusions or contrast with sibling tools such as send_invoice or convert_estimate_to_invoice, leaving routing decisions to the agent's inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_invoiceA

Email an invoice to its customer. Looks up the sender address and the customer's email automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNo
subjectNo
invoiceIdYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavior beyond the annotations by disclosing that sender address and customer email are resolved automatically. But it does not mention consequences of sending, such as whether the invoice is marked as sent or whether repeated sends are possible, despite idempotentHint being false and no readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, information-dense sentences. The first delivers the core action, and the second adds the key lookup behavior with no filler or redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a side-effecting external communication tool with no output schema and minimal annotations, yet the description does not explain success/failure behavior, what happens after sending, or error cases like a missing customer email. It is enough to select the tool but not to invoke it confidently in all realistic situations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it does not explain the optional body and subject parameters or their defaults. It only implicitly clarifies that invoiceId identifies the invoice and that recipient addresses are auto-resolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Email an invoice to its customer.' This clearly distinguishes it from sibling tools like send_estimate and update_invoice, and the added detail about automatic address lookup sharpens what the tool uniquely does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes it clear this tool is for emailing invoices, which implies the appropriate context. However, it does not explicitly state when not to use it or mention alternatives such as send_estimate, nor does it note preconditions like the invoice being finalized or the customer having a valid email.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test_connectionA
Read-only

Verify the API base URL and token work by calling /me.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already indicates this is a safe read operation. The description adds that the tool calls the /me endpoint, disclosing the concrete mechanism and reinforcing that no mutation occurs. It provides useful behavioral context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, direct sentence conveys both the purpose and the implementation. It is front-loaded with the action and resource, with no filler or redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only connectivity check, the description is sufficiently complete. It does not describe response formats or error behavior, but given the tool's simplicity and the readOnlyHint annotation, this is a minor gap rather than a critical omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline of 4 applies. The description appropriately focuses on the tool's behavior rather than parameter details, which are nonexistent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Verify') and a specific resource ('API base URL and token') while naming the exact endpoint used ('/me'). This distinguishes it from data-fetching siblings like get_user or get_users, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this tool to validate that the API base URL and token are configured correctly. It does not explicitly name alternatives or exclusion criteria, but its role as a connection check is evident from the wording and the sibling tool list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_customerB
Idempotent

Update fields on an existing customer.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
emailNo
phoneNo
customerIdYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description 'Update fields' is consistent with these annotations and adds no extra behavioral context, such as whether it does a partial update or what happens to omitted fields. Since annotations carry the burden, a baseline 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no filler. It front-loads the core action and resource, and every word earns its place. There is no unnecessary detail or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and 0% schema description coverage, the description is too thin. It does not explain partial-update semantics, error behavior, or what happens to fields not provided. While the required parameter is in the schema, an agent would benefit from explicit guidance on how the update behaves. The description leaves significant gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description mentions no parameters at all. The property names (name, email, phone, customerId) are self-explanatory, but the description does not clarify that only customerId is required or that the other fields are optional updates. The description adds no value beyond the schema, failing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Update' and the resource 'customer', specifying that it modifies an existing customer. This distinguishes it from create_customer and delete_customer, though it doesn't explicitly contrast with those siblings. The resource is unambiguous, so an agent can determine the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like create_customer or delete_customer. The description implies it is for existing customers, but it does not explicitly state when to choose update over create, nor does it mention any prerequisites or limitations. An agent would have to infer usage from the sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_estimateB
Idempotent

Update fields on an existing estimate.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsNo
notesNo
estimateIdYes
expiry_dateNo
estimate_dateNo
template_nameNo
reference_numberNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare idempotentHint=true and destructiveHint=false, which covers the safety profile. The description adds that this updates fields rather than performing some other action, but it does not clarify whether the update is partial or full replacement, whether totals are recalculated, or what side effects occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The sentence is short, clear, and front-loaded with the core verb and resource. It contains no filler, though it is so concise that it omits useful detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 inputs, no output schema, and zero parameter descriptions, a single generic sentence is inadequate. An agent would not know update merging semantics, accepted date formats, whether items replace the full list, or what a successful response looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description names none of the 7 parameters. It does not even mention that estimateId is required or that all other fields are optional. The description adds no meaning beyond the raw schema property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update') and a clear resource ('existing estimate'), which distinguishes it from siblings like create_estimate, delete_estimate, send_estimate, and convert_estimate_to_invoice. The high-level operation is immediately obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing estimate' conveys that this tool is for modifying an already-created estimate, providing clear context and differentiating it from creation and deletion tools. However, it does not explicitly name alternatives or state when not to use this tool, so it falls short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_invoiceC
Idempotent

Update fields on an existing invoice.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsNo
notesNo
due_dateNo
invoiceIdYes
invoice_dateNo
template_nameNo
reference_numberNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond the word 'update' – it doesn't disclose whether this is a partial update (only specified fields changed) or a full replacement, nor does it mention that unspecified fields remain untouched. This is a significant behavioral gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise and front-loaded with the core action. However, it is under-specified; it doesn't provide any of the details an agent needs. While it is not verbose, it sacrifices usefulness for brevity, making it merely adequate in structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, one required, and no output schema, the description should provide more context. It doesn't mention whether updating is partial, what happens to missing fields, error conditions, or any preconditions. The schema alone is insufficient, and the description does not fill the gap, making the tool incomplete for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for most properties (only 'price' has a description). The description does not compensate by explaining any parameter meaning – it only says 'Update fields' without listing or elaborating on any of the 7 parameters. The agent must rely on the bare schema names, which is insufficient for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Update') and a specific resource ('existing invoice'), making the core action clear. However, it doesn't differentiate itself from sibling tools like update_estimate beyond the resource name, which is obvious from the name itself. It also doesn't specify which fields can be updated, leaving the agent to infer from the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It doesn't mention that an invoice must already exist (require invoiceId) or that you should fetch the invoice first. No exclusions or alternative tool recommendations are provided, leaving the agent to guess when this is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.1.0
    • Changedcreate_invoice1 field changed
      • addedInput schema / properties / tax_percent
        Added value: +{
        +  "description": "Adds the existing tax type with this percentage (e.g. 19) on the whole invoice. Omit for no tax.",
        +  "type": "number"
        +}
  2. 23 tool updatesv1.0.0
    • First observedconvert_estimate_to_invoice
    • First observedcreate_customer
    • First observedcreate_estimate
    • First observedcreate_invoice
    • First observeddelete_customer
    • First observeddelete_estimate
    • First observeddelete_invoice
    • First observedget_customer
    • First observedget_customer_invoices
    • First observedget_customers
    • First observedget_dashboard_stats
    • First observedget_estimate
    • First observedget_estimates
    • First observedget_invoice
    • First observedget_invoices
    • First observedget_user
    • First observedget_users
    • First observedsend_estimate
    • First observedsend_invoice
    • First observedtest_connection
    • First observedupdate_customer
    • First observedupdate_estimate
    • First observedupdate_invoice

TDQS

B3.3/5.0

Scored across 23 tools

Disambiguation4/5

Most tools map cleanly to a specific resource and action (get_invoices vs get_invoice vs get_customer are clear). The only real overlap is get_customer_invoices, which largely duplicates get_invoices with a customer filter, though the descriptions still make the intent understandable.

Naming Consistency5/5

Tool names consistently follow verb_noun snake_case (get_, create_, update_, delete_, send_, convert_). Singular/plural distinctions are used predictably for list vs detail endpoints, and compound actions like convert_estimate_to_invoice follow the same pattern.

Tool Count3/5

23 tools is on the heavy side, within the 16-25 borderline range. The count is justified by full CRUD for invoices, estimates, and customers plus send/convert/dashboard tools, but it is more than a typical focused MCP server.

Completeness3/5

The core invoice, estimate, and customer lifecycle has solid create/read/update/delete coverage plus send and convert actions. However, there are notable gaps for an invoicing domain: no payment recording, invoice line-item management, or user administration beyond read-only lookup.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Production-grade MCP server for FreshBooks. 25 tools for invoices, clients, expenses, payments, time tracking, projects, estimates, and financial reports. OAuth2 with automatic token refresh.
    25
    3
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides 55 tools for managing QuickBooks entities like customers, invoices, and bills via any MCP-compatible client, built on Cloudflare Workers with OAuth 2.0 authentication.
    8
    Apache 2.0
  • F
    license
    B
    quality
    D
    maintenance
    Bridges to a simple_invoicing FastAPI backend, exposing invoice, product, ledger, inventory, buyer, and payment management as MCP tools for use with MCP-compatible clients.
    15
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Comprehensive MCP server for Wave Accounting, providing 45+ tools across invoicing, customers, products, transactions, bills, estimates, taxes, and financial reporting, plus 17 pre-built UI workflows.
    5
    -