BillingEngine MCP server
OfficialYou can manage BillingEngine customers, draft invoices, and payments through an MCP client.
Search, create, and update customers.
Create and update draft invoices with line items, dates, due days, and service periods.
List, filter, and get invoices by status, customer, number, or date.
Download invoice PDFs or XML e-invoices to your computer.
Record payments for sent open invoices and list payments.
Limitations: invoices are created as drafts, nothing can be deleted, and sending invoices is done in BillingEngine. Requires a paid plan and API key.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@BillingEngine MCP serverCreate a draft invoice for Malermeister Müller for 100 euros."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
BillingEngine MCP server
Create customers and invoices, look up open invoices and record payments in BillingEngine by talking to Claude, Cursor or any other MCP client:
"Create an invoice for Acme Painting over 100 euros."
"Which invoices are overdue?"
"Acme paid invoice 2610030028, record the payment."
BillingEngine is invoicing software for freelancers and small businesses with e-invoicing (EN 16931, XRechnung, ZUGFeRD / Factur-X). The server is a thin layer over the BillingEngine API and runs on your computer.
Requirements
Node.js 20 or newer
A BillingEngine account on a paid plan. The API is not available in the Free plan.
Your API key from Settings > API in BillingEngine
Related MCP server: bill4time-mcp
Setup
Claude Desktop
Add this to claude_desktop_config.json:
{
"mcpServers": {
"billingengine": {
"command": "npx",
"args": ["-y", "billingengine-mcp"],
"env": { "BILLINGENGINE_API_KEY": "your-api-key" }
}
}
}Claude Code
claude mcp add billingengine --env BILLINGENGINE_API_KEY=your-api-key -- npx -y billingengine-mcpCursor and other clients
Use the same command, args and env as above in the MCP settings of your client.
Tools
Tool | What it does |
| Searches customers by company, name or email |
| Creates a customer |
| Changes a customer |
| Creates a draft invoice for a customer |
| Changes a draft invoice |
| Lists invoices, filtered by status, customer, number or date |
| Returns one invoice with items and payment status |
| Saves the PDF or the XML e-invoice to your Downloads folder |
| Records a payment for sent invoices |
| Lists payments, optionally for one invoice |
What it does and does not do
Invoices are created as drafts. Nothing is sent to your customers; you review and send the invoice in BillingEngine. Every result contains a link to open the draft.
Prices are net. Number, language, texts, VAT treatment and payment terms come from your account settings, exactly as in the invoice form.
Nothing can be deleted through this server.
New invoices count towards the monthly invoice maximum of your plan.
The API allows 60 requests per minute per key.
Your key stays on your computer and is only sent to BillingEngine.
If a key is ever exposed, generate a new one under Settings > API. The old key stops working immediately. The settings also show when the key was last used.
Configuration
Variable | Meaning |
| Your API key (required) |
| Base URL, defaults to |
Development
npm install
npm run build
BILLINGENGINE_API_KEY=... node dist/index.jsLicense
MIT. BillingEngine itself is a commercial service; see billingengine.com.
Available Tools
10 toolscreate_customerCreate a customerA
Creates a customer. Needs a company or a last name and a full address (line 1, postal code, city, country).
| Name | Required | Description | Default |
|---|---|---|---|
| No | |||
| phone | No | ||
| address | Yes | ||
| company | No | ||
| last_name | No | ||
| first_name | No | ||
| vat_number | No | ||
| buyer_reference | No | Leitweg-ID or buyer reference for e-invoices |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, establishing this as a non-destructive write. The description adds the useful required-field constraint ('Needs a company or a last name and a full address'), which is genuine behavioral context. However, it omits what happens on duplicate detection, what the response contains, or any rate/permission requirements – gaps that matter for a creation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, no wasted words, and the required-field constraint is front-loaded after the purpose statement. Efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be described. But with 8 parameters at 13% schema coverage, the description only covers a subset of semantics, leaving the agent to guess the meaning of several optional fields. A creation tool of this complexity warrants more parameter and behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 13%, so the description must carry weight. It does explain the key business rules: an address requires line 1, postal code, city, and country, and either company or last name is needed. But with 8 parameters, several (email, phone, vat_number, buyer_reference, first_name) have no explanation in either schema or description. Partial compensation, not full.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a customer') that clearly distinguishes it from the sibling update_customer and find_customers. An agent can identify this as the record-creation tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives like find_customers (to check for duplicates first) or update_customer. The description gives no context about preconditions or when this tool should be selected over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_invoiceCreate a draft invoiceA
Creates a draft invoice for a customer. Number, language, texts and VAT treatment come from the account settings. The invoice is not sent; it counts towards the monthly invoice maximum of the plan.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Invoice date, defaults to today | |
| items | Yes | ||
| due_days | No | Days until payment is due | |
| signature | No | ||
| customer_id | Yes | ||
| introduction | No | ||
| service_period_end_date | No | ||
| service_period_start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations confirm it is a write operation (readOnlyHint=false) but do not mention side effects like counting towards a monthly maximum. The description adds useful behavioral context: number/language/texts/VAT come from account settings, the invoice is not sent, and it counts toward the plan's invoice limit. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the core action, then scoping details. Every sentence carries useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers high-level behavior and side effects, and an output schema exists (so return values need not be explained). However, with 8 parameters and only 25% schema coverage, the description omits parameter guidance entirely, leaving gaps for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate. It does not describe any parameters—not even the required customer_id or items. The schema has some descriptions, but the description adds no meaning beyond what little is there.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Creates a draft invoice') and clarifies that the invoice is not sent. It implicitly distinguishes from update_invoice and record_payment by calling it a 'draft invoice', but does not explicitly name siblings or differentiate beyond that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (creating a draft, not sending, counting towards limit) but provides no explicit when-to-use vs alternatives guidance. It does not mention when to use create_invoice instead of update_invoice or how it relates to sending an invoice later.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_invoiceDownload an invoice fileAIdempotent
Saves the PDF or the XML e-invoice of an invoice on this computer and returns the file path. Defaults to the Downloads folder.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| format | No | ||
| directory | No | Target folder, defaults to ~/Downloads |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuine context: the side effect is writing a file to local disk, the call returns a file path, and the default destination is the Downloads folder. It does not, however, mention overwrite behavior or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and outcome, with the default-location detail appended. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema and annotations cover the return value and safety profile, so those need no description. The gaps are the undocumented 'id' parameter and the absence of any routing guidance among the many sibling invoice tools, leaving the definition only minimally complete for a 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 33% (only 'directory' is documented). The description partially compensates by clarifying the pdf/xml choices and the Downloads default, matching both the format enum and directory default, but it says nothing about the required 'id' parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (saves) and resource (PDF/XML e-invoice) and clarifies the outcome (file on this computer plus returned path). This implicitly separates it from get_invoice, which retrieves invoice data rather than writing a file, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no prerequisites, and no direction on choosing between this and get_invoice or list_invoices. The context is inferable but the definition never states it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_customersFind customersARead-only
Lists customers, optionally filtered by a search text that matches company, names and email.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of the list, starting at 1 | |
| query | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the match semantics of the filter but says nothing about pagination behavior, result limits, or ordering that an agent might need when listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the filter behavior follows immediately after the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the return shape need not be explained, and annotations cover safety. The remaining gap is that paging limits/ordering are left entirely to the schema, but for a simple two-param list tool this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% – 'page' is documented in the schema but 'query' has no description. The description compensates by explaining what 'query' matches (company, names, email), which is meaningful added value beyond the bare string type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists customers') plus the scope of filtering ('search text that matches company, names and email'). The read/list verb inherently distinguishes it from the create_customer, update_customer, and record_payment siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'optionally' implies filtering is a choice, but there is no explicit guidance on when to use this versus other tools or on prerequisites. Usage is only inferable from the self-evident purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_invoiceGet an invoiceBRead-only
Returns one invoice with its items, amounts and payment status.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered structurally. The description enumerates return contents (items, amounts, payment status), but an output schema exists that already conveys this, so the added behavioral value is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though it spends words restating return contents that the output schema already exposes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema and annotations present, the bar is lower, and the description is adequate for a simple lookup. It omits any mention of error behavior (e.g. nonexistent id) and leaves the id parameter entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required 'id' parameter, so the description must compensate and does not — it never states that id identifies the invoice to fetch or what form it takes. The reference to 'one invoice' hints at lookup but adds no real semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (get) plus resource (invoice), and the phrase 'one invoice' distinguishes it from the list_invoices sibling. It does not, however, differentiate itself from download_invoice, which also operates on a single invoice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named. With siblings like list_invoices, create_invoice, update_invoice and download_invoice in the set, the agent gets no help deciding between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_invoicesList invoicesBRead-only
Lists invoices, newest first. Status: draft, open (sent and unpaid), overdue (open and past due) or paid.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of the list, starting at 1 | |
| after | No | Invoice date from | |
| before | No | Invoice date until | |
| number | No | ||
| status | No | ||
| per_page | No | ||
| customer_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=true, openWorldHint=false), so the bar is lower. The description still adds real context: the newest-first ordering and, importantly, disambiguates the status enum by defining open as 'sent and unpaid' and overdue as 'open and past due', which the schema enum does not explain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core purpose front-loaded and no filler. It is efficient, though the terseness leaves room to add filtering/pagination context that an agent would need.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations carry the safety profile. However, for a 7-parameter listing tool the description omits any mention of date filters, customer filtering, or pagination behavior, which are central to correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43% (only page, after, before are documented), so the description should compensate more. It clarifies the status values, which is genuinely useful, but says nothing about number, per_page, or customer_id filtering, leaving those parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists invoices') and adds the scope detail of ordering ('newest first'). An agent can distinguish it from the singular get_invoice and from list_payments, though it does not explicitly name a sibling to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of alternatives like get_invoice for a single invoice or find_customers, and no stated exclusions or prerequisites. Usage is only implied by the verb 'Lists'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_paymentsList paymentsARead-only
Lists recorded payments, newest first, optionally only those for one invoice.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of the list, starting at 1 | |
| invoice_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds one genuine behavioral trait beyond the schema: results are returned newest first. However, it says nothing about pagination behavior despite a page parameter, so it is adequate rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, followed by sort order and the optional filter. No filler and nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and annotations cover safety. The description supplies sort order and filter semantics; only pagination behavior is left implicit, which is a minor gap for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (invoice_id has no description), and the description compensates by explaining that invoice_id restricts results to a single invoice's payments. The page parameter is already documented in the schema, so the description covers the actual gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists recorded payments'), plus scope details (newest first, optional invoice filter). It is distinguishable from siblings like list_invoices and record_payment by resource name, though it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the filter applies ('optionally only those for one invoice') but gives no explicit guidance on when to use this tool versus list_invoices or other siblings, and no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_paymentRecord a paymentA
Records a payment for one or more sent, open invoices of the same customer and tax rate. The amount is gross.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Defaults to today | |
| amount | Yes | Gross amount received | |
| invoice_ids | Yes | ||
| payment_method | No | Defaults to transfer |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-readonly, non-destructive, closed-world mutation, so the safety profile is covered. The description adds the gross-amount convention and the same-customer/same-tax-rate eligibility rule, but says nothing about what happens to the invoices afterward (state change, partial payment handling) or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action and the gross-amount convention front-loaded. Every clause carries information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and annotations cover the mutation safety profile. The description supplies the operation, its scope limits, and the amount convention - enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the description still adds meaning beyond it: the eligibility constraints on invoice_ids (sent, open, same customer and tax rate) are not captured in the schema, and "amount is gross" reinforces the gross-amount semantics. Date and payment_method defaults are left to the schema, which documents them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Records a payment") and narrows scope precisely: payments against one or more sent, open invoices belonging to the same customer and tax rate. No sibling tool creates payments, so the operation is unambiguous in context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description embeds clear preconditions for use - invoices must be sent and open, and must share the same customer and tax rate - which tells the agent when this call is valid. It does not name an alternative tool or state when-not to use it, but no competing payment-creation sibling exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_customerUpdate a customerAIdempotent
Changes the given fields of a customer; fields that are left out stay as they are.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| No | |||
| phone | No | ||
| address | No | ||
| company | No | ||
| last_name | No | ||
| first_name | No | ||
| vat_number | No | ||
| buyer_reference | No | Leitweg-ID or buyer reference for e-invoices |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation as non-readOnly, non-destructive and idempotent, so the safety profile is covered. The description adds genuine behavioral context beyond that: omitted fields are preserved unchanged, clarifying the merge/patch semantics. It does not describe error or permission behavior, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the verb and resource before the partial-update qualifier. Every clause earns its place with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema means return values need not be explained, and the annotations cover the safety profile. But with 9 parameters at 11% schema coverage and a nested address object, the near-total absence of parameter semantics leaves a meaningful gap for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 11%, so the description carries most of the burden for parameter meaning, and it names no fields at all. The generic phrase 'given fields' adds nothing about id, email, phone, the nested address object, or their formats. The low-coverage gap is not compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Changes ... a customer'. An agent can distinguish it from create_customer and find_customers by the modification verb, though no sibling is named explicitly. Clear but without explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb (modify an existing customer), and the partial-update note signals this is a field-level edit rather than a full replacement. However, there is no explicit guidance on when to use this versus create_customer or find_customers, and no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_invoiceUpdate a draft invoiceAIdempotent
Changes a draft invoice. Passing items replaces all existing items, so send the complete list. Sent invoices cannot be changed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| date | No | Invoice date, defaults to today | |
| items | No | ||
| due_days | No | Days until payment is due | |
| signature | No | ||
| introduction | No | ||
| service_period_end_date | No | ||
| service_period_start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnlyHint=false, destructiveHint=false, idempotentHint=true), and the description adds genuinely non-derivable behavior: items replacement overwrites the full list and only drafts are mutable. This is real context beyond the annotations, though it omits what happens to omitted optional fields and error behavior on non-draft invoices.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the highest-risk gotcha (items replacement), then the precondition. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description covers the mutation scope and the two main hazards. Partial-update semantics for omitted fields and the failure mode for a non-draft id are left implicit, which is the only meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25% across 8 parameters, so the description must add meaning. It does add the critical items semantics ('passing items replaces all existing items, so send the complete list'), but says nothing about date, due_days, signature, introduction, or the service period fields, leaving half the surface undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Changes a draft invoice'), with the scope qualifier 'draft' that separates it from create_invoice and the read-only invoice siblings. It does not explicitly name a sibling to route away from, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-not condition: 'Sent invoices cannot be changed,' which tells the agent the tool only applies to drafts. It doesn't name an alternative for the sent-invoice case, but no obvious sibling covers that path, so this is a strong but not exhaustive usage statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
create_customer - First observed
create_invoice - First observed
download_invoice - First observed
find_customers - First observed
get_invoice - First observed
list_invoices - First observed
list_payments - First observed
record_payment - First observed
update_customer - First observed
update_invoice
TDQS
Scored across 10 tools
Each tool targets a clearly distinct resource and action: customers (find/create/update), invoices (create/update/list/get/download), and payments (record/list). There is no meaningful overlap between tools, and an agent can easily select the right one based on the requested resource and operation.
All tools use snake_case with a verb_noun pattern, which is highly consistent. The only minor deviation is using 'find_customers' for listing customers while invoices and payments use 'list_*', but this does not create confusion.
With 10 tools, the set is well-scoped for a billing engine, covering the core customer, invoice, and payment operations without excessive bloat. Each tool has a clear place in the domain.
The surface covers create/read/update for customers and invoices plus payment recording, but it lacks a tool to send a draft invoice, which is a critical step because record_payment only works on sent/open invoices. Deletion operations for any entity are also absent, creating notable workflow gaps.
Maintenance
Related MCP Connectors
API-first CRM for LLMs - contacts, companies, deals and activities over a native MCP server.
Invoicing you drive by talking to your AI: log time, raise invoices and track what's owed via MCP.
Create and manage invoices and customers on Jupiter Invoice (MCP, API-key auth).
Hosted MCP server for Mini Accountant: invoices, expenses, customers, analytics, tax estimates.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server for the Billingo invoicing API (v3) that enables AI assistants like Claude to manage invoices, partners, products, expenses, and more through natural language.4212MIT
- AlicenseCqualityAmaintenanceMCP server for Bill4Time providing API coverage for legal billing and time tracking, enabling natural language interaction with clients, projects, time entries, invoices, payments, and more from Claude Desktop.59125 PyPI1MIT

makeleaps-mcpofficial
AlicenseAqualityCmaintenanceUnofficial MCP server to operate MakeLeaps clients, quotes, and invoices from LLMs via the MakeLeaps API, with local execution and no telemetry.8MIT- FlicenseNot gradedqualityDmaintenanceHosted MCP server that enables AI assistants to manage clients, invoices, and expenses via the Invox API, supporting actions like drafting, sending, cancelling, and marking invoices as paid, as well as logging expenses and updating client information.-