Billy MCP
Billy MCP is a local MCP server that provides read-only accounting insights and controlled write operations for Billy bookkeeping, including receipt management, vendor tracking, proposal-based execution, and batch handling.
Read company status, write switches, and receipt inbox (
billy_status).List and fetch typed Billy resources with filters and pagination (
billy_list,billy_get).Get period overview of unreconciled bank lines, postings, and receipts (
billy_period_overview).Run trial balance and profit/loss reports using live chart data (
billy_trial_balance,billy_profit_loss).View outstanding unpaid documents and period expense totals (
billy_outstanding,billy_period_expenses).Import and archive receipts (PDF/PNG/JPEG) with deduplication and provenance (
billy_import_receipt,billy_receipts).Save and list vendor billing portal info (
billy_save_vendor,billy_vendors).Prepare concrete write proposals (create bills, journals, payments, approve, sales invoices, credit notes, contacts, upload receipts, reconcile) (
billy_prepare).Inspect, refresh, and execute saved proposals with hash verification and approval modes (
billy_plan,billy_refresh_plan,billy_execute).Prepare, inspect, refresh, and execute ordered purchase batches (1–10 cases) with staged progress (
billy_batch_prepare,billy_batch_get,billy_batch_refresh,billy_batch_execute).Review execution journal including rejected/uncertain operations (
billy_journal).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Billy MCPImport this receipt and prepare a draft bill"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Billy MCP
A local MCP server for Billy bookkeeping: complete accounting reads, receipt collection/provenance, reviewed purchase and sales drafts, credit notes, payments, reports and reconciliation of existing bank postings. Generic company support: the API token determines the company. This is an independent community project, not an official Billy or Shine product.
Run
Version 0.2.1 is distributed as an installable tarball on GitHub Releases. The scoped package is also published on npm as @pete-life/billy-mcp@0.2.1. Use npm install -g @pete-life/billy-mcp@0.2.1, the release tarball or the repository build below; see installation and agent setup.
Requires Node.js 22.13+ (built-in SQLite).
git clone https://github.com/pete-life/billy-mcp.git
cd billy-mcp
npm ci --ignore-scripts
npm run build
npm run setup
npm startCreate a company API token in Billy under Settings → Access tokens. Setup asks for it with hidden terminal input, fetches the connected company and asks you to confirm that identity. It writes an owner-readable credentials.env outside the repository, under ~/.local/share/billy-mcp/. Do not paste tokens in chat. npm start loads that file automatically.
Alternatively set BILLY_ACCESS_TOKEN in the MCP process environment. The company ID is discovered automatically. BILLY_ORGANIZATION_ID is optional: set it to enforce an expected company in addition to the token. A mismatched token is rejected. Setup records this extra guard automatically.
Each company has separate local receipt metadata, plans and vendor registry rows. Use a different BILLY_DATA_DIR for fully separate account profiles. Do not share a profile directory across machines or network filesystems; execution locking is local SQLite.
Generic MCP client configuration for this checkout:
{
"mcpServers": {
"billy": {
"command": "node",
"args": ["/absolute/path/to/billy-mcp/dist/launch.js"]
}
}
}command must resolve to Node 22.13+ in your desktop client's environment; an absolute Node path is safest. This file is a configuration example, not an automatic edit to your client settings.
Related MCP server: billy-mcp
Included tools
Tool | Purpose |
| Connected company, write switches and receipt inbox |
| Typed resource selection, supported filters, full pagination |
| Unreconciled bank lines, possible existing postings and receipts |
| Account balances and period profit/loss using the live chart |
| Current unpaid documents and period expense totals |
| Archive originals, deduplicate bytes and retain provenance |
| Vendor billing accounts and retrieval/access status |
| Preview a concrete financial operation and current record snapshots |
| Inspect or refresh an unexecuted proposal |
| Execute a reviewed proposal once, then read back the outcome |
| Inspect completed, rejected and uncertain operations |
| Preflight and inspect 1–10 ordered purchase cases |
| Refresh unstarted scope, approve once and resume guarded stages |
Read tools default to compact, sanitized responses. Lists return {records,count,complete}; individual reads return {record,complete}. Use verbose:true for additional sanitized fields. Every record is retained; compact output does not truncate rows or change internal verification snapshots.
The bookkeeping-period MCP prompt loads the bookkeeping skill. It orchestrates Gmail/Drive/local files/vendor portals through the agent's existing connectors and browser tools. Those external tools and logged-in sessions are not bundled in this MCP server. Original files are downloaded into the configured inbox, read by the agent and imported with source and extracted invoice metadata. This project does not contain a universal authenticated vendor scraper or built-in OCR.
Reusable bookkeeping skill
The bookkeeping skill contains tool recipes, exception handling and an optional learning routine. The MCP prompt loads the same file. To use it directly in a skill-aware client, install or link the skills/billy-bookkeeping directory according to that client's instructions.
Choose your own model. Store company-specific mappings, mailbox selection and billing portal access outside the public repository. External mail, file and browser tools must be supplied by the calling agent.
The learning routine records evidence privately under the configured data directory. Any reusable public documentation change must exclude company records and secrets. Learning does not authorize additional accounting actions.
Supported operations
billy_prepare accepts a validated discriminated operation:
upload_receipt: upload an archived PDF/PNG/JPEG and confirm the attachment exists.create_contact: create a customer/supplier with duplicate name/registration checks.create_bill: create a draft, using an uploaded receipt, supplier, invoice date/number, currency, explicit tax mode and account/tax-coded lines. Supplier name and extracted totals must match.create_journal: create a balanced draft, with receipt or explicit no-receipt reason plus bank-line identity. Tax-coded journal expansion is not live-verified; use a purchase bill for VAT-coded expenses.approve: approve a reviewed bill, invoice or journal draft. Bills require attached evidence.create_sales_invoice/update_draft_invoice: create a reviewed sales draft or edit its documented header fields. Invoice lines cannot be replaced after creation.send_invoice: send an approved sales invoice to one reviewed contact-person email in a separate operation.create_customer_credit_note/create_supplier_credit_note: create an original-linked, amount-limited draft credit note. Supplier credits require their original credit document.create_payment: register full/partial same-currency payments, supported supplier FX settlements and evidenced fees against an exact unused bank line. Fee-bearing payments require a base-currency cash account; ambiguous FX rounding is rejected before writing. This records money already moved, not a bank transfer.reconcile: associate one bank line's existing match with an existing bank-account posting and approve it. No expense is created. This is separately gated and live-verified for the documented single-line DKK flow.
See operation contracts and limitations for required evidence, supported currency combinations and verification status. Purchase batches execute original upload, draft, optional approval, payment and reconciliation with persistent per-stage progress.
Execution guarantees and limits
Writes are off by default. To enable them after setup/acceptance, set BILLY_ALLOW_WRITES=true in the local credentials file or process environment. This enables the capability; it does not authorize arbitrary financial actions. Default BILLY_APPROVAL_MODE=confirm requires a user-facing MCP form approving the exact single operation or complete batch. Unsupported clients, decline and cancellation fail closed. An operator can explicitly choose BILLY_APPROVAL_MODE=trusted_automation in their private local profile for authorized automation. The authorization text remains an audit note, never proof of consent. See approval modes.
Each preview has a persisted ID and hash binding the operation and the current records. Execution claims a company-wide SQLite lock, validates fresh data, writes and reads back. Repeated execution returns stored evidence. Another changed request is still another operation: callers must not use different wording/amounts to bypass an uncertain write.
No POST/PUT automatic retries. Read-only 429/temporary server errors have bounded retries and a timeout.
A crash, timeout, malformed response or failed verification after sending a write is unknown, never “failed, safe to retry”. Unknown/executing plans block subsequent writes for that company. A definitive first-write 400/401/403/404/409/422/429 rejection is retryable only after a fresh preview. Recovery of genuinely uncertain writes requires operator investigation against Billy and the interactive local command described below.
Preview expires after 30 minutes; changed data requires refreshing and reviewing the new hash.
Receipt import restricts paths to the configured inbox, rejects escaping symlinks, checks PDF/PNG/JPEG signatures and limits size to 20 MB. SHA-256 identifies the original bytes. It does not prove invoice validity or correct extraction.
Duplicate bill checks compare invoice references across contacts with matching name/registration number, plus same-date/amount candidates. They cannot prove uniqueness when existing supplier identity, references and dates are all different or missing. Inspect existing postings first; journal/payment bank-line IDs prevent local duplicate actions.
A sequence of remote API calls is not an atomic Billy transaction. If reconciliation stops after association creation, it is unknown and must be investigated.
Concurrent manual edits in Billy cannot be fully locked by this server. Fresh snapshots reduce but cannot eliminate that race.
Live verification status
Live purchase flow verified on 2026-09-22: company-token connection, original PDF upload, attachment linking, Danish VAT purchase draft, approval and independent balanced-ledger readback were verified through the live API. Supplier legal-name changes can be matched by verified country and registration number. A same-currency DKK supplier payment and single-posting bank reconciliation were also live-verified on 2026-09-22, including independent balanced-ledger and zero-balance checks; see live acceptance.
BILLY_ALLOW_BANK_MATCHING defaults to false independently of other writes. Public documentation labels match relationships read-only and does not explain the full approval sequence. The implemented single-posting sequence follows the separate documented association resource; the single-line DKK flow was subsequently live-validated; other variants remain unverified.
All four v0.2 reports completed in a GET-only live probe on 2026-09-23. The sales, credit-note, partial-FX, fee, approval and batch paths are fixture tested; they are not newly live-verified. Current exceptions include FX customer receipts and non-base-bank FX settlement, ambiguous cent allocations, grouped/split bank matches, credit-note settlement/refunds, automatic VAT filing or settlement, subscription changes, bank transfers, remote hosting, scheduling, remote batch atomicity and automatic recovery of unknown writes. Tax-coded journal expansion remains unverified.
Operator recovery
After inspecting the real Billy records and resolving any partial operation in Billy, run:
npm run recover -- PLAN_ID applied
# Or, only after establishing that no remote changes occurred:
npm run recover -- PLAN_ID not_appliedThis is an interactive local operator command, never an MCP tool. It requires an evidence note and confirmation of the plan ID, records the decision and never writes to Billy. It refuses recovery of an executing plan while the recorded executor process is alive. applied becomes a terminal state that cannot replay; not_applied allows a refreshed preview. Incorrect operator assertions are not detectable automatically, so inspect the entire operation first.
Development
npm run checkTests use temporary local files and simulated Billy responses. The stdio test starts the compiled MCP server and performs protocol initialization, tool discovery, validation errors and prompt retrieval. No test touches production accounting data.
Official references: Billy API, MCP TypeScript SDK. API documentation checked 2026-09-23.
License and contributions
MIT licensed. See LICENSE. Contributions are welcome; read CONTRIBUTING.md. For security reports, see SECURITY.md. This is an early release with deliberately limited accounting flows; review the supported operations and verification limits before enabling writes.
Available Tools
21 toolsbilly_batch_executeADestructive
Approve the complete ordered company batch once and execute its stages sequentially through guarded plans. A stop returns explicit partial progress. Resume the same hash after inspecting the cause.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | ||
| expectedHash | Yes | ||
| authorization | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the description correctly reflects a mutating operation. It adds valuable behavioral insight beyond annotations: that a stop yields explicit partial progress and that resumption uses the same hash. This clarifies failure/resume semantics without contradicting the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences front-load the primary action and then add a critical behavioral note about stops/resume. No filler or redundancy; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core operation and resume behavior, but with no output schema and incomplete parameter documentation, it lacks essential details like success indicators, error/failure formats, and what 'guarded plans' entail. It is adequate for a rough understanding but not complete for dependable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter meaning. The description only implicitly references expectedHash via 'same hash' and 'guarded plans', but gives no explanation of batchId or authorization. It fails to meaningfully compensate for the undocumented parameters, leaving an agent to guess their purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Approve... and execute') on a specific resource ('complete ordered company batch') and distinguishes it from the singular 'billy_execute' sibling by explicitly mentioning batch and sequential guarded stages. It is clear, actionable, and non-tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: it is for batch execution, and hints at a resume workflow after a stop ('Resume the same hash after inspecting the cause'). However, it does not explicitly differentiate this tool from siblings like billy_execute or billy_batch_prepare, nor does it state when NOT to use it. The guidance is present but limited.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_batch_getBRead-onlyIdempotent
Inspect saved batch, child plan IDs, partial progress and stop reason.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering side effects. The description adds value by specifying the data aspects returned (child plan IDs, partial progress, stop reason) but does not disclose error handling or response format, which is acceptable given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no filler. It front-loads the verb 'Inspect' and packs key specifics, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with one parameter and no output schema, the description covers the core purpose and data returned. However, it omits guidance on when to use it relative to siblings and does not describe response shape or error behavior, leaving notable gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explain the batchId parameter beyond its self-evident name. It implies the batchId identifies the saved batch but provides no additional details about semantics, constraints, or usage beyond the schema pattern.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Inspect' and identifies the resource ('saved batch') plus distinctive details (child plan IDs, partial progress, stop reason). It clearly differentiates from sibling tools like billy_batch_execute or billy_batch_refresh by emphasizing inspection, though it doesn't explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as billy_status or billy_get. The description only states what the tool does without contextual cues for selection or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_batch_prepareADestructive
Preflight and save an ordered purchase batch (1-10 cases). Each case binds an original receipt, exact draft lines and explicit approval/payment/reconciliation stages. No Billy writes.
| Name | Required | Description | Default |
|---|---|---|---|
| cases | Yes | ||
| reason | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=false and destructiveHint=true. The description adds the key context that despite the batch being saved, no writes occur in Billy, and that each case passes through explicit approval/payment/reconciliation stages. This complements rather than contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the core purpose is front-loaded and the behavioral caveat is in the second sentence. Every clause adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a deeply nested input schema, 0% schema descriptions, and no output schema, this overview is insufficient. The agent still lacks guidance on the `reason` field, the structure of `bill.lines`, payment optionality, and what the tool returns or does on preflight failure. It is a useful summary, not a complete spec.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It does map top-level concepts to parameters: cases, receipt, draft lines, approval/payment/reconciliation stages. However, it never mentions the `reason` parameter or explains the exact required fields (contactId, lines, etc.), leaving the agent to infer from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Preflight and save an ordered purchase batch (1-10 cases),' a specific action on a specific resource. It also states 'No Billy writes,' which clearly separates this from execution tools. This is enough to distinguish it from siblings like billy_batch_execute or billy_prepare.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'No Billy writes' gives a clear exclusion: this tool is for preparation/saving only, not for committing to Billy. It implies use before execution tools, but it does not name alternatives or explicitly state when not to use it. Thus it provides clear context but no explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_batch_refreshADestructive
Refresh initial evidence for a batch before any stage starts; clears prior client approval.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as destructive and not read-only. The description adds the specific destructive effect ('clears prior client approval'), which is valuable behavioral detail beyond the generic destructiveHint flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with two tightly related clauses. The main action and timing are front-loaded, and the destructive warning follows naturally without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, low-complexity tool with annotations covering safety (destructive, not idempotent, open world), the description provides enough operational context: what it does, when to do it, and what side effect to expect. It does not detail return behavior, but no output schema exists and none is strictly needed here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description technically doesn't explain batchId. However, the only parameter is self-evident from its name and UUID format, and the description's 'a batch' reinforces the target. This is adequate but not exemplary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Refresh initial evidence'), a clear object ('a batch'), and a temporal context ('before any stage starts'). It also names a distinctive side effect ('clears prior client approval'), which separates it from sibling batch tools like billy_batch_prepare or billy_refresh_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before any stage starts' gives explicit usage timing. It does not name alternatives or exclusions, but the timing constraint is clear enough context for an agent to select this tool for pre-stage refresh operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_executeADestructive
Execute the exact reviewed proposal once. Default approval uses the MCP client form; a supplied authorization note is only an audit assertion. Unknown outcomes block further writes.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| expectedHash | Yes | ||
| authorization | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the annotations: it explains the approval flow, clarifies that a supplied authorization note is only an audit assertion, and warns that unknown outcomes block further writes. These are valuable operational details not present in the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place. The core action is front-loaded, followed by critical authorization and failure behavior. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers key behavioral aspects such as idempotency, authorization, and failure blocking, which is strong for a destructive tool. However, with no output schema and two undocumented required parameters, an agent still lacks enough detail to fully understand expectedHash and planId semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It partially addresses the 'authorization' parameter by explaining its audit-only role, but it does not explain planId or expectedHash, leaving two required parameters semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Execute the exact reviewed proposal once.' This clearly distinguishes the tool from siblings like billy_batch_execute by emphasizing single, exact execution of a reviewed proposal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage after a proposal has been reviewed and approved, and it clarifies the authorization mechanism. However, it does not explicitly state when to prefer this tool over alternatives such as billy_batch_execute, nor does it mention exclusions or recovery paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_getARead-onlyIdempotent
Read a Billy record and optionally sideload relationships. Default response is compact; verbose returns sanitized detail.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| include | No | ||
| verbose | No | ||
| resource | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds useful behavioral context beyond annotations: default compact responses, verbose mode returning sanitized detail, and optional relationship sideloading. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no fluff. The core action is front-loaded, and the optional behaviors are stated efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple get-by-id tool, the description covers the main behaviors: reading, sideloading, and response verbosity. The lack of an output schema is partially mitigated by the compact/verbose distinction, though specific return fields are not described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify 'include' as relationship sideloading and 'verbose' as compact vs sanitized detail, but it does not describe the resource or id parameters beyond what their names and patterns imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description names a specific verb ('Read') and a resource ('a Billy record'), and the singular framing distinguishes it from billy_list. It does not explicitly mention 'by id', though the required id parameter makes this inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus siblings like billy_list or billy_batch_get. The context is implied by the singular 'record', but no when-to-use or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_import_receiptBDestructive
Archive an original PDF/PNG/JPEG from the configured inbox. Provide extracted invoice metadata and provenance from Gmail, Drive, local files or vendor portal. Same bytes are deduplicated. Does not upload to Billy.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| filePath | Yes | ||
| metadata | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds useful behavioral details: 'Same bytes are deduplicated' and 'Does not upload to Billy.' These clarify side effects and boundaries. It does not, however, explain what destructive action occurs (e.g., whether the source file is deleted or moved), which would be relevant given destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler; the core action and boundary are front-loaded. The second sentence is slightly awkward in its phrasing ('from the configured inbox' vs. 'from Gmail, Drive, local files or vendor portal'), but overall it is compact and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should indicate what the tool returns or what success looks like, but it does not. It also omits the required metadata fields and how deduplication affects the result. For a write/destructive tool with required nested parameters, this is a significant completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three required nested parameters. It gives broad hints by mentioning file formats, source kinds, and extracted invoice metadata, but it does not enumerate required metadata fields or explain the meaning of source.reference and vendorId. The schema itself is fairly descriptive, but the description adds insufficient parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Archive an original PDF/PNG/JPEG from the configured inbox' and clarifies the source types (Gmail, Drive, local files, vendor portal). The phrase 'Does not upload to Billy' helps distinguish this from other Billy-related write tools, although it does not name a specific sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when a receipt/invoice needs archiving with extracted metadata and provenance. However, it does not explicitly explain when not to use it or name alternatives like billy_prepare or billy_execute, leaving the routing decision somewhat to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_journalARead-onlyIdempotent
Read the last 200 proposals, including rejected/uncertain writes. An unknown outcome requires reconciliation against live Billy records before recovery.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, open-world behavior. The description adds valuable non-obvious traits: a fixed 200-entry cap, inclusion of rejected/uncertain writes, and the warning that unknown outcomes require reconciliation against live records before recovery. These go beyond the structured metadata and are consistent with openWorldHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences carry meaningful, non-redundant information. The primary action is front-loaded, and the crucial recovery caveat is placed second without padding. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only journal tool with simple optional parameters and no output schema, the description conveys the scope, content, and a critical operational caveat. It does not explain what a 'proposal' is or describe the return shape, but the sibling tool names and the Billy domain provide enough context for an agent to use it correctly. The undocumented 'verbose' parameter is a minor omission here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'verbose' has no schema description and schema coverage is 0%. The description does not mention or explain 'verbose' at all, leaving the agent to infer its meaning from the parameter name. Since the description fails to compensate for the missing schema documentation, the semantics are under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation ('Read') and the resource ('the last 200 proposals'), and adds the key distinguishing detail that rejected/uncertain writes are included. It does not explicitly name a sibling, but the 'proposals' framing with rejected/uncertain writes separates it from plain list/get tools like billy_list and billy_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when this tool is useful (inspecting recent proposals, including rejected/uncertain ones) and gives a follow-up directive for unknown outcomes. However, it does not explicitly state when to prefer billy_journal over billy_list, billy_get, or billy_status, nor does it mention alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_listARead-onlyIdempotent
List every page of a Billy resource. Default response is compact and complete; verbose returns sanitized detail. Unknown filters are rejected.
| Name | Required | Description | Default |
|---|---|---|---|
| filters | No | ||
| verbose | No | ||
| resource | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, non-destructive, idempotent behavior. The description adds valuable behavioral detail beyond those annotations: default responses are compact and complete, verbose returns sanitized detail, and unknown filters are rejected. This meaningfully helps an agent anticipate runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying distinct information: the core operation, response behavior, and a validation constraint. No filler or redundancy, and the most important purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter list tool with no output schema, the description covers the essential usage aspects: resource choice, filters, verbose mode, and rejection of unknown filters. It stops short of detailing pagination mechanics or response shape, but the provided annotations for safety and idempotence reduce the need for further explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description partially compensates by explaining verbose behavior and filter validation. However, it does not explain how filters should be structured or which filter fields are supported, and resource semantics are left entirely to the schema enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('List every page') and a clear object ('a Billy resource'), making it obvious this is a paginated listing tool. It doesn't explicitly name sibling tools like billy_get, but 'every page' implies collection-wide listing rather than fetching a single item.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The wording implies use when an agent needs all pages of a Billy resource, along with optional filtering and verbose output. However, it does not explicitly contrast with alternatives such as billy_get or billy_period_overview, leaving the when-not-to-use guidance to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_outstandingARead-onlyIdempotent
Current approved unpaid bills and invoices, grouped by original currency. Not a historical as-of report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds meaningful output behavior beyond those hints: only approved unpaid items, grouped by original currency, and current rather than historical. It does not describe return formatting or pagination, but those are less critical given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The core scope is front-loaded, and the clarifying 'not a historical as-of report' comes second as a precise caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only report tool with no output schema, this description is sufficient: it states the content, the approval/unpaid filter, the currency grouping, and the temporal nature. Nothing needed to decide whether to call it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so there is nothing for the description to add about parameter semantics. The baseline of 4 applies because no parameter documentation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource: current approved unpaid bills and invoices, and adds grouping by original currency. It also explicitly rules out historical as-of semantics, which helps distinguish it from sibling report tools. It lacks an explicit verb like 'list' or 'get,' but the purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Current... Not a historical as-of report' gives an explicit temporal scope and a when-not signal, so an agent can avoid using it for historical snapshots. It does not name sibling alternatives like billy_period_overview, but the current-vs-historical distinction provides clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_period_expensesBRead-onlyIdempotent
Net expense postings on debit-normal accounts under a live account-nature reportType (default incomeStatement), optionally narrowed to selected expense account IDs; credits reduce expense.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | ||
| start | Yes | ||
| verbose | No | ||
| accountIds | No | ||
| reportType | No | incomeStatement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, and non-destructive behavior, so the description adds value by explaining that credits reduce expense, only debit-normal accounts are included, and a live account-nature reportType is required. The phrase 'live account-nature' is somewhat cryptic, preventing a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence packs the core semantics and front-loads the most important qualifier, 'net expense postings.' The 'live account-nature' parenthetical is dense and could be clearer, but there is no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, no output schema, and a nonstandard term ('live account-nature reportType'), yet the description never explains what 'live' means, what the response looks like, or what happens if no postings exist. It is adequate for a simple read-only expense query but leaves operational details unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter semantics. It clarifies accountIds as expense account narrowing and reportType as defaulting to incomeStatement with account-nature semantics, and 'period' implies start/end. However, start, end, and verbose are not explicitly explained, leaving a meaningful gap for the required date range.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines the tool's output: net expense postings for a period, scoped to debit-normal accounts and optionally narrowed by accountIds, with a default reportType. This is more specific than a bare 'expenses' label and helps separate it from similar reporting tools, though it lacks an explicit verb like 'list' or 'fetch'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No alternative tool is mentioned and there is no when-to-use or when-not-to-use guidance. It implies the caller wants period expense postings and can optionally filter by accountIds, but it does not tell an agent when to prefer this over billy_period_overview or billy_profit_loss.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_period_overviewARead-onlyIdempotent
Inventory unreconciled bank lines, existing postings, bills and collected receipts for a period. Candidate matches are suggestions, not an accounting verdict.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | ||
| start | Yes | ||
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior, so the description does not need to restate that. It adds the important behavioral caveat that candidate matches are suggestions, not an accounting verdict, which is meaningful context beyond the annotations. This helps the agent interpret results correctly without overstepping.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff or restatement. The first sentence front-loads the action and scope, and the second quickly delivers the key interpretive caveat. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent tool with a simple three-parameter schema, the description communicates what is inventoried and how to interpret the output. The absence of an output schema makes the item enumeration and 'not an accounting verdict' caveat valuable. It does not specify exact output structure or period inclusivity, but the annotations and simple schema reduce the need for more.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It partially does by clarifying that start and end define a period and that the output includes specific record types, but it does not explain the optional verbose flag or date-boundary behavior. The schema provides patterns and defaults, and the parameter names are fairly intuitive, making this adequate though not thorough.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete action ('Inventory') and specifies the resource scope: unreconciled bank lines, postings, bills, and collected receipts over a period. This clearly distinguishes it from more generic sibling tools like billy_list or billy_get, though it does not explicitly call them out. The caveat about candidate matches further sharpens its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for a period' implies this tool is intended for period-scoped overviews of reconciliation-related entities, and the listed item types suggest when it would apply. However, there is no explicit guidance on when not to use it or which sibling tool to prefer for different needs, such as billy_outstanding or billy_get. This is adequate but relies on the agent to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_planBRead-onlyIdempotent
Inspect a saved proposal and its execution evidence. Default output retains operation, hash and snapshot hashes.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare read-only, idempotent, and non-destructive behavior, so the bar for disclosure is lower. The description adds useful context by stating what the default output contains ('operation, hash and snapshot hashes'), which goes beyond annotation data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the core purpose without filler. It is appropriately concise and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only two-parameter tool, the safety profile is well covered by annotations五星. However, the effect of 'verbose' and the full output structure are not explained, and there is no output schema to fill that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It only hints that planId identifies 'a saved proposal' and leaves the 'verbose' parameter completely unexplained, without any format or effect details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('Inspect') and names a specific resource ('a saved proposal and its execution evidence'). This distinguishes it from generic lookups like billy_get or billy_status, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus siblings such as billy_get, billy_status, or billy_execute. The agent must infer its role solely from the tool name and short description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_prepareBDestructive
Validate and persist a concrete write proposal without modifying Billy. Review the returned operation, reason, ID and hash. Receipts must be uploaded before preparing a booking. Reconciliation is restricted to matching existing bank-account postings.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| verbose | No | ||
| operation | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate non-read-only, destructive, and open-world behavior. The description adds meaningful behavioral context by stating that the tool does not modify Billy, that it returns an operation/reason/ID/hash for review, and that receipts must be uploaded first. These go beyond the annotations and help set expectations. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tight sentences. The first front-loads the core purpose, the second tells the agent what to review, and the third gives key constraints. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the enormous polymorphic input schema, no output schema, and annotations that flag destructive behavior, the description is far too sparse. It does not explain how to construct the `operation` object, what the returned hash/ID are for, whether execution is required afterward, or how errors are surfaced. An agent would be left guessing about most of the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the `operation`, `reason`, and `verbose` parameters. It does not: the mention of 'operation, reason, ID and hash' refers to the returned values, not the input. The only parameter-related hints are indirect constraints about receipts and reconciliation. This is far below what the complex polymorphic operation schema requires.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Validate and persist a concrete write proposal') with a resource and a key constraint ('without modifying Billy'). It distinguishes the tool as a prepare step, though it does not explicitly name alternatives like billy_batch_prepare or billy_execute, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a preparatory role and gives preconditions ('Receipts must be uploaded before preparing a booking', 'Reconciliation is restricted to matching existing bank-account postings'), but it never states when to use this tool versus siblings such as billy_batch_prepare or billy_execute. No alternatives or exclusions are mentioned, so usage guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_profit_lossARead-onlyIdempotent
Period P&L from non-voided postings. Defaults to live incomeStatement account natures, or accepts explicit reportType or revenue and expense account IDs; no chart numbers are assumed.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | ||
| start | Yes | ||
| verbose | No | ||
| reportType | No | ||
| expenseAccountIds | No | ||
| revenueAccountIds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds meaningful behavioral detail beyond those hints: it reports only non-voided postings, defaults to live incomeStatement account natures, and explicitly warns that no chart numbers are assumed. This reduces the risk of incorrect expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences deliver the core behavior first and the parameter flexibility second, with no wasted words. Every clause adds semantic value, and the structure supports quick agent comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the read-only annotations, six parameters, and no output schema, the description is largely complete: it explains data inclusion, defaults, and override behavior. It could optionally clarify verbose or the exact response shape, but these are not essential for a safe and correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It does for the most important parameters: reportType and revenue/expenseAccountIds indicate alternative selection modes, and 'Period' implies start/end. Only verbose is left unexplained, but its boolean nature and default false make that a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific deliverable ('Period P&L') and a precise data source ('non-voided postings'). It also distinguishes itself by explaining the default account-nature basis versus explicit reportType or revenue/expense account IDs, which differentiates it from account-code-driven siblings like billy_trial_balance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: use it for period profit/loss reporting, with sensible defaults applied automatically and explicit overrides available when needed. It does not explicitly name sibling alternatives or state when not to use the tool, but the provided context is sufficient for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_receiptsCRead-onlyIdempotent
List locally archived supporting documents and their Billy attachment IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is known. The description adds context about the scope (locally archived) and the output (attachment IDs), which is useful beyond annotations. However, it does not disclose any behavioral nuances such as pagination, sorting, or the effect of the 'verbose' parameter, so it is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tight sentence that front-loads the action and resource. There is no redundant wording, and it conveys the essence of the tool efficiently. Structure is simple and effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description covers the primary purpose but misses important context: it does not explain the 'verbose' parameter, nor does it clarify what 'locally archived' means or what 'supporting documents' encompass. The annotations cover safety, but usage and parameter semantics are incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has one optional boolean parameter 'verbose' with no description, and the tool description does not mention it at all. Since schema description coverage is 0%, the description must explain the parameter but fails to do so. The agent is left to guess what 'verbose' does, which is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action (List) and a resource (locally archived supporting documents) with a clear output (Billy attachment IDs). It is distinct enough from siblings like billy_import_receipt and billy_list, though it doesn't explicitly differentiate from billy_list. The phrase 'locally archived' gives context that it is a specific subset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of which circumstances warrant this tool instead of billy_list or billy_get, nor any exclusions. The description implies it is for archived documents but does not state it explicitly or provide decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_refresh_planADestructive
Refresh an unexecuted/rejected proposal after changed data or expiry; review it again before execution.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal destructive (destructiveHint: true) and read-write (readOnlyHint: false). The description adds meaningful context: the target state (unexecuted/rejected) and the reason for refreshing (data changes/expiry), plus the recommendation to review again. This goes beyond the annotations without contradicting them. It does not detail side effects (e.g., what exactly is overwritten), but the bar is lower given annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the action and resource, and packs in the state condition and workflow hint. No filler words; every clause adds information. It is optimally concise and structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers the core purpose and usage scenario. However, it omits any explanation of the optional 'verbose' parameter and the exact refresh semantics (what changes are applied). Given the openWorldHint, an agent might need to know if there are side effects beyond the plan refresh. Thus, it is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the parameters at all. planId is implied to be the proposal ID but is not explicitly stated, and the optional 'verbose' flag is completely undocumented. The description fails to compensate for the lack of schema descriptions, leaving the agent to infer parameter meanings from the tool name and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('refresh') and a specific resource ('unexecuted/rejected proposal'). It clarifies the condition ('after changed data or expiry') and the intended follow-up ('review it again before execution'). This clearly distinguishes it from sibling tools like billy_batch_refresh (batch vs singular) and billy_execute (execution vs refresh).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the trigger condition: refresh when a proposal is unexecuted or rejected and data changed or expired. It implies the workflow: refresh then review before execution. It does not explicitly name alternatives or state when not to use, but the condition is clear enough for an agent to decide based on the tool's state. Sibling tools like billy_prepare or billy_execute have different purposes, so the guidance is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_save_vendorADestructive
Store a vendor billing portal location and retrieval status. This registry guides the agent browser/connector; it does not log in itself. Never include secrets or session URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| name | Yes | ||
| notes | No | ||
| status | Yes | ||
| portalUrl | Yes | ||
| accountLabel | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a write operation (readOnlyHint false) and destructiveHint true, so the description adds value by stating it does not log in itself and by warning against including secrets or session URLs. This provides important security context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, with the core purpose stated upfront and the security note following. No waste, and the key constraint is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and no parameter details in the description, it is incomplete. It provides the purpose and a security constraint, but omits parameter semantics, usage conditions, and expected side effects, leaving gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description should compensate by explaining parameters. It only references 'portal location' and 'retrieval status' which loosely map to portalUrl and status, but it does not describe any of the six parameters, leaving the agent to infer meaning from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the verb 'Store' and the resource 'vendor billing portal location and retrieval status', which is specific. It also clarifies the tool's role in guiding the agent browser/connector, distinguishing it from tools that might perform login. However, it does not explicitly differentiate from sibling tools beyond this context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting it guides the browser/connector and does not log in, but it lacks explicit when-to-use guidance or alternatives. No mention of prerequisites or conditions for using this tool over siblings like billy_prepare or billy_execute.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_statusBRead-onlyIdempotent
Configuration and connection status. Never exposes credentials.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the call as read-only, idempotent, and non-destructive; the description adds a meaningful behavioral guarantee beyond those: 'Never exposes credentials.' It also tells the agent the output scope, configuration and connection state. This is good supplementary transparency for a safe status check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the tool's purpose and followed by a high-value safety note. No filler or repetition of schema/annotation information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status tool with strong annotations, the description covers the essential scope and credential safety. However, with no output schema, it doesn't describe what fields or connection details the status actually contains, and it leaves the meaning of verbose to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, verbose, has no description in the schema and schema description coverage is 0%. The tool description does not compensate by explaining what verbose changes about the output or how it should be used, so the description adds no parameter meaning beyond the raw name and boolean type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool surfaces configuration and connection status, which is a specific resource and a clear function. It doesn't explicitly name a verb like 'get' or contrast itself with siblings, but 'status' is unambiguous enough to distinguish it from the listed Billy tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to check status versus using a sibling tool, and no context about how this relates to refresh, execute, or batch operations. The only implicit signal is the word 'status', which suggests checking but doesn't explain prerequisites or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_trial_balanceBRead-onlyIdempotent
Balances by live account through an inclusive posting entry date. Uses base currency and rejects FX postings without a base amount.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | ||
| accountIds | No | ||
| includeZero | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, non-destructive behavior. The description adds valuable behavioral context beyond that: the date filter is inclusive, the tool operates in base currency, and FX postings without a base amount are rejected. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The core action is front-loaded, and the second sentence adds distinct behavioral detail about currency handling and FX rejection. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, three-parameter query, the description covers the essential behavior: inclusive as-of date, live-account grouping, base-currency handling, and FX rejection. It does not describe the output shape, but no output schema exists and the tool name plus 'balances' makes the return concept clear. Minor gaps around optional filters are acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that asOf is an inclusive posting entry date and that balances are per live account, which partially maps to accountIds, but it does not explain includeZero or the exact semantics of accountIds filtering. Compensation is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Balances') and resource ('by live account') with a temporal scope ('through an inclusive posting entry date'), so the core action is clear. It does not explicitly contrast itself with siblings like billy_profit_loss or billy_period_expenses, so it misses full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool—when you need account balances through a date—but gives no explicit when-to-use guidance and names no alternatives. With many financial sibling tools present, an agent gets little help choosing between this and billy_profit_loss or billy_journal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billy_vendorsBRead-onlyIdempotent
List vendor billing portals and access exceptions for receipt retrieval.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds the purpose 'for receipt retrieval', which provides context on why this tool exists but doesn't disclose any additional behavioral aspects like data freshness, pagination, or rate limits. Given annotations cover the main concerns, this score is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action ('List vendor billing portals and access exceptions') and adds a purpose ('for receipt retrieval'). There is zero waste and no unnecessary detail. It achieves maximum conciseness for the information it does provide.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema, the description is minimally adequate: it states what it lists and hints at why. However, it fails to explain the 'verbose' parameter or clarify what 'access exceptions' means, leaving some ambiguity. Given the tool's simplicity, a complete description would be concise but should at least mention the parameter's effect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has one optional boolean parameter 'verbose' with default false, but the schema description coverage is 0% (no property description). The tool description does not mention this parameter at all, so an agent has no idea what 'verbose' does or how it affects the output. The description should have compensated for the schema gap but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and a specific resource ('vendor billing portals and access exceptions'). It distinguishes itself from siblings like billy_list or billy_receipts by focusing on vendors, though it doesn't explicitly name alternatives. The purpose is clear enough for an agent to infer it's about vendor-related billing data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus siblings like billy_list, billy_receipts, or billy_save_vendor. There is no mention of prerequisites, typical use cases, or exclusions. An agent would have to rely on the name alone, which is insufficient for proper tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
v0.2.1- Added
billy_batch_execute - Added
billy_batch_get - Added
billy_batch_prepare - Added
billy_batch_refresh - Changed
billy_get2 fields changed- changed
Input schema / properties / resource / enumPrevious value: -[ - "accounts", - "contacts", - "invoices", - "products", - "bills", - "bankPayments", - "bankLines", - "bankLineMatches", - "bankLineSubjectAssociations", - "daybooks", - "daybookTransactions", - "postings", - "taxRates", - "attachments", - "files", - "salesTaxReturns", - "transactions" -]New value: +[ + "accounts", + "accountGroups", + "accountNatures", + "contacts", + "contactPersons", + "invoices", + "products", + "bills", + "bankPayments", + "bankLines", + "bankLineMatches", + "bankLineSubjectAssociations", + "daybooks", + "daybookTransactions", + "postings", + "taxRates", + "salesTaxRulesets", + "attachments", + "files", + "salesTaxReturns", + "transactions" +] - added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
billy_import_receipt2 fields changed- added
Input schema / properties / metadata / properties / creditedInvoiceNumberAdded value: +{ + "maxLength": 160, + "minLength": 1, + "type": "string" +} - added
Input schema / properties / metadata / properties / documentTypeAdded value: +{ + "enum": [ + "invoice", + "creditNote" + ], + "type": "string" +}
- Changed
billy_journal1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
billy_list2 fields changed- changed
Input schema / properties / resource / enumPrevious value: -[ - "accounts", - "contacts", - "invoices", - "products", - "bills", - "bankPayments", - "bankLines", - "bankLineMatches", - "bankLineSubjectAssociations", - "daybooks", - "daybookTransactions", - "postings", - "taxRates", - "attachments", - "files", - "salesTaxReturns", - "transactions" -]New value: +[ + "accounts", + "accountGroups", + "accountNatures", + "contacts", + "contactPersons", + "invoices", + "products", + "bills", + "bankPayments", + "bankLines", + "bankLineMatches", + "bankLineSubjectAssociations", + "daybooks", + "daybookTransactions", + "postings", + "taxRates", + "salesTaxRulesets", + "attachments", + "files", + "salesTaxReturns", + "transactions" +] - added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Added
billy_outstanding - Added
billy_period_expenses - Changed
billy_period_overview1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
billy_plan1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
billy_prepare2 fields changed- changed
Input schema / properties / operation / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "contactId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "currencyId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "entryDate": { - "pattern": "^\\d{4}-\\d{2}-\\d{2}$", - "type": "string" - }, - "kind": { - "const": "create_bill", - "type": "string" - }, - "lines": { - "items": { - "additionalProperties": false, - "properties": { - "accountId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "amount": { - "maximum": 10000000000, - "minimum": 0, - "type": "number" - }, - "description": { - "maxLength": 1000, - "minLength": 1, - "type": "string" - }, - "taxRateId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - } - }, - "required": [ - "accountId", - "taxRateId", - "description", - "amount" - ], - "type": "object" - }, - "maxItems": 100, - "minItems": 1, - "type": "array" - }, - "receiptId": { - "pattern": "^[a-f0-9]{64}$", - "type": "string" - }, - "suppliersInvoiceNo": { - "maxLength": 160, - "minLength": 1, - "type": "string" - }, - "taxMode": { - "enum": [ - "incl", - "excl" - ], - "type": "string" - } - }, - "required": [ - "kind", - "receiptId", - "contactId", - "entryDate", - "currencyId", - "suppliersInvoiceNo", - "taxMode", - "lines" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "bankLineId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "daybookId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "description": { - "maxLength": 1000, - "minLength": 1, - "type": "string" - }, - "entryDate": { - "pattern": "^\\d{4}-\\d{2}-\\d{2}$", - "type": "string" - }, - "kind": { - "const": "create_journal", - "type": "string" - }, - "lines": { - "items": { - "additionalProperties": false, - "properties": { - "accountId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "amount": { - "maximum": 10000000000, - "minimum": 0, - "type": "number" - }, - "currencyId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "side": { - "enum": [ - "debit", - "credit" - ], - "type": "string" - }, - "taxRateId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "text": { - "maxLength": 500, - "minLength": 1, - "type": "string" - } - }, - "required": [ - "accountId", - "text", - "amount", - "side", - "currencyId" - ], - "type": "object" - }, - "maxItems": 100, - "minItems": 2, - "type": "array" - }, - "noReceiptReason": { - "maxLength": 1000, - "minLength": 10, - "type": "string" - }, - "receiptId": { - "pattern": "^[a-f0-9]{64}$", - "type": "string" - } - }, - "required": [ - "kind", - "daybookId", - "entryDate", - "description", - "lines" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "bankLineId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "cashAccountId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "cashAmount": { - "maximum": 10000000000, - "minimum": 0, - "type": "number" - }, - "cashExchangeRate": { - "exclusiveMinimum": 0, - "maximum": 100000000, - "type": "number" - }, - "cashSide": { - "enum": [ - "debit", - "credit" - ], - "type": "string" - }, - "entryDate": { - "pattern": "^\\d{4}-\\d{2}-\\d{2}$", - "type": "string" - }, - "kind": { - "const": "create_payment", - "type": "string" - }, - "subjectAmount": { - "maximum": 10000000000, - "minimum": 0, - "type": "number" - }, - "subjectCurrencyId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "subjectReference": { - "pattern": "^(bill|invoice):[A-Za-z0-9_-]{1,160}$", - "type": "string" - } - }, - "required": [ - "kind", - "entryDate", - "cashAmount", - "cashSide", - "cashAccountId", - "bankLineId", - "subjectReference" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "id": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "kind": { - "const": "approve", - "type": "string" - }, - "resource": { - "enum": [ - "bills", - "invoices", - "daybookTransactions" - ], - "type": "string" - } - }, - "required": [ - "kind", - "resource", - "id" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "kind": { - "const": "upload_receipt", - "type": "string" - }, - "receiptId": { - "pattern": "^[a-f0-9]{64}$", - "type": "string" - } - }, - "required": [ - "kind", - "receiptId" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "countryId": { - "pattern": "^[A-Z]{2}$", - "type": "string" - }, - "email": { - "format": "email", - "pattern": "^(?:[A-Za-z0-9_'+\\-]+\\.)*[A-Za-z0-9_'+\\-]*[A-Za-z0-9_+-]@(?:[A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$", - "type": "string" - }, - "isCustomer": { - "type": "boolean" - }, - "isSupplier": { - "type": "boolean" - }, - "kind": { - "const": "create_contact", - "type": "string" - }, - "name": { - "maxLength": 200, - "minLength": 1, - "type": "string" - }, - "registrationNo": { - "maxLength": 100, - "type": "string" - } - }, - "required": [ - "kind", - "name", - "countryId", - "isSupplier", - "isCustomer" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "bankLineId": { - "pattern": "^[A-Za-z0-9_-]{1,160}$", - "type": "string" - }, - "kind": { - "const": "reconcile", - "type": "string" - }, - "subjectReference": { - "pattern": "^(invoice|bill|posting|daybookTransaction|bankPayment):[A-Za-z0-9_-]{1,160}$", - "type": "string" - } - }, - "required": [ - "kind", - "bankLineId", - "subjectReference" - ], - "type": "object" - } -]New value: +[ + { + "additionalProperties": false, + "properties": { + "contactId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "currencyId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "entryDate": { + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" + }, + "kind": { + "const": "create_bill", + "type": "string" + }, + "lines": { + "items": { + "additionalProperties": false, + "properties": { + "accountId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "amount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "description": { + "maxLength": 1000, + "minLength": 1, + "type": "string" + }, + "taxRateId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + } + }, + "required": [ + "accountId", + "taxRateId", + "description", + "amount" + ], + "type": "object" + }, + "maxItems": 100, + "minItems": 1, + "type": "array" + }, + "receiptId": { + "pattern": "^[a-f0-9]{64}$", + "type": "string" + }, + "suppliersInvoiceNo": { + "maxLength": 160, + "minLength": 1, + "type": "string" + }, + "taxMode": { + "enum": [ + "incl", + "excl" + ], + "type": "string" + } + }, + "required": [ + "kind", + "receiptId", + "contactId", + "entryDate", + "currencyId", + "suppliersInvoiceNo", + "taxMode", + "lines" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "bankLineId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "daybookId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "description": { + "maxLength": 1000, + "minLength": 1, + "type": "string" + }, + "entryDate": { + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" + }, + "kind": { + "const": "create_journal", + "type": "string" + }, + "lines": { + "items": { + "additionalProperties": false, + "properties": { + "accountId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "amount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "currencyId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "side": { + "enum": [ + "debit", + "credit" + ], + "type": "string" + }, + "taxRateId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "text": { + "maxLength": 500, + "minLength": 1, + "type": "string" + } + }, + "required": [ + "accountId", + "text", + "amount", + "side", + "currencyId" + ], + "type": "object" + }, + "maxItems": 100, + "minItems": 2, + "type": "array" + }, + "noReceiptReason": { + "maxLength": 1000, + "minLength": 10, + "type": "string" + }, + "receiptId": { + "pattern": "^[a-f0-9]{64}$", + "type": "string" + } + }, + "required": [ + "kind", + "daybookId", + "entryDate", + "description", + "lines" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "bankLineId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "cashAccountId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "cashAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "cashExchangeRate": { + "exclusiveMinimum": 0, + "maximum": 100000000, + "type": "number" + }, + "cashSide": { + "enum": [ + "debit", + "credit" + ], + "type": "string" + }, + "entryDate": { + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" + }, + "feeAccountId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "feeAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "kind": { + "const": "create_payment", + "type": "string" + }, + "subjectAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "subjectCurrencyId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "subjectReference": { + "pattern": "^(bill|invoice):[A-Za-z0-9_-]{1,160}$", + "type": "string" + } + }, + "required": [ + "kind", + "entryDate", + "cashAmount", + "cashSide", + "cashAccountId", + "bankLineId", + "subjectReference" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "id": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "kind": { + "const": "approve", + "type": "string" + }, + "resource": { + "enum": [ + "bills", + "invoices", + "daybookTransactions" + ], + "type": "string" + } + }, + "required": [ + "kind", + "resource", + "id" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "kind": { + "const": "upload_receipt", + "type": "string" + }, + "receiptId": { + "pattern": "^[a-f0-9]{64}$", + "type": "string" + } + }, + "required": [ + "kind", + "receiptId" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "countryId": { + "pattern": "^[A-Z]{2}$", + "type": "string" + }, + "email": { + "format": "email", + "pattern": "^(?:[A-Za-z0-9_'+\\-]+\\.)*[A-Za-z0-9_'+\\-]*[A-Za-z0-9_+-]@(?:[A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$", + "type": "string" + }, + "isCustomer": { + "type": "boolean" + }, + "isSupplier": { + "type": "boolean" + }, + "kind": { + "const": "create_contact", + "type": "string" + }, + "name": { + "maxLength": 200, + "minLength": 1, + "type": "string" + }, + "registrationNo": { + "maxLength": 100, + "type": "string" + } + }, + "required": [ + "kind", + "name", + "countryId", + "isSupplier", + "isCustomer" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "bankLineId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "kind": { + "const": "reconcile", + "type": "string" + }, + "subjectReference": { + "pattern": "^(invoice|bill|posting|daybookTransaction|bankPayment):[A-Za-z0-9_-]{1,160}$", + "type": "string" + } + }, + "required": [ + "kind", + "bankLineId", + "subjectReference" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "contactId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "contactMessage": { + "maxLength": 2000, + "type": "string" + }, + "currencyId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "entryDate": { + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" + }, + "expectedNetAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTaxAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTotalAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "kind": { + "const": "create_sales_invoice", + "type": "string" + }, + "lines": { + "items": { + "additionalProperties": false, + "properties": { + "description": { + "maxLength": 1000, + "minLength": 1, + "type": "string" + }, + "expectedTaxRateId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "productId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "quantity": { + "exclusiveMinimum": 0, + "maximum": 100000000, + "type": "number" + }, + "salesTaxRulesetId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "unitPrice": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + } + }, + "required": [ + "productId", + "salesTaxRulesetId", + "expectedTaxRateId", + "quantity", + "unitPrice" + ], + "type": "object" + }, + "maxItems": 100, + "minItems": 1, + "type": "array" + }, + "taxMode": { + "enum": [ + "incl", + "excl" + ], + "type": "string" + } + }, + "required": [ + "kind", + "contactId", + "entryDate", + "currencyId", + "taxMode", + "lines", + "expectedNetAmount", + "expectedTaxAmount", + "expectedTotalAmount" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "contactMessage": { + "maxLength": 2000, + "type": "string" + }, + "expectedNetAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTaxAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTotalAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "id": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "kind": { + "const": "update_draft_invoice", + "type": "string" + }, + "paymentTermsDays": { + "description": "Net days from invoice entryDate; sets paymentTermsMode to net and verifies the computed dueDate", + "maximum": 3650, + "minimum": -365, + "type": "integer" + }, + "taxMode": { + "enum": [ + "incl", + "excl" + ], + "type": "string" + } + }, + "required": [ + "kind", + "id", + "expectedNetAmount", + "expectedTaxAmount", + "expectedTotalAmount" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "contactPersonId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "emailBody": { + "maxLength": 10000, + "minLength": 1, + "type": "string" + }, + "emailSubject": { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + "id": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "kind": { + "const": "send_invoice", + "type": "string" + }, + "recipientEmail": { + "format": "email", + "pattern": "^(?:[A-Za-z0-9_'+\\-]+\\.)*[A-Za-z0-9_'+\\-]*[A-Za-z0-9_+-]@(?:[A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$", + "type": "string" + } + }, + "required": [ + "kind", + "id", + "contactPersonId", + "recipientEmail", + "emailSubject", + "emailBody" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "entryDate": { + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" + }, + "expectedNetAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTaxAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTotalAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "kind": { + "const": "create_customer_credit_note", + "type": "string" + }, + "lines": { + "items": { + "additionalProperties": false, + "properties": { + "description": { + "maxLength": 1000, + "minLength": 1, + "type": "string" + }, + "expectedTaxRateId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "originalLineId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "productId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "quantity": { + "exclusiveMinimum": 0, + "maximum": 100000000, + "type": "number" + }, + "salesTaxRulesetId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "unitPrice": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + } + }, + "required": [ + "productId", + "salesTaxRulesetId", + "expectedTaxRateId", + "quantity", + "unitPrice", + "originalLineId" + ], + "type": "object" + }, + "maxItems": 100, + "minItems": 1, + "type": "array" + }, + "originalInvoiceId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + } + }, + "required": [ + "kind", + "originalInvoiceId", + "entryDate", + "lines", + "expectedNetAmount", + "expectedTaxAmount", + "expectedTotalAmount" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "entryDate": { + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" + }, + "expectedNetAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTaxAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "expectedTotalAmount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "kind": { + "const": "create_supplier_credit_note", + "type": "string" + }, + "lines": { + "items": { + "additionalProperties": false, + "properties": { + "accountId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "amount": { + "maximum": 10000000000, + "minimum": 0, + "type": "number" + }, + "description": { + "maxLength": 1000, + "minLength": 1, + "type": "string" + }, + "originalLineId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "taxRateId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + } + }, + "required": [ + "originalLineId", + "accountId", + "taxRateId", + "description", + "amount" + ], + "type": "object" + }, + "maxItems": 100, + "minItems": 1, + "type": "array" + }, + "originalBillId": { + "pattern": "^[A-Za-z0-9_-]{1,160}$", + "type": "string" + }, + "receiptId": { + "pattern": "^[a-f0-9]{64}$", + "type": "string" + } + }, + "required": [ + "kind", + "receiptId", + "originalBillId", + "entryDate", + "lines", + "expectedNetAmount", + "expectedTaxAmount", + "expectedTotalAmount" + ], + "type": "object" + } +] - added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Added
billy_profit_loss - Changed
billy_receipts1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
billy_refresh_plan1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
billy_status1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
- Added
billy_trial_balance - Changed
billy_vendors1 field changed- added
Input schema / properties / verboseAdded value: +{ + "default": false, + "type": "boolean" +}
13 tool updates
v0.1.0- First observed
billy_execute - First observed
billy_get - First observed
billy_import_receipt - First observed
billy_journal - First observed
billy_list - First observed
billy_period_overview - First observed
billy_plan - First observed
billy_prepare - First observed
billy_receipts - First observed
billy_refresh_plan - First observed
billy_save_vendor - First observed
billy_status - First observed
billy_vendors
TDQS
Scored across 21 tools
The toolset separates generic reads, financial reports, receipt/vendor management, and proposal/batch write workflows. Some overlap exists among the reporting tools and between refresh_plan/batch_refresh or plan/batch_get, but the descriptions are specific enough to guide selection.
All tools share the billy_ prefix and snake_case, but the action pattern is inconsistent: bare verbs, noun phrases, verb_noun, and object_verb forms are mixed. This is readable but lacks a uniform verb_noun convention.
21 tools is on the heavy side and spans multiple subdomains such as reporting, receipts, vendors, single proposals, and batches. Each tool appears purposeful, but the count is at the borderline where agents may have trouble finding the right tool.
The surface covers status, generic reads, financial reports, receipt archival, vendor registry, single-plan lifecycle, and batch lifecycle, with a journal for recovery after unknown outcomes. Minor gaps like explicit cancellation/voiding or direct upload-to-Billy are not clearly exposed, but the proposal workflow avoids dead ends.
Maintenance
Related MCP Connectors
Headless API-first double-entry accounting & bookkeeping engine. 84 MCP tools over HTTP.
Hosted MCP server for Mini Accountant: invoices, expenses, customers, analytics, tax estimates.
Bookkeeping for owner-operated businesses. Query transactions, invoices, and reports.
- ManiloOAuthapp.ledgy.api
Log, query, and edit expenses, budgets, and accounts in Manilo (formerly Ledgy) from any MCP-compatible AI assistant.
Related MCP Servers
- FlicenseAqualityDmaintenanceProvides double-entry accounting ledger creation, transaction recording, and financial reporting capabilities via MCP.7-
- AlicenseBqualityDmaintenanceAn MCP server for Danish accounting via Billy.dk API, enabling natural-language control over invoices, bank lines, reports, and more, with a write-guard for safety.651MIT
- AlicenseAqualityDmaintenanceConnects AI assistants to the Danish accounting platform Billy (billy.dk) for managing invoices, journal entries, balances, and receipts, with built-in human approval for all write operations.4047 npm1MIT
- AlicenseAqualityDmaintenanceIntegrates Billy's accounting system with MCP, providing tools to manage invoices, contacts, products, payments, and more via natural language.152MIT