TapTax
Server Details
UK self-employed tax: tax so far, receipts, mileage, invoices and Making Tax Digital updates.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 15 tools
Each tool targets a distinct resource and action: transactions (add/list/update), invoices (create/list/send/mark_paid), mileage (log/summary), and tax checks (status, code, estimate, position). The MTD tools (check_mtd_status vs get_mtd_deadlines vs prepare_quarterly_update) and the two estimation tools (check_tax_code vs estimate_self_employed_tax) overlap conceptually but descriptions clearly separate eligibility, deadlines, and filing prep from PAYE codes and sole-trader estimates.
Every tool follows a consistent snake_case verb_noun pattern (add_transaction, check_tax_code, create_invoice, list_invoices, send_invoice, update_transaction). No camelCase or mixed-convention deviations are present.
15 tools sit within the well-scoped range and each maps to a meaningful capability across invoicing, transactions, mileage, and MTD/tax checks. Nothing appears redundant or padded.
Core lifecycle coverage is strong: transaction create/list/update, invoice create/list/send/mark_paid, and mileage log/summary form coherent workflows. Gaps exist around deletion/cancellation (no delete_transaction, void/delete invoice, or edit/delete mileage trip), but these are minor and workable via the app.
Available Tools
15 toolsadd_transactionAdd income or an expenseAInspect
Record a business income or expense in the signed-in person's TapTax, such as a receipt or a payment. The category is the HMRC self-employment box it belongs in. The same entry added twice within 10 minutes is recorded once.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | The date on the receipt or of the payment | |
| type | Yes | ||
| amount | Yes | In pounds, VAT included | |
| category | Yes | For an expense, its HMRC self-employment box (SA103F): costOfGoods: materials, parts and stock used in the work or sold on, e.g. trade and building suppliers such as Screwfix, Toolstation, City Electrical Factors, Edmundson Electrical, Travis Perkins, Wickes; constructionIndustryScheme: payments to subcontractors under the Construction Industry Scheme; staffCosts: wages, salaries and pension contributions for employees, and payments to non-CIS subcontractors; travelCosts: car, van and travel: fuel, vehicle insurance, servicing and repairs, road tax, parking, tolls, public transport and taxi fares, and hotels and meals on overnight business trips; premisesRunningCosts: rent, business rates, power, heating, water, insurance and security for business premises (never vehicles); maintenanceCosts: repairs and maintenance of business premises, machinery and equipment; adminCosts: phone, broadband, postage, stationery, printing, small office items, and software and app subscriptions such as accounting or tax software (TapTax, Xero, QuickBooks); advertisingCosts: advertising, marketing and website costs; businessEntertainmentCosts: entertaining clients or customers; interest: interest on business loans and overdrafts; financialCharges: bank, overdraft, card and payment-processing charges; badDebt: money owed to the business that will not be paid and was counted as income; professionalFees: accountants, solicitors and other professional fees, and professional indemnity insurance; depreciation: depreciation of equipment (rare in a bank transaction); other: other allowable business costs, e.g. protective clothing, trade and professional body memberships, public liability insurance. For income: turnover (sales and fees) or other (other business income). | |
| personal | No | True for a personal (non-business) payment | |
| description | Yes | Supplier or client, and what it was for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-read-only, non-destructive, openWorld=false mutation. The description adds genuinely new behavior beyond them: a 10-minute deduplication window ('The same entry added twice within 10 minutes is recorded once'). That is exactly the kind of side-effect nuance the annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the category semantics, then the dedup caveat. No filler and no repetition of schema content; every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation with no output schema, the description covers the action, the categorization rule and the dedup caveat, while annotations carry the safety profile and the schema carries parameter detail. It stops short of stating auth/account requirements or what a successful response confirms, which is a minor gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83% (above the 80% baseline), and the category enum is already documented exhaustively in the schema. The description only adds the framing that a category is 'the HMRC self-employment box it belongs in' and that examples are receipts/payments, which is marginal over the structured data. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record a business income or expense in the signed-in person's TapTax') and concretizes it with examples (a receipt or a payment). The verb 'Record' cleanly separates it from list_transactions, update_transaction and create_invoice, so an agent can place it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (recording business receipts/payments for tax categorization) but never states when to prefer this over create_invoice or prepare_quarterly_update, nor any exclusions. Usage is inferable rather than stated, which is the definition of a 3.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_mtd_statusCheck if Making Tax Digital appliesARead-onlyInspect
Say whether and when Making Tax Digital for Income Tax applies to someone, from their qualifying income: self-employment turnover plus property income, before expenses, as on their Self Assessment return. Employment and pension income do not count.
| Name | Required | Description | Default |
|---|---|---|---|
| propertyIncome | No | ||
| selfEmploymentTurnover | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds domain semantics (income counted before expenses) but says nothing about what the determination returns — a boolean, a date, a threshold comparison — which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the determination being made before the qualifying-income rule. Every clause carries information the agent needs and nothing is repeated from the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only check with no output schema, the description supplies the essential rule for producing a correct answer (what counts toward qualifying income). The main gap is the absence of any hint about the output form or applicable thresholds, though the annotation set covers safety concerns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it does: it defines qualifying income as self-employment turnover plus property income, before expenses, and states what is excluded. It does not clarify that propertyIncome defaults to 0, nor units/currency, but the meaning of both parameters is substantially conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('say whether and when') and a precise resource (Making Tax Digital for Income Tax applicability), plus the basis for the determination (qualifying income). This distinguishes it cleanly from siblings like get_mtd_deadlines (dates) and estimate_self_employed_tax (tax estimate) without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for invoking it (anyone whose MTD status must be determined from qualifying income) and rules out misuse by stating employment and pension income do not count. It does not, however, route the agent explicitly to alternatives such as get_mtd_deadlines or prepare_quarterly_update when those are the better fit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_tax_codeCheck a PAYE tax codeARead-onlyInspect
Check whether a UK PAYE tax code (e.g. 1257L, BR, K475, 1257L W1/M1) looks right for a salary and circumstances, explain what it means, and estimate any over- or under-payment. Flat-rate codes (BR, D0, D1) are only judged when hasMultipleJobs is known; without it the verdict is needs_info. Returns a verdict tier; refund figures are estimates, not amounts owed.
| Name | Required | Description | Default |
|---|---|---|---|
| taxCode | Yes | The code as printed on a payslip or P45 | |
| taxYear | No | Tax year the code applies to; defaults to the current one | |
| scotland | No | True for a Scottish taxpayer | |
| grossSalary | Yes | Annual gross pay from this job, in pounds | |
| monthsOnCode | No | For emergency codes: how many monthly payslips were paid on it so far | |
| hasCompanyCar | No | ||
| hasMultipleJobs | No | Whether the person has another job or pension; omitted when unknown | |
| hasPrivateMedical | No | ||
| hasMarriageAllowance | No | ||
| receivesStatePension | No | ||
| marriageAllowanceDirection | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false, destructiveHint=false), so the description correctly focuses on behavior the annotations cannot express: a verdict-tier output, the needs_info fallback when hasMultipleJobs is absent, and the caveat that refund figures are estimates rather than amounts owed. It does not describe the full shape of the verdict or its other fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core purpose before the edge-case rule and the estimate caveat. Every sentence carries information, though the middle sentence is packed and could be split for scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 parameters, no output schema, and no annotations covering inputs, the description leaves significant ground uncovered: it never explains the purpose of the many circumstance flags or how taxYear/scotland/monthsOnCode affect the result, and it only gestures at the return structure via 'verdict tier'. Adequate for the core case, incomplete for a tool of this parameter complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 55% across 11 parameters, so the schema documents roughly half. The description adds real value by giving concrete tax-code formats (1257L, BR, K475, 1257L W1/M1) and by explaining the role of hasMultipleJobs, but it is silent on the many boolean circumstance flags (hasCompanyCar, hasPrivateMedical, hasMarriageAllowance, receivesStatePension, scotland, monthsOnCode) that materially change the verdict.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (UK PAYE tax code), plus the three things it does: validate, explain, and estimate over/under-payment. Sibling tools are all invoicing/transaction/mileage oriented, so this tool is unambiguously distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit conditional rule for a hard case: flat-rate codes (BR, D0, D1) are only judged when hasMultipleJobs is known, otherwise the verdict is needs_info. That tells the agent when the tool can and cannot give a real answer. It does not name alternative tools or state broader when-not conditions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_invoiceCreate an invoiceAInspect
Create a draft invoice in the signed-in person's TapTax, numbered in their sequence and using their saved business details. The invoice is created as a draft and is not sent. Free plans include 3 invoices; paid plans are unlimited.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| notes | No | ||
| dueDate | No | Exact due date; otherwise dueInDays applies | |
| dueInDays | No | Payment terms in days when no dueDate is given (0 = due on receipt) | |
| issueDate | No | Defaults to today | |
| clientName | Yes | ||
| clientEmail | No | The client email; required to send the invoice | |
| clientAddress | No | ||
| vatRatePercent | No | VAT rate for a VAT-registered business, e.g. 20 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly false, destructive false, idempotent false, openWorld false), yet the description adds meaningful behavior beyond that: the invoice is a draft and unsent, it is auto-numbered in sequence, it uses saved business details, and free plans have a 3-invoice quota while paid plans are unlimited. This is exactly the kind of side-effect, quota, and post-condition context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences with the core action front-loaded and no filler. The quota sentence is relevant but slightly separated from the main behavioral statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool, the description covers the key behavioral traits an agent needs: draft state, auto-numbering, business detail reuse, and plan limits. Annotations carry the safety profile and the schema covers many parameters, so no critical gap remains, though more parameter context and an explicit alternative for sending would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not discuss any of the 9 parameters, so it adds no meaning beyond what the schema provides. Schema description coverage is only 56%, leaving required fields like clientName and the items array without descriptive guidance in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'Create a draft invoice' — and adds scope ('in the signed-in person's TapTax') and a key distinguishing property ('created as a draft and is not sent'). This clearly separates it from siblings like send_invoice and mark_invoice_paid without requiring the schema to be opened.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when this tool applies: the invoice is created as a draft and is not sent. This implicitly routes sending to send_invoice, but no alternative tool is named explicitly and there are no stated exclusions or prerequisites beyond the free-plan quota note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_self_employed_taxEstimate self-employed taxARead-onlyInspect
Estimate UK income tax and Class 4 National Insurance for a sole trader from turnover and allowable expenses, with the take-home figure, effective and marginal rates, payments on account and tailored tips, from HMRC's published rates for the tax year. Works from the figures given, without a TapTax account.
| Name | Required | Description | Default |
|---|---|---|---|
| taxYear | No | UK tax year, e.g. "2026-27" (6 April 2026 to 5 April 2027). Defaults to the current one. | |
| expenses | No | Allowable business expenses, in pounds (excluding mileage if businessMiles is given) | |
| turnover | Yes | Self-employed income before expenses, in pounds | |
| otherIncome | No | Other taxable income such as a salary or pension, in pounds | |
| businessMiles | No | Business miles driven in a car or van, claimed at the approved mileage rate | |
| studentLoanPlan | No | ||
| taxDeductedAtSource | No | Tax already paid through PAYE or CIS deductions this year | |
| pensionContributions | No | Personal pension contributions paid (relief at source), in pounds |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: figures come from HMRC's published rates for the stated tax year, no account is required, and the response contains effective/marginal rates and payments on account. It doesn't discuss rounding, rate-update lag, or limitations of the estimate, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence, front-loaded with the core purpose and inputs before the output list. Every clause carries information, though the trailing enumeration of outputs makes it slightly long; still efficient overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, no output schema, and read-only annotations, the description compensates well by enumerating the returned figures (take-home, effective/marginal rates, payments on account, tips) and clarifying the no-account scenario. Minor gaps remain around assumptions and edge cases for out-of-range or unsupported tax years.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 88%, so the schema already documents nearly every parameter, including enums for taxYear and studentLoanPlan. The description only echoes 'turnover and allowable expenses' and adds no format, unit, or interaction detail (e.g. how businessMiles interacts with expenses) beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Estimate), precise resource (UK income tax and Class 4 NI for a sole trader), and the inputs it works from (turnover, allowable expenses). It also enumerates what comes back (take-home, effective/marginal rates, payments on account, tips), which clearly separates it from sibling tools like get_tax_position or check_tax_code.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'Works from the figures given, without a TapTax account' tells the agent when this tool is appropriate: standalone what-if estimation rather than account-based lookups. It does not explicitly name an alternative sibling (e.g. get_tax_position) or state exclusions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mileage_summaryGet my mileageBRead-onlyInspect
The signed-in person's business mileage for a tax year: total miles, the deductible amount at the approved rates, and the most recent trips.
| Name | Required | Description | Default |
|---|---|---|---|
| taxYear | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds useful scope ('the signed-in person's', deductible amounts at approved rates), but omits what happens when taxYear is omitted (required parameters = 0), which is a real behavioral gap for a 0-required-param tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with no filler, leading with the scope (signed-in person's business mileage) and then listing exactly what is returned. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully enumerates the return contents. Combined with read-only annotations, an agent has almost everything it needs; only the omitted-parameter default is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It does say the summary is 'for a tax year', tying to the taxYear parameter, but it never states the default when taxYear is omitted or clarifies the enum values beyond their self-evident format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names the specific resource (business mileage for a tax year) and enumerates the payload: total miles, deductible amount at approved rates, recent trips. An agent can distinguish it from log_mileage_trip, which writes trips rather than summarizing them. It is a noun phrase rather than an explicit verb, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named. It never says to use log_mileage_trip first to record trips, nor when this summary is preferable to estimate_self_employed_tax or prepare_quarterly_update, all of which touch mileage-related tax figures.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mtd_deadlinesMaking Tax Digital deadlinesARead-onlyInspect
List the standard Making Tax Digital quarterly update periods and deadlines for a tax year, and which one is next. Updates are cumulative from 6 April.
| Name | Required | Description | Default |
|---|---|---|---|
| taxYear | No | Defaults to the current tax year |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, non-destructive, and closed-world, so the safety profile is covered. The description adds useful domain context ('Updates are cumulative from 6 April') and implies a next-deadline indicator, but says nothing about return shape, ordering, or behavior for years outside the enum.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence plus one short supporting clause, front-loaded with the core action and resource. No filler; the cumulative-updates note earns its place by explaining the deadline structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-optional-parameter lookup with full annotation coverage, the description gives enough to call it correctly. Without an output schema it could say a bit more about what the returned list looks like, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single taxYear parameter carries its enum and default ('Defaults to the current tax year') in the schema. The description's 'Updates are cumulative from 6 April' adds interpretive context for the returned periods but no parameter-level detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (Making Tax Digital quarterly update periods and deadlines), plus a distinctive extra output ('which one is next'). The scope is clear enough to separate it from check_mtd_status and prepare_quarterly_update, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is for looking up statutory MTD deadlines, but there is no explicit when-to-use, when-not-to-use, or pointer to an alternative such as prepare_quarterly_update for acting on a deadline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_tax_positionGet my tax positionARead-onlyInspect
The signed-in person's tax year so far, as the TapTax dashboard shows it: business income, expenses and profit counted by the filing rule, the estimated income tax and Class 4 NI, how many transactions still need a business-or-personal decision, which quarterly updates they have filed through TapTax, and the next deadline.
| Name | Required | Description | Default |
|---|---|---|---|
| taxYear | No | Defaults to the current tax year |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds genuine behavior beyond that: the data is scoped to the signed-in person only, values are 'counted by the filing rule' (a computation caveat), and it discloses that undecided transactions are surfaced as a count. It does not mention caching, latency, or cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense but well-structured sentence, front-loaded with the identity of the result ('the signed-in person's tax year so far') before the itemised contents. Every clause maps to a distinct returned field with no filler, though the comma-list is long enough that it reads as a run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must define the return payload itself, and it does so exhaustively: income, expenses, profit, estimated income tax, Class 4 NI, undecided-transaction count, filed quarterly updates, and next deadline, plus the signing-rule caveat. Nothing material is missing for a single-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 100% and an enumerated taxYear parameter whose default is documented ('Defaults to the current tax year'), the schema fully carries parameter meaning. The description only implies a year scope via 'tax year so far' and adds no enum values, format, or behavior for out-of-range years, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (the signed-in person's tax position for the year) and enumerates the exact scope of that position: income, expenses, profit, tax and Class 4 NI, undecided transactions, filed quarterly updates, next deadline. It is highly concrete, but it never names or contrasts with the overlapping siblings estimate_self_employed_tax, check_mtd_status, or get_mtd_deadlines, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the description frames the tool as 'the TapTax dashboard view', which signals it is the aggregate overview rather than a point calculation. There is no explicit when-to-use, when-not-to-use, or alternative ('use X for a single estimate'), and the overlap with estimate_self_employed_tax and get_mtd_deadlines is left unresolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_invoicesList my invoicesARead-onlyInspect
The signed-in person's invoices, newest first, with each one's live status (a sent invoice past its due date reads as overdue) and the totals outstanding and overdue.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No | unpaid = sent, viewed or overdue | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the read-only safety profile, but the description meaningfully adds behavior: results are ordered newest first, status is computed live (a sent invoice past due reads as overdue), and totals outstanding/overdue are returned. These are real behavioral traits not derivable from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-formed sentence that front-loads what is returned and packs ordering, computed status, and totals without a wasted word.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of explaining the return payload (invoices, live status, outstanding/overdue totals), and annotations cover safety. It omits pagination/limit behavior and the per-invoice field list, leaving a minor gap for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the status enum is explained in the schema, but limit is undocumented anywhere. The description adds conceptual clarity around 'overdue' and 'live status' that touches the status filter, but it never addresses limit or the parameter surface directly, so it neither compensates fully nor ignores the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (the signed-in person's invoices) and implies the list verb clearly with scope ('newest first'). It is obviously distinct from create_invoice, send_invoice, and mark_invoice_paid, but it does not explicitly name or contrast with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or any mention of alternatives such as mark_invoice_paid or send_invoice. Usage is only inferable from the resource name; nothing tells the agent when this is the right choice over other invoice tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_transactionsList my transactionsARead-onlyInspect
Search the signed-in person's income and expense transactions, newest first. Filter by tax year or dates, type, text in the description, or needsReview=true for the ones still waiting for a business-or-personal decision (the review queue).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | On or before this date | |
| from | No | On or after this date | |
| type | No | ||
| limit | No | ||
| search | No | Text to find in the description | |
| taxYear | No | ||
| needsReview | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered by structured data. The description adds useful ordering context ('newest first') and explains what needsReview=true represents (the review queue). It does not disclose pagination behavior or default ordering interaction with limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and scope, then the filter options. No filler and every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description covers what is retrieved, ordering, and the filter semantics that matter most. It is slightly thin on default result size/pagination, but annotations carry the safety profile and the filter list is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43%, so the description must compensate, and it partly does by naming the filter dimensions (tax year, dates, type, description text, needsReview) and clarifying that needsReview means 'still waiting for a business-or-personal decision'. However, limit, the enum values, and default value semantics are left to the schema, leaving gaps in the 7-parameter set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (income/expense transactions) scoped to the signed-in person, and explicitly distinguishes itself from the mutation siblings (add_transaction, update_transaction) via the read-only 'signed-in person's' framing. An agent can identify this as the read/list counterpart to the write tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions for using it: filter by tax year, dates, type, text, or needsReview=true for the review queue. It implicitly routes agents needing to add or modify transactions to siblings, though it does not name them explicitly or state when NOT to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
log_mileage_tripLog a business tripAInspect
Record a business journey in the signed-in person's mileage log. TapTax prices it at the approved mileage rate for the tax year (tiered after 10,000 car miles). Commuting to a permanent workplace is not business mileage. Mileage tracking is on the Pro plan.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| date | Yes | ||
| from | Yes | ||
| miles | Yes | Business miles for this trip | |
| purpose | Yes | ||
| vehicle | No | Vans count as car | car |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the generic mutation profile (readOnly=false, idempotent=false, destructive=false); the description adds real behavioral context: automatic pricing at the approved rate, tiering after 10,000 car miles, and a Pro-plan gating requirement. It stops short of saying what the response contains or whether entries can be amended.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the action and then progressively adding pricing, the eligibility exclusion, and the plan gate. Every sentence carries distinct information, though the pricing/tiering detail is useful context rather than action-critical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation with no output schema, the description covers the semantic rules an agent most needs: what counts as business mileage, how it is priced, and the plan restriction. Remaining gaps (return value, permissions beyond plan gating) are moderate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (33%), so the description carries more weight than usual. Its mention of 'car miles' and the tiering rule indirectly signals that the vehicle type affects pricing, but it adds nothing about from/to/date/purpose, leaving most parameters documented only by their bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Record a business journey in the signed-in person's mileage log'), clearly distinct from the read-only sibling get_mileage_summary. An agent can identify this as the write-side mileage tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear exclusion rule ('Commuting to a permanent workplace is not business mileage') and a prerequisite ('Mileage tracking is on the Pro plan'). It doesn't explicitly name get_mileage_summary as the read counterpart, so routing guidance is good but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_invoice_paidMark an invoice paidADestructiveIdempotentInspect
Record that the client has paid an invoice (dated today). Stops any payment reminders.
| Name | Required | Description | Default |
|---|---|---|---|
| invoiceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and non-read-only, so the safety profile is covered. The description adds genuinely new behavior: the payment date is forced to today and payment reminders are stopped as a side effect — context an agent cannot derive from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste; the core action is front-loaded and the side effect follows. Nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation tool with annotations covering idempotency and destructiveness, the description supplies the two non-obvious facts (date defaults to today, reminders stop). Only minor gaps remain, such as authorization requirements or behavior when the invoice is already marked paid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single invoiceId parameter has no textual explanation, but its meaning is self-evident from the name and the format/pattern constraints in the schema. The description adds only an implicit reference to 'an invoice' rather than parameter-level semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record that the client has paid an invoice') with clear scope, and the sibling tools (create_invoice, send_invoice, list_invoices) are distinct enough that the action is unambiguous. It does not explicitly name a sibling to differentiate against, keeping it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this vs alternatives such as add_transaction, nor any prerequisites or exclusions. The mention of '(dated today)' implies a usage constraint but is stated as a side effect, not as usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_quarterly_updatePrepare my quarterly updateARead-onlyInspect
Get the signed-in person's Making Tax Digital quarterly update ready: the cumulative income, expenses and profit from 6 April to the end of the quarter, what still needs reviewing, and a card with a button that opens TapTax to check and file it. Filing itself always happens in the TapTax app, never from the chat, so HMRC receives the details of the device the person files from.
| Name | Required | Description | Default |
|---|---|---|---|
| quarter | No | 1 = to 5 July, 2 = to 5 October, 3 = to 5 January, 4 = to 5 April | |
| taxYear | No | Defaults to the current tax year |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavior: the output is a card with a button that opens TapTax, filing is out-of-band, and HMRC receives device details from the filing device. This goes well beyond the structured fields, though it doesn't discuss auth or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose followed by the filing constraint. Every clause earns its place, though the second sentence is dense and could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does so by naming the figures returned and the card/button artifact. Combined with the filing-location disclosure, an agent has what it needs, leaving only minor gaps around permissions or freshness of the data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both quarter and taxYear are fully documented in the schema (quarter maps to 5 July/5 Oct/5 Jan/5 Apr, taxYear defaults to current). The description's mention of '6 April to the end of the quarter' loosely reinforces the period semantics but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get the signed-in person's Making Tax Digital quarterly update ready') and then enumerates exactly what it produces: cumulative income, expenses, profit for the period, review items, and a TapTax card. This clearly separates it from siblings like check_mtd_status or get_mtd_deadlines, which report status/deadlines rather than assembling an update.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes a clear boundary — it prepares the update but filing 'always happens in the TapTax app, never from the chat' — which tells the agent not to expect to file. It lacks explicit routing guidance against alternatives (e.g. when to prefer check_mtd_status), so it is clear context without named exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_invoiceEmail an invoice to the clientADestructiveInspect
Email an invoice with its PDF to the client address on it, from the person's business name. This sends a real email to the client.
| Name | Required | Description | Default |
|---|---|---|---|
| invoiceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, destructive, non-idempotent, and open-world behavior, so the bar is lower. The description still adds concrete value: what is transmitted (the invoice plus its PDF), who receives it (the address stored on the invoice), and who it appears to come from (the person's business name), plus the explicit warning that a real email goes out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core action is front-loaded ahead of the sender/recipient detail. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema action whose annotations already communicate the safety profile, the description covers what is sent, to whom, and from whom. It omits failure conditions (e.g., invoice lacking a client email address) and whether the send can be undone, which would be minor additions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole parameter, invoiceId, is never explained in the description. The text implies an invoice (and its stored client address) is the subject of the call, but adds no format or identification detail beyond that implication; a single obvious parameter keeps this at an acceptable middle level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Email an invoice') and adds scope detail: the PDF, the client address on the invoice, and the sender identity. An agent can separate this from create_invoice, list_invoices, and mark_invoice_paid without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use framing or named alternative among the sibling tools. The sentence 'This sends a real email to the client' functions as an implicit caution that the tool should only be invoked when an actual client-facing send is intended, but the agent is left to infer when this is appropriate versus other invoice tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_transactionUpdate or review a transactionADestructiveIdempotentInspect
Change a transaction the signed-in person owns: mark it business or personal (this is how the review queue is cleared), or correct its category, description, amount or date.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| date | No | ||
| amount | No | ||
| category | No | ||
| description | No | ||
| businessOrPersonal | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=true, idempotentHint=true), and the description is consistent with them; it adds the ownership prerequisite, which is useful context. However, it never explains the destructive aspect or whether omitted fields are left unchanged, so it adds only modest value beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the action and scope front-loaded and the review-queue insight parenthetically attached; no filler or restated name/title content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers purpose, ownership restriction, mutable fields, and the review-queue workflow, while annotations carry the safety profile. The remaining gap is partial-update semantics (what happens to fields omitted from the call).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it names five of the six parameters (businessOrPersonal, category, description, amount, date), with id implied by the ownership constraint. It adds no format detail (date pattern, amount bounds, allowed category values), so it compensates well but not completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Change a transaction') tied to the resource and constrains it to records 'the signed-in person owns', then enumerates exactly which fields are mutable (business/personal flag, category, description, amount, date). This clearly separates it from add_transaction and list_transactions without needing the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies real usage context: marking business/personal 'is how the review queue is cleared', which tells the agent when this tool is the right choice. It stops short of naming alternatives or stating when not to use it, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
- First observed
add_transaction - First observed
check_mtd_status - First observed
check_tax_code - First observed
create_invoice - First observed
estimate_self_employed_tax - First observed
get_mileage_summary - First observed
get_mtd_deadlines - First observed
get_tax_position - First observed
list_invoices - First observed
list_transactions - First observed
log_mileage_trip - First observed
mark_invoice_paid - First observed
prepare_quarterly_update - First observed
send_invoice - First observed
update_transaction
Related MCP Connectors
Verified Canadian self-employed tax math: SE tax, CPP, GST/HST, instalments, CRA deadlines. Free.
Calculate sales tax & VAT, record transactions and refunds, manage products and customers.
2026 quarterly tax estimates, SE tax, safe harbor and mileage for 1099 gig workers.
Receipt tracker with no friction: receipts arrive by email and file themselves
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceCalculate income tax (UK/US brackets), EU VAT, UK corporation tax, and capital gains tax. Provides estimates only - not professional tax advice.5 npm30 PyPIMIT
- AlicenseNot gradedqualityCmaintenanceEnables 1099 gig workers and freelancers to calculate 2026 quarterly federal estimates, self-employment tax, safe-harbor minimum payments, and mileage or actual-expense deductions, including the mid-year business mileage rate change from 72.5 to 76 cents. It also compares standard mileage against actual vehicle costs and exposes every tax constant with its source.MIT
- AlicenseAqualityBmaintenanceLogs deductible driving trips, manages effective-dated mileage rates, computes per-category summaries, and exports CSV for accountants, all locally with no network calls.9MIT
- FlicenseNot gradedqualityBmaintenanceQuery current and historical UK official figures (tax bands, minimum wage, benefits, energy price cap and 100+ more) with effective dates and links to official government sources. Data refreshed whenever the official sources change.-
Glama MCP Gateway
Add one secure layer between your agents and this server.