Skip to main content
Glama

Server Details

Verifiable US tax oracle for AI agents: cited, machine-checkable federal and state tax computation

Status
Healthy
Last Tested
Transport
Streamable HTTP
URL

Glama MCP Gateway

Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.

MCP client
Glama
MCP server

Full call logging

Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.

Tool access control

Enable or disable individual tools per connector, so you decide what your agents can and cannot do.

Managed credentials

Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.

Usage analytics

See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.

100% free. Your data is private.
Tool DescriptionsA

Average 4.2/5 across 15 of 15 tools scored.

Server CoherenceA
Disambiguation5/5

Every tool targets a distinct tax concept or operation: individual vs. business vs. fiduciary tax, complete return computation, dependency analysis, rule lookups, and verification. There is no overlap in purpose, making tool selection unambiguous for an agent.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (e.g., calculate_tax, compute_return, verify_fact). The verb clearly indicates the action, and the noun the target. No mixing of conventions or vague verbs.

Tool Count5/5

With 15 tools, the server covers the core tax domain without being overwhelming. Each tool serves a specific, justified purpose, from calculation to verification to rule exploration. The count feels well-scoped for a comprehensive tax assistant.

Completeness4/5

The toolset covers individual, business, fiduciary taxes, federal and state returns, dependency, parameter lookups, rule search, and fact verification. Minor gaps exist (e.g., no explicit tool for payroll or gift tax), but the core workflows are complete and well-integrated.

Available Tools

15 tools
calculate_business_taxAInspect

Compute US federal BUSINESS-ENTITY tax from the same cited corpus: check-the-box entity classification, Form 1120 corporate income tax (§ 179/168(k)/174A/163(j)/DRD/NOL, § 250, GBC/FTC/BEAT), S-corp entity taxes, corporate estimates, the § 4501 buyback excise, AET and PHC taxes. Individual returns → calculate_tax. Unknown keys are rejected; unmodeled territory refuses loudly with the reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNoREQUIRED for computation: the law-in-force date — use the intended tax year's year-end (e.g. "2025-12-31" for TY2025). Omitting it is an error, never a default.
targetNorule to derive (default: us.federal.corp.entity_level_income_tax — the classification-aware entity income tax). Other targets: us.federal.corp.entity_classification, .taxable_income, .income_tax_after_credits, .beat, .estimated.quarterly_payment, .stock_buyback_excise, .accumulated_earnings_tax, .phc_tax, .s_corp_entity_taxes
corpNCTINoTY2025: the § 951A GILTI inclusion (with its § 78 gross-up); TY2026+: net CFC tested income (NCTI, OBBBA). The § 250 deduction applies 50% (2025) / 40% (2026+). The inclusion itself is GROSS INCOME (§ 951A(a)) — it must also be in corpGrossIncome with its § 78 gross-up; this fact drives only the § 250 deduction and FTC basket. The per-CFC tested-income aggregation is not modeled. In dollars.
corpFDDEINoTY2025: foreign-derived intangible income (FDII); TY2026+: foreign-derived deduction eligible income (FDDEI, OBBBA — QBAI abolished). The § 250 deduction applies 37.5% (2025) / 33.34% (2026+). In dollars.
corpPHCIncomeNoPersonal holding company income (§ 543: dividends, interest, royalties, annuities, certain rents). In dollars.
llcMemberCountNoNumber of members (owners) of the LLC — one member defaults to disregarded-entity treatment, two or more to partnership (Treas. Reg. § 301.7701-3(b)(1)).
qreCurrentYearNoQualified research expenses for the current year (§ 41(b); § 41(d) qualification attested). In dollars.
corpGrossIncomeNoThe corporation's gross income (§ 61), INCLUDING any dividends received, any § 951 subpart F and § 951A NCTI/GILTI inclusions with their § 78 gross-ups (§ 951A(a) is a gross-income INCLUSION — the § 250 deduction is computed separately from corpNCTI), and any § 245A-eligible foreign-sub dividends. In dollars.
corpIsREITorRICNoThe corporation is a real estate investment trust (§ 856) or regulated investment company (§ 851). REFUSES — their dividends-paid deduction and distribution requirements are not modeled.
entityLegalFormNoThe business's state-law legal form: a limited liability company, or a state-law corporation (a per-se corporation under Treas. Reg. § 301.7701-2(b)(1)).
corpCapitalGainsNoThe corporation's capital gains for the year (§ 1211(a): losses offset only these; net gain is ordinary-rate income for a corporation). In dollars.
corpPriorYearTaxNoTax shown on the corporation's preceding-year return (§ 6655(d) prior-year prong; unavailable if that year showed zero tax or was short). In dollars.
corpCapitalLossesNoThe corporation's capital losses for the year, including prior-year § 1212(a) carryovers being used (allowed only to the extent of capital gains). In dollars.
corpDividendsPaidNoDividends paid during the year (the § 561 dividends-paid deduction for the accumulated-earnings computation). In dollars.
corpTaxableIncomeNoThe C corporation's taxable income BEFORE the § 250 deduction, if already computed — used as-is when provided. Leave at 0 to have the engine compute it from corpGrossIncome and the deduction components (charitable/DRD/NOL machinery). CONTRACT: the § 250 deduction is computed separately from corpFDDEI/corpNCTI and subtracted by the tax rule — an AS-FILED Form 1120 line 30 already nets out § 250, so when providing corpFDDEI/corpNCTI enter the pre-§ 250 amount here (line 30 plus the § 250 deduction as filed), never the net. In dollars.
qreAvgPrior3YearsNoAverage annual qualified research expenses over the 3 preceding years — 0 means no prior QREs (the 6% startup rate of § 41(c)(4)(B) applies). In dollars.
corpSection179CostNoCost of § 179 property the corporation elects to expense — including qualified real property (roofs, HVAC, fire/security systems on nonresidential real property, § 179(d)(1)(B)(ii)) that § 168(k) cannot reach. Must NOT also be in corpEquipmentPurchases; EXCLUDE passenger automobiles and sport utility vehicles (the § 280F caps and the § 179(b)(5) SUV cap are not modeled). In dollars.
corpStockIssuedFMVNoFair market value of stock issued by the corporation during the taxable year (including to employees) — netted against repurchases under § 4501(c)(3). In dollars.
sCorpGrossReceiptsNoThe S corporation's gross receipts for the year (§ 1375). In dollars.
corpFiscalYearFilerNoThe corporation uses a FISCAL taxable year (or files a § 443 short-period return). Fiscal and short years are not modeled — the OBBBA parameters (§ 250 rates, § 59A 10.5%, § 960(d) 90%, § 448(c) $32M, the § 170(b)(2) 1% floor) apply by the taxable year's BEGINNING date, and § 443(b) annualization / § 15 proration are not encoded.
corpForeignResearchNoFOREIGN research or experimental expenditures paid this year — capitalized and amortized over 15 years (§ 174; first-year deduction is 1/30 under the midpoint convention). In dollars.
corpNOLCarryforwardNoNet operating loss carryforward available this year (§ 172: deduction limited to 80% of taxable income before the NOL; post-TCJA, no carrybacks). In dollars.
employeeAnnualWagesNoOne employee's annual wages, for the employer-side payroll-tax target (§ 3111 FICA + FUTA). In dollars.
corpDomesticResearchNoDomestic research or experimental expenditures — currently deductible under § 174A (OBBBA, permanent from 2025). In dollars.
corpDrdOwnershipTierNoOwnership of the dividend-paying corporation: under 20% (50% DRD), 20–80% (65% DRD), or 80%+ affiliated (100% DRD, § 243(a)(3)).
corpForeignTaxesNCTINoForeign taxes attributable to the § 951A basket (GILTI/NCTI) — the § 960(d) deemed-paid credit takes the 80% (2025) / 90% (2026+, OBBBA) allowance, no carryovers. In dollars.
corpTIThrough3MonthsNoCorporate taxable income for the first 3 months (§ 6655(e) annualization, installments 1-2). In dollars.
corpTIThrough6MonthsNoCorporate taxable income for the first 6 months (§ 6655(e) annualization, installment 3). In dollars.
corpTIThrough9MonthsNoCorporate taxable income for the first 9 months (§ 6655(e) annualization, installment 4). In dollars.
corpDividendsReceivedNoDividends received from other taxable domestic corporations (§ 243 DRD; must also be included in corpGrossIncome). Enter only dividends on stock meeting the § 246(c) holding period (held more than 45 days during the 91-day window around the ex-dividend date; 90/181 for certain preferred) — attested; § 1059 extraordinary-dividend basis reduction not modeled. In dollars.
sCorpNetPassiveIncomeNoPassive investment income net of directly-connected deductions (§ 1375(b)(2)). In dollars.
sCorpShareholderCountNoNumber of shareholders, counting married couples and § 1361(c)(1) family members as one (§ 1361(b)(1)(A): may not exceed 100).
sCorpTaxableIncomeAsCNoThe S corporation's taxable income computed as if it were a C corporation (§§ 1374(b)(1)/1375(b)(1)(B) cap). In dollars.
corpBaseErosionTestMetNoThe corporation's base erosion percentage is 3% or more (2% for banks/securities dealers) — one of the two § 59A applicable-taxpayer tests. BEAT applies only to $500M+ multinationals.
corpEquipmentPurchasesNoCost of qualified § 168(k) property acquired AND placed in service this year (acquired after January 19, 2025 — 100% bonus depreciation, OBBBA-permanent). EXCLUDE passenger automobiles (the § 280F luxury-auto caps are not modeled) and anything entered in corpSection179Cost. In dollars.
corpIsLargeCorporationNoThe corporation had taxable income of $1,000,000 or more in any of the 3 preceding taxable years (§ 6655(g)(2) 'large corporation' — may not use the prior-year safe harbor).
corpOrdinaryDeductionsNoOrdinary business deductions (salaries, rents, prior-year amortization, …) — everything EXCEPT charitable contributions, the dividends-received deduction, and NOLs, which have their own limited rules. Enter compensation already limited by § 162(m) (no deduction for a covered employee's remuneration over $1,000,000 at a publicly held corporation — not modeled, attested). In dollars.
corpOwnedByFiveOrFewerNoMore than 50% of the stock's value was owned (directly or via § 544 attribution) by 5 or fewer individuals during the last half of the year (§ 542(a)(2)).
filedForm2553SElectionNoThe entity filed a timely Form 2553 S election under § 1362(a)(1) (for an eligible entity this also deems association classification, Reg. § 301.7701-3(c)(1)(v)(C)).
generalBusinessCreditsNoAggregate current-year § 38(b) general business credits (e.g. the § 41 research credit target's result) — limited under § 38(c). In dollars.
corpAvgGrossReceipts3yrNo3-year-average annual gross receipts (§ 448(c) test: $31M for 2025, $32M for 2026 — at or below it the § 163(j) limit does not apply). In dollars.
corpForeignTaxesGeneralNoCreditable foreign income taxes in the § 904(d) GENERAL basket. In dollars.
corpStockRepurchasedFMVNoFair market value of the corporation's own stock repurchased (§ 317(b) redemptions and economically similar transactions) during the taxable year, for the § 4501 excise. In dollars.
corpIsCoveredCorporationNoThe corporation is a 'covered corporation' for the § 4501 stock-repurchase excise tax: a domestic corporation whose stock is traded on an established securities market (§ 4501(b)).
corpSection245ADividendsNoForeign-source portion of dividends received from specified 10-percent-owned foreign corporations, eligible for the § 245A participation-exemption DRD (100%). Must also be included in corpGrossIncome. Attested by entry: US-shareholder status, NOT a § 245A(e) hybrid dividend, and the § 246(c)(5) 365-day holding period met; no foreign tax credit is allowed for the deducted portion (§ 245A(d)) — keep these taxes out of the FTC inputs. In dollars.
sCorpHasAccumulatedEandPNoThe S corporation has accumulated earnings and profits from C-corporation years at the close of the year (§ 1375 applies only then).
corpAccumulatedEandPStartNoAccumulated earnings and profits at the close of the PRECEDING year (§ 535(c)(2) minimum-credit offset). In dollars.
corpIsPersonalServiceCorpNoThe corporation's principal function is services in health, law, engineering, architecture, accounting, actuarial science, performing arts, or consulting (§ 535(c)(2)(B): $150,000 minimum credit instead of $250,000).
filedForm8832CorpElectionNoThe entity filed a Form 8832 election to be classified as an association taxable as a corporation (Treas. Reg. § 301.7701-3(c)).
corpBaseErosionTaxBenefitsNoBase erosion tax benefits for the year (§ 59A(c)(2)) — include the base-erosion percentage of any NOL deduction (§ 59A(c)(1)(B)). Added back to reach modified taxable income. In dollars.
corpUndistributedPHCIncomeNoUndistributed personal holding company income (§ 545: taxable income adjusted, less federal taxes and the dividends-paid deduction). In dollars.
sCorpRecognizedBuiltInGainNoNet recognized built-in gain during the § 1374(d)(7) 5-year recognition period after a C-to-S conversion (0 if the period has passed or there was no conversion). In dollars.
corpBusinessInterestExpenseNoBusiness interest expense (§ 163(j): limited to 30% of EBITDA-based ATI unless the § 448(c) gross-receipts test is met). In dollars.
corpCharitableContributionsNoThe corporation's charitable contributions — current-year gifts plus allowable prior-year § 170(d)(2) carryovers being used (both subject to the same ceiling and, from 2026, the OBBBA floor). In dollars.
corpFilesConsolidatedReturnNoThe corporation joins a consolidated return (§§ 1501-1504) — intercompany eliminations and SRLY rules are not modeled, so this refuses.
corpReasonableNeedsRetentionNoEarnings retained for the reasonable needs of the business (§§ 535(c)(1), 537 — documented needs; part of the accumulated earnings credit). In dollars.
sCorpHasMultipleStockClassesNoThe corporation has more than one class of stock (§ 1361(b)(1)(D); differences in voting rights alone do not create a second class, § 1361(c)(4)).
sCorpPassiveInvestmentIncomeNoThe S corporation's passive investment income — royalties, rents, dividends, interest, annuities (§ 1375(b)(3)). In dollars.
sCorpHasIneligibleShareholderNoAny shareholder is ineligible under § 1361(b)(1)(B)–(C): a nonresident alien, or an entity other than an estate or eligible trust/exempt organization.
corpForeignSourceIncomeGeneralNoForeign-source taxable income in the general basket (§ 904 limitation numerator; § 861 expense allocation attested). In dollars.
corpAdjustedOrdinaryGrossIncomeNoAdjusted ordinary gross income (§ 543(b)(2)) — the 60% test base. In dollars.
corpPortfolioDebtFinancedPercentNoAverage indebtedness percentage (0-100) of debt-financed portfolio stock (§ 246A) — reduces the 50%/65% DRD proportionally; 0 = not debt-financed.
corpAvgAdjustedFinancialStatementIncomeNo3-year-average adjusted financial statement income (§ 56A) — over $1 billion triggers the corporate AMT, which this engine refuses to approximate. In dollars.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries full burden. It discloses that unknown keys are rejected and unmodeled territory refuses with a reason, which is helpful behavioral information. However, it does not mention whether the computation is read-only or has side effects. For a compute tool, this is reasonable, slightly above average.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads the main function and includes a clear directive for individual returns. It is concise given the tool's complexity, though the list of tax items could be overwhelming. Efficient but slightly noisy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (63 parameters, no output schema), the description provides a good overview of what taxes are computed but does not explain the output format or return value. This leaves some ambiguity for the agent. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description itself does not add parameter-level detail beyond listing categories. The mention of 'unknown keys are rejected' reinforces schema validation but adds minimal value. Score is at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it computes US federal business-entity tax, listing specific forms and components (Form 1120, S-corp, buyback excise, etc.). It also distinguishes from the sibling tool calculate_tax by directing individual returns there, providing a clear resource and differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use the alternative tool (calculate_tax for individual returns) and warns that unknown keys are rejected and unmodeled territory refuses with reason. This gives clear guidance on both when to use and what to expect on input validation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calculate_fiduciary_taxAInspect

Compute US federal income tax for an ESTATE or TRUST (Form 1041): the § 1(e) compressed brackets and § 642(b) exemption. Input is taxable income before the exemption, after the §§ 651/661 distribution deduction. Retained capital gains refuse loudly (§ 1(h) trust breakpoints not modeled). Grantor trusts belong on the grantor's individual return via calculate_tax.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNoREQUIRED for computation: the law-in-force date — use the intended tax year's year-end (e.g. "2025-12-31" for TY2025). Omitting it is an error, never a default.
targetNorule to derive (default: us.federal.fiduciary.income_tax).
fiduciaryTypeNoForm 1041 filer type for the § 642(b) exemption: estate ($600), simple trust required to distribute all income currently ($300), or complex trust ($100). Grantor trusts do not file their own tax — use the grantor's individual return.
fiduciaryLongTermGainsNoNet long-term capital gain retained by the estate/trust. Any positive amount REFUSES — the § 1(h) preferential breakpoints for estates and trusts are not modeled. In dollars.
fiduciaryIncomeBeforeExemptionNoThe estate/trust's taxable income BEFORE the § 642(b) exemption but AFTER the §§ 651/661 income-distribution deduction (the DNI machinery is attested by this input). In dollars.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description fully discloses key behaviors: long-term gains cause failure ('refuse loudly'), missing asOf is an error, and grantor trusts are out of scope. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, extremely concise and front-loaded. Every word adds value, no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of trust taxation and lack of output schema, the description covers purpose, inputs, and limitations well. Minor gap: no explicit mention of what the function returns (presumably a tax amount), but overall complete enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline 3. Description adds value by explaining purpose of each parameter (e.g., fiduciaryType for exemption, fiduciaryLongTermGains refusal), going beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it computes US federal income tax for estates or trusts (Form 1041), mentioning specific sections and exemptions, and distinguishes from sibling tools like calculate_tax (for individuals) and calculate_business_tax.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use ('Compute US federal income tax for an ESTATE or TRUST') and when not to use ('Grantor trusts belong on the grantor's individual return via calculate_tax'). Also warns that retained long-term capital gains cause refusal, providing clear usage boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calculate_taxAInspect

Compute US federal INDIVIDUAL income tax (or balance due if withholding is given) from a content-addressed corpus of cited rules. NEVER estimate tax yourself — call this, and report ONLY numbers returned by oracle calls made with the real facts (never hand-check or approximate a line the oracle can compute: your recalled parameters may be stale). Negative result = refund. Returns the answer, every assumption made, and hashes that let anyone re-verify the full derivation offline. Facts are grouped (filing, income, retirement, credits, …) — fill the groups that apply; unknown keys are rejected, and the engine names any missing fact the target needs. When source documents CONFLICT on a value, do not silently pick one: compute both branches, disclose the conflict and your choice; an interview/confirmation answer (rollover, conversion, taxable-amount screens) usually reflects taxpayer intent better than a payer form's box code — prefer it and disclose. That heuristic covers FACTS only: LEGAL classifications (qualifying child vs other dependent, filing status, SSTB) follow the statute's tests, not intake checkbox labels — a generic 'claim dependent credit' flag does not convert a qualifying child into an ODC dependent. TRANSCRIBE documented amounts as given even when they look anomalous (e.g. state withholding in a no-income-tax state): disclose the anomaly, never delete or 'correct' a documented number from outside knowledge. If you believe an oracle result is wrong, report the ORACLE's number and note your dissent — never substitute your own: the corpus is primary-source-verified and your recollection is not. Business entities → calculate_business_tax; estates/trusts → calculate_fiduciary_tax; § 152 dependency → determine_dependent.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNoREQUIRED for computation: the law-in-force date — use the intended tax year's year-end (e.g. "2025-12-31" for TY2025). Omitting it is an error, never a default.
stateNostate taxable income for the state tax targets (us.ca/us.va/us.il.income_tax; parameters via lookup_tax_parameter)
filingNowho is filing: status, age/blindness, dependency, student status
incomeNowages, interest, capital gains, unemployment, foreign earned income
targetNorule to derive (default: net tax; balance due when payments_estimates.federalTaxWithheld is given). Determinations: us.federal.eligible.tips_deduction, us.federal.estimated.quarterly_payment, us.federal.estimated.safe_harbor_met
creditsNoCTC/ODC counts, dependent care, saver's, adoption, education
itemizedNoSchedule A: SALT, mortgage, medical, charitable
documentsNoRAW document transcription (preferred over hand-mapped facts): W-2 boxes, 1099-R boxes/codes, SSA-1099 boxes, dependents' birth dates. SSA-1099s are first-class: box 5 sums into socialSecurityBenefits and box 6 into withholding, so the § 86 taxable-benefits worksheet runs on the transcribed total instead of a hand-mapped guess. The tool derives wages/withholding (incl. Form 8959 Part IV), box-3/5 wage coordination, dependent classifications, age facts, and early-distribution penalties deterministically — and errors if the same value is also passed as a hand-mapped fact.
kiddie_taxNoForm 8615 inputs for a child subject to § 1(g)
retirementNosocial security, IRA/pension distributions, early-distribution penalty
adjustmentsNoIRA/HSA contributions, student-loan and car-loan interest
investor_amtNoAMT preferences (ISO spread) and § 1202 QSBS exclusion
tips_overtimeNo§ 224 tips and § 225 overtime deductions (OBBBA)
healthcare_ptcNo§ 36B premium tax credit / Form 1095-A reconciliation
rentals_passiveNoSchedule E rentals/royalties + § 469 passive-loss netting (Form 8582 via us.federal.passive_loss_allowed)
self_employmentNoSchedule C / K-1, QBI inputs, SE deductions, home office
household_employerNoSchedule H nanny/household-employee taxes
payments_estimatesNowithholding, prior-year safe harbor, annualized installments
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavior: it computes from a cited corpus, never estimates, returns answer with assumptions and verifiable hashes, negative result means refund, and details handling of conflicts, transcription rules, and heuristics. All behavioral traits are clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very long and detailed, containing multiple paragraphs and inline examples. While front-loaded with purpose, it includes extensive guidance that could be more concise. However, the complexity of the tool may justify the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description only briefly mentions the return value ('answer, every assumption made, and hashes') without specifying structure. It covers input complexity well but leaves output format vague, which is a gap for such a complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds overall context but does not delve into individual parameter semantics beyond what the schema already provides. The rich schema descriptions carry the parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it computes US federal individual income tax and specifically differentiates from sibling tools (calculate_business_tax, calculate_fiduciary_tax, determine_dependent). It uses a specific verb+resource and clarifies scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides extensive guidance: never estimate tax yourself, call this tool; when to use alternatives (business entities, estates, dependency); how to handle conflicts, anomalies, and transcriptions; and explicit do's and don'ts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_filing_statusesAInspect

Compute the answer under every filing status for the same facts — e.g. to answer 'should we file jointly or separately?'. Statuses that need more facts report their error instead of guessing.

ParametersJSON Schema
NameRequiredDescriptionDefault
factsYesfacts for the computation: either flat corpus fact ids (see list_input_facts) or the same group objects calculate_tax accepts (filing, income, retirement, …), plus optional target and asOf. Business/fiduciary/dependent facts are accepted flat. Unknown keys are rejected by name — nothing is ever silently dropped.
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries full burden. It discloses that the tool reports errors for missing facts rather than guessing, and notes that unknown keys are rejected. However, it does not mention destructive actions, authorization requirements, or output behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences that front-load the purpose and an example. Every word adds value, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description adequately explains the tool's core function, it lacks details about the output format (no output schema exists). The tool likely returns results per status, but this is not stated, leaving the agent to infer. Additionally, the handling of missing facts is addressed but other edge cases are not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for the single 'facts' parameter. The tool description adds no additional parameter details beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Compute the answer under every filing status for the same facts' with a concrete example ('should we file jointly or separately?'). This distinguishes it from sibling tools like calculate_tax, which compute under a single status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context on when to use the tool (e.g., comparing filing statuses) and mentions that statuses needing more facts report an error. However, it lacks explicit guidance on when not to use it (e.g., when a single status calculation suffices) and does not name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compute_returnAInspect

Compute the COMPLETE Form 1040 bottom-line set in one call — the 17 lines that determine the return (1a, 9, 10, 11, 12e, 15, 16, 17 AMT, 19, 22, 23, 24, 25d, 27a, 28, 32, 33, 34/37), each whole-dollar rounded by the engine. Takes the SAME input as calculate_tax (prefer the documents block: transcribe W-2/1099-R/SSA-1099 boxes and dependent birth dates — SSA-1099s are first-class, box 5 and box 6 are summed for you; the tool derives ages, classifications, Part IV withholding, and penalties deterministically). Never assemble return lines by hand — this tool is the return. TRANSCRIPTION CONVENTIONS: (1) a PRIOR-YEAR Form 1040 in the file supplies CONTINUING conditions the current-year interview omits — the 'Someone can claim: You as a dependent' checkbox and the blindness boxes carry forward unless the current-year data contradicts them; (2) COMMUNITY PROPERTY: do NOT split income 50/50 between MFS spouses when they lived apart all year with no transfers (§ 66(a) allocates earned income to the earner) or when a written separation agreement ended the community — transcribe each document to its named earner; never invent a Form 8958 split the intake does not request; (3) bonus depreciation for assets placed in service 1/1-1/19/2025 is 40% (§ 168(k) pre-OBBBA phase-down; 100% only for property ACQUIRED after 1/19/2025).

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNoREQUIRED for computation: the law-in-force date — use the intended tax year's year-end (e.g. "2025-12-31" for TY2025). Omitting it is an error, never a default.
stateNostate taxable income for the state tax targets (us.ca/us.va/us.il.income_tax; parameters via lookup_tax_parameter)
filingNowho is filing: status, age/blindness, dependency, student status
incomeNowages, interest, capital gains, unemployment, foreign earned income
targetNorule to derive (default: net tax; balance due when payments_estimates.federalTaxWithheld is given). Determinations: us.federal.eligible.tips_deduction, us.federal.estimated.quarterly_payment, us.federal.estimated.safe_harbor_met
creditsNoCTC/ODC counts, dependent care, saver's, adoption, education
itemizedNoSchedule A: SALT, mortgage, medical, charitable
documentsNoRAW document transcription (preferred over hand-mapped facts): W-2 boxes, 1099-R boxes/codes, SSA-1099 boxes, dependents' birth dates. SSA-1099s are first-class: box 5 sums into socialSecurityBenefits and box 6 into withholding, so the § 86 taxable-benefits worksheet runs on the transcribed total instead of a hand-mapped guess. The tool derives wages/withholding (incl. Form 8959 Part IV), box-3/5 wage coordination, dependent classifications, age facts, and early-distribution penalties deterministically — and errors if the same value is also passed as a hand-mapped fact.
kiddie_taxNoForm 8615 inputs for a child subject to § 1(g)
retirementNosocial security, IRA/pension distributions, early-distribution penalty
adjustmentsNoIRA/HSA contributions, student-loan and car-loan interest
investor_amtNoAMT preferences (ISO spread) and § 1202 QSBS exclusion
tips_overtimeNo§ 224 tips and § 225 overtime deductions (OBBBA)
healthcare_ptcNo§ 36B premium tax credit / Form 1095-A reconciliation
rentals_passiveNoSchedule E rentals/royalties + § 469 passive-loss netting (Form 8582 via us.federal.passive_loss_allowed)
self_employmentNoSchedule C / K-1, QBI inputs, SE deductions, home office
household_employerNoSchedule H nanny/household-employee taxes
payments_estimatesNowithholding, prior-year safe harbor, annualized installments
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description must fully disclose behavior. It does so thoroughly: it describes rounding, deterministic derivation of ages and classifications, first-class handling of SSA-1099s, and key conventions. It also warns about errors like omitting asOf. However, it does not explicitly state whether the tool is read-only (it logically is, but not stated), and some behavioral details could be more structured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough but overly verbose. The top-level description and many parameter descriptions could be condensed. While every sentence adds value, the sheer length and repeated emphasis on legal details reduce conciseness. Front-loading the core purpose helps, but the description still feels dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the extreme complexity of the tool (18 parameters, nested objects), the description is very complete. It covers edge cases, legal references, and usage conventions. However, the absence of an output schema means the agent must infer the return format from the description, which only lists line numbers but not their mapping to output fields. This slight gap prevents a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. However, the descriptions of parameters add immense value beyond the schema by providing legal citations, examples, constraints, and practical usage notes. For instance, the wages parameter includes format examples, and many state-specific parameters include cross-references and computation rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool computes the complete Form 1040 bottom-line set in one call, listing the 17 key lines. It distinguishes itself from the sibling calculate_tax by noting it takes the same input but handles the full return, saying 'Never assemble return lines by hand — this tool is the return.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear when-to-use guidance: it is for the full return computation, not for manual assembly. It references calculate_tax for tax calculation only. Additionally, it includes detailed transcription conventions (prior-year Form 1040, community property, bonus depreciation) that govern usage, and explicitly states that omitting asOf is an error.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compute_state_returnAInspect

Compose a STATE return's printed-form line set deterministically (2025 IL-1040 / VA 760 / CA 540 / NY IT-201 / PA-40) — correct line NUMBERS from the printed forms and whole-dollar rounding, with the state tax computed by the oracle targets internally. PA is CLASS-BASED: transcribe the pa* class fields (Box 16 compensation, per-spouse loss classes) — federalAGI is NOT the PA base; the composer runs the class netting, Schedule O, and Tax Forgiveness targets itself, and reports the WPTC as a note (no printed line). Workflow: run compute_return first for the federal substrate, compute any state-specific components the citations describe (additions, subtractions, credits without targets — disclose each), then call this ONCE and report its line set VERBATIM. Never hand-assemble state line numbers: transposed lines on correct dollars are the dominant state error mode. ALWAYS pass taxableSocialSecurity and unemploymentCompensation when nonzero (VA/CA/NY subtractions are applied by the composer). ALWAYS transcribe the intake's state-specific block (e.g. ca_tax_return.ca_form540_schca: AB 5 employee-classification additions; va_sch_a fields; county/use-tax questions) — those fields drive composer inputs. For VA MFJ, pass vaYourVagi/vaSpouseVagi (the separate-VAGI worksheet) so the composer can run the Spouse Tax Adjustment worksheet itself.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfYesyear-end date, e.g. 2025-12-31 — REQUIRED
caAmtNoOVERRIDE for Form 540 line 61 — prefer caIsoPreference + caAmtTaxesAddback so the composer builds Schedule P AMTI and evaluates us.ca.amt itself; a passed caAmt wins.
wagesNofederal line 1a wages (NY IT-201 line 1)
caBhstNous.ca.bhst result (pass the oracle target's answer) — Form 540 line 62 Behavioral Health Services Tax (R&TC § 17043, 1% of CA taxable income over $1,000,000); added into line 64 total tax when nonzero.
caYCTCNous.ca.yctc result (pass the oracle target's answer)
useTaxNoconsumer use / sales-use tax owed on the return
additionsNototal state additions to federal AGI (e.g. NY 414(h) A-104 + IRC-125 A-101; VA Schedule ADJ line 2 codes). GATE RULE: coded addition/subtraction line-item arrays sitting under a false 'do you have additions/subtractions' boolean are inactive template rows (especially $1-$4 placeholder amounts) — transcribe $0 for them and disclose; the gate controls for these arrays
caCalEITCNous.ca.caleitc result (pass the oracle target's answer)
filingHohNo
dependentsNodependent count (CA dependent exemption credits; NY $1,000 exemptions)
exemptionsNopersonal + dependent exemption COUNT (self + spouse + dependents)
federalAGINofederal Form 1040 line 11 (from compute_return, verbatim). REQUIRED for il/va/ca/ny — the composer refuses without it. NOT used by PA (class-based: pass the pa* class fields instead).
paGamblingNoPA-40 line 8: gambling and lottery winnings net of wager costs (noncash PA Lottery prizes exempt; cash prizes taxable)
paInterestNoPA-40 line 2: PA-taxable interest (gross class — no expenses; includes commercial-annuity interest taxable as PA interest)
vaYourVagiNoPRIMARY taxpayer's separate VAGI (MFJ only — the 760 instructions' 'Worksheet for Determining Separate Virginia Adjusted Gross Income': own wages/SE/pensions, own share of joint items 50/50, own age deduction and subtractions). Enables the composer's Spouse Tax Adjustment worksheet (Form 760 line 17). vaYourVagi + vaSpouseVagi must equal line 9 VAGI.
federalEITCNofederal EIC, line 27a (from compute_return)
filingJointNo
paDividendsNoPA-40 line 3: PA-taxable dividends INCLUDING mutual-fund capital-gain distributions (PA classifies them as dividends, not gains)
vaItemizingNotaxpayer itemized federally (VA requires the same election, Va. Code § 58.1-322.03(1)) — enables the VA Schedule A computation from the component inputs below; Form 760 line 10 replaces the line 11 standard deduction
filingStatusNoREQUIRED in practice: the federal filing status — drives the state bracket schedule, standard deduction column, and exemption structure. The filingJoint/filingHoh/filingHohOrQss booleans are legacy aliases; when filingStatus is present it wins.
jurisdictionYes
subtractionsNototal state subtractions OTHER than the automatic ones (taxable social security / unemployment have their own inputs below; e.g. NY S-136 alimony paid, IL retirement subtraction)
vaFamilyVagiNoSchedule ADJ line 10 total family VAGI (you + spouse + dependents' VAGI) for the Credit for Low-Income Individuals poverty test; defaults to line 9 VAGI when omitted
vaSpouseVagiNospouse's separate VAGI for the STA worksheet (Form 760 line 17 box; MFJ only)
ilK12ExpensesNoIL qualified K-12 education expenses (before the $250 floor)
paBusinessNetNoPA-40 line 4, TAXPAYER's own net business/profession/farm income or LOSS (negative allowed; within-class netting of the taxpayer's own activities only — a loss never crosses classes or spouses)
paEstateTrustNoPA-40 line 7: estate/trust income (PA Schedule J; an estate or trust cannot distribute a loss — never negative)
paPropertyNetNoPA-40 line 5, taxpayer's own net gain/loss from sale/exchange/disposition of property (negative allowed; no carryover)
caHsaDeductionNofederal HSA deduction (Form 8889 line 13) — California does not conform to § 223: the composer ADDS it back for CA
filingHohOrQssNo
ilChildUnder12NoIL CTC gate: a QUALIFYING CHILD (§ 152(c) lineage — child/stepchild/foster/sibling or their descendants) under age 12. A qualifying-relative/ODC-only dependent does NOT satisfy this even if under 12; leave false.
ilEitcOverrideNous.il.eitc oracle target's answer (35 ILCS 5/212(a)(vi), (b-5), (b-10): 20% of the federal EITC recomputed WITHOUT the § 32(c)(1)(A)(ii) childless age gate) — WINS over the generic 20%-of-federalEITC line-29 computation when present. MUST be used (not merely optional) for a taxpayer age 18-24 or 65+ with NO qualifying children: federalEITC alone is correctly $0 for that population under federal law, so line 29 = 20% x federalEITC would wrongly zero out Illinois' decoupled credit — pass us.il.eitc's computed answer instead. Safe to pass for every IL EITC claimant (agrees with the generic computation outside the decoupled population).
vaAgeDeductionNoOVERRIDE ONLY — pass vaAgeQualifyingFull/vaAgeQualifyingTested instead and the composer computes the age deduction itself (including the AFAGI social-security exclusion agents routinely miss). When splitting an odd joint total between spouses (Form 760 lines 4a/4b), the odd dollar goes to the SPOUSE.
ageOrBlindBoxesNocount of age-65+/blind boxes checked (taxpayer/spouse, per box)
caIsoPreferenceNoISO exercise spread AMT preference (§ 56(b)(3) as modified by R&TC § 17062) — with caAmtTaxesAddback this lets the composer BUILD Schedule P AMTI itself (AMTI = line 19 taxable income + taxes deducted in the CA itemized deduction + this preference; standard deduction added back instead when not itemizing) and evaluate us.ca.amt internally; caAmt (a precomputed answer) wins if both are given
caRentersCreditNous.ca.renters_credit result (pass the oracle target's answer) — nonrefundable, joins the exemption credits in the line-48 subtraction from tax.
cityWithholdingNoNY line 73 NYC withholding
vaItemizedOtherNoVA Sch A other itemized deductions
extensionPaymentNopayment made with an extension request. VA 760 line 22 (its own line, never folded into the estimated-payments line). NY IT-201 line 75 is the COMBINED line — 'estimated tax payments and amount paid with Form IT-370' — so for NY this is added into the same line as estimatedPayments, not kept separate.
nycTaxableIncomeNoNYC taxable income (IT-201 line 47) if NYC resident
paRentRoyaltyNetNoPA-40 line 6, taxpayer's own net rents/royalties/patents/copyrights (short-term rentals under 30 days are BUSINESS income, line 4)
paResidentCreditNoPA-40 line 22: resident credit for tax paid other states (Schedule G-L; not for reciprocal-state compensation: IN/MD/NJ/OH/VA/WV). Subtracts BEFORE Tax Forgiveness — the composer handles the ordering.
stateWithholdingNostate income tax withheld (IL line 25 / VA 19a / CA 71 / NY 72). CONVENTIONS: IL line 25 sums state withholding from EVERY document (W-2s + all 1099s). NY line 72 = W-2 box 17 NYS withholding PLUS NY-coded state withholding from 1099s whose PAYER has an in-state (NY) address; NY-coded withholding printed by an OUT-OF-STATE-addressed payer is NOT included; disclose any excluded amount in notes. VA 19a = the PRIMARY taxpayer's withholding from EVERY document type (W-2, 1099, VK-1 — Form 760 line 19 instructions name all three; the payer's address does NOT matter for VA, unlike NY); a jointly-issued document's state withholding splits 50/50 between 19a/19b with the odd dollar to the primary.
vaRefundableEitcNoOVERRIDE ONLY — the composer now computes the Form 760 line 23 credit itself from federalEITC + the eligibility inputs below. If passed, this refundable amount wins over the computed selection.
yonkersSurchargeNous.ny.yonkers_surcharge result (pass the oracle target's answer) — 16.75% of the Yonkers worksheet's netted base (nyYonkersBase). Added into line 62's total and printed on its own line (IT-201 LINE 55, not 54 — line 54 is MCTMT) when nonzero.
caAmtTaxesAddbackNotaxes actually included in the CA itemized deduction (property taxes etc. surviving the Schedule CA SALT adjustments) — the Schedule P line 2 addback used when the composer builds AMTI from caIsoPreference; $0 when not itemizing (the composer adds back the standard deduction instead)
estimatedPaymentsNostate estimated payments ONLY (extension payments and prior-year credited overpayments have their own lines where the form provides them)
ilPropertyTaxPaidNoIL property tax on principal residence, net of business-use portion
ilTeacherExpensesNoIL Schedule 1299-C educator materials expenses
nyHouseholdCreditNoNYS household credit from table 2 (us.ny.parameters citation)
paNrk1WithholdingNoPA-40 line 17: nonresident tax withheld from PA Schedule(s) NRK-1
refundableCreditsNostate refundable credits, e.g. the NY credit block: ESCC + NYS EIC + IT-216 + NYC EIC + NYC school tax + NYC child care (WITHOUT their own oracle target, self-computed per the us.ny.parameters citation and disclosed) PLUS us.ny.it214 (the Real Property Tax Credit, which DOES have an oracle target as of TY2025 v5 — pass its computed answer here, not a hand-derived percentage of rent)
vaItemizedMedicalNoVA Sch A line 1: total medical/dental expenses BEFORE any floor (VA applies its own 10%-of-FAGI floor — Virginia deconforms from the federal 7.5% floor)
claimedAsDependentNosomeone else can claim this taxpayer as a dependent (carry the prior-year 1040 'Someone can claim: You as a dependent' checkbox forward as a continuing condition unless the current-year interview contradicts it). IL: zeroes the line 10 exemption allowance when base income exceeds the exemption amount. VA: limits the standard deduction to earned income.
nycHouseholdCreditNoNYC household credit from table 5
pa529ContributionsNoSchedule O code T: § 529 contributions, ALREADY capped at $19,000 per beneficiary per taxpayer-spouse (2025); no deduction for rollovers/beneficiary changes
paScheduleDcCreditNoPA-40 line 23 component: the Child and Dependent Care Enhancement credit — pass us.pa.cdcc's computed answer (= 100% of the federal Form 2441 line 9a tentative credit; refundable)
vaItemizedCasualtyNoVA Sch A casualty/theft losses (protected from the overall limitation)
vaItemizedGamblingNoVA Sch A gambling losses (§ 165(d), limited to winnings; protected from the overall limitation)
yonkersWithholdingNoNY line 74 Yonkers withholding (W-2 box 19 with a Yonkers locality)
paAbleContributionsNoSchedule O code A: PA ABLE contributions, capped at the federal gift-tax exclusion ($19,000 for 2025)
paGrossCompensationNoPA-40 line 1a: W-2 BOX 16 total (NOT Box 1 — 401(k)/elective deferrals are PA-taxable; eligible retirement distributions are exempt and excluded). Falls back to the shared wages input when omitted (composer discloses). Include taxable early-distribution amounts under the cost-recovery method.
paPenaltiesInterestNoPA-40 line 27: penalties and interest incl. estimated-underpayment penalty (REV-1630)
paScheduleOcCreditsNoPA-40 line 23 component: Schedule OC restricted credits total (transcribed; no oracle target)
paSpouseBusinessNetNoPA-40 line 4, SPOUSE's own net business income or loss (kept separate: PA never nets one spouse's loss against the other's income)
paSpousePropertyNetNoPA-40 line 5, spouse's own net property gain/loss
vaAgeQualifyingFullNocount of filers (taxpayer/spouse) born ON OR BEFORE January 1, 1939 — each gets the UNCONDITIONAL $12,000 age deduction (no income test)
vaYourAgeBlindBoxesNoSTA worksheet Part 1 line 2: PRIMARY taxpayer's 65+/blind box count (0-2) — per-spouse exemption = boxes x $800 + $930
caAb5NetLossAdditionNonet losses from businesses where the worker is an employee for California (intake ca_form540_schca.add_net_loss) — the federal Schedule C loss is disallowed for CA: Schedule CA BUSINESS addition, col C
caItemizedDeductionsNoCA itemized deduction total (Schedule CA Part II, line 29) — agent-computed per Schedule CA's own itemized rules WITH disclosure (differs from the federal Schedule A: no SALT cap, mortgage/medical add-backs, etc.). Form 540 line 18 takes the GREATER of this or the CA standard deduction; omit to use the standard deduction only.
nonrefundableCreditsNostate NONREFUNDABLE credits without an oracle target — capped at the state tax due by the composer (an excess never creates a refund). VA: do NOT put the low-income credit or any VA EITC election here — passing it forces a legacy capped path; instead pass federalEITC (+ vaFamilyVagi if testing the low-income credit) and the composer computes and SELECTS the Form 760 line 23 credit itself (TY2025 refundable VA EITC = 20% of federal EIC, uncapped — it dominates whenever federal EITC > 0). IL: pass ICR raw inputs instead where fields exist.
vaItemizedCharitableNoVA Sch A charitable contributions (federal Schedule A amount)
vaItemizedOtherTaxesNoVA Sch A line 6 other taxes (foreign income tax etc.)
vaItemizedSalesTaxesNoVA Sch A line 5a when the general SALES tax election was made federally — capped at the Virginia SALT cap ($40,000; $20,000 MFS for TY2025)
paEligibilityAddbacksNoSchedule SP Section III nontaxable add-backs (gifts, inheritances, insurance proceeds, non-PA income, nontaxable military pay, excluded home-sale gain, educational assistance, outside cash support). NOT Social Security, eligible retirement benefits, child support, or workers' comp.
paMsaHsaContributionsNoSchedule O codes M/H: MSA + HSA contributions at the federally-allowed amounts
paSpDependentChildrenNoSchedule SP dependent CHILDREN count (child/stepchild/adopted; grandchild of a grandparent; foster child of a foster parent — never other relatives) claimable as federal dependents; each adds $9,500 to the Tax Forgiveness eligibility-income threshold
paStudentLoanInterestNoSchedule O code S: student loan interest PAID (new deduction for 2025; the composer caps at $2,500 — pass the uncapped amount)
taxableSocialSecurityNofederally TAXABLE social security (Form 1040 line 6b, from compute_return). REQUIRED whenever nonzero: VA (760 line 5 subtraction), CA (Schedule CA line 6 col B), and NY (IT-201 line 27) all subtract it — the composer applies the subtraction automatically; do NOT also fold it into the generic subtractions total.
vaAgeQualifyingTestedNocount of filers born January 2, 1939 - January 1, 1961 (65+ for 2025 but income-tested): the composer computes $12,000 each, reduced dollar-for-dollar by AFAGI over $50,000 single / $75,000 married — where AFAGI = federal AGI MINUS the federally taxable social security (the SS exclusion is the step agents miss; Va. Code § 58.1-322.03(2))
vaSpouseAgeBlindBoxesNoSTA worksheet Part 1 line 2: spouse's 65+/blind box count (0-2)
vaSpouseTaxAdjustmentNoOVERRIDE ONLY — the composer now computes the VA Spouse Tax Adjustment worksheet itself when vaYourVagi/vaSpouseVagi are provided. If passed, this amount wins.
caDepreciationAdditionNoCA depreciation-difference addition: federal depreciation (with § 168(k) bonus, which California NEVER conforms to) minus CA depreciation (plain MACRS on the same asset). Positive = CA income addition (Schedule CA col C on the business/rents line). Compute per-asset and disclose.
paSpouseRentRoyaltyNetNoPA-40 line 6, spouse's own net rent/royalty amount
paUnreimbursedExpensesNoPA-40 line 1b: Schedule UE unreimbursed employee business expenses (a compensation-class expense, never a line-10 deduction)
spouseStateWithholdingNoVA line 19b spouse withholding (spouse's own W-2/1099/VK-1 boxes + spouse's half of jointly-issued documents' withholding, odd dollar to the primary)
vaScheduleAdjDeductionsNoSchedule ADJ line 9 total deductions (deduction CODES like 105 continuing-teacher-education, 199 other) — prints on Form 760 line 13; these are DEDUCTIONS from VAGI, never income subtractions on line 7
caAb5GrossIncomeAdditionNogross income from businesses where the worker is classified as an EMPLOYEE for California (AB 5/Dynamex reclassification; the intake's ca_form540_schca.add_gross_income field) — Schedule CA WAGE addition, col C
caHsaTaxableDistributionNoHSA distribution amount taxed federally (Form 8889 line 16) — not income for California: the composer SUBTRACTS it for CA
unemploymentCompensationNounemployment compensation included in federal AGI (Schedule 1 line 7). REQUIRED whenever nonzero: VA fully subtracts it (Va. Code § 58.1-322.02(9), Schedule ADJ) and CA excludes it (Schedule CA line 7 col B) — the composer subtracts automatically for those states; do NOT also fold it into the generic subtractions total. IL and NY tax it (no subtraction).
vaItemizedRealEstateTaxesNoVA Sch A line 5b real estate taxes — NOT subject to the SALT cap for Virginia
caEducatorExpensesDeductedNofederal educator-expense deduction claimed (§ 62(a)(2)(D)) — California does NOT conform: the composer ADDS it back on Schedule CA (line 11 col C). Pass the federal amount actually deducted (both spouses' combined).
caTaxableEarlyDistributionNoretirement-plan early distribution amount subject to the FEDERAL § 72(t) additional tax — California imposes its own 2.5% additional tax on the same base (R&TC § 17085(c)(1), FTB 3805P); the composer computes 2.5% and prints it on Form 540 line 63
vaItemizedMortgageInterestNoVA Sch A home mortgage interest and points (federal Schedule A amount)
priorYearOverpaymentCreditedNoprior-year state overpayment applied toward this year's estimated tax. VA 760 line 21 (its own printed line — never fold into line 20 estimated payments). Other states: folded into the estimated-payments line.
vaItemizedInvestmentInterestNoVA Sch A investment interest (protected from the overall limitation)
vaItemizedPersonalPropertyTaxesNoVA Sch A line 5c personal property taxes — NOT subject to the SALT cap for Virginia
vaItemizedStateLocalIncomeTaxesNoVA Sch A line 5a when INCOME taxes are claimed (mutually exclusive with sales taxes)
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses deterministic behavior, internal computations (oracle targets for state tax), and specific handling for each state. However, it does not explicitly state whether the tool has side effects or is read-only, and lacks a clear statement on mutability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but densely packed with essential information. It is front-loaded with key workflow and common errors. Some repetition or over-explanation could be trimmed, but overall it earns its length given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 98 parameters, 5 states, no output schema, the description covers workflow, state-specific rules, parameter dependencies, and common errors. It doesn't detail the output structure (line set format) but implies a verbatim line set. Slightly more on output format would raise completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite 96% schema description coverage, the description adds significant value beyond the schema. It explains why parameters are needed (e.g., 'caIsoPreference + caAmtTaxesAddback so the composer builds Schedule P AMTI'), gives interaction rules, and provides nuanced guidance (e.g., 'passed caAmt wins' when both given).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Compose a STATE return's printed-form line set deterministically' for specific states, clearly identifying the tool's purpose. It distinguishes from sibling tools like compute_return by emphasizing state-level composition and warning against hand-assembling line numbers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear workflow: run compute_return first, compute state-specific components, then call this tool once. It includes explicit when-to-use instructions (e.g., 'ALWAYS pass taxableSocialSecurity and unemploymentCompensation when nonzero') and when not to (e.g., 'federalAGI is NOT the PA base').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

determine_dependentAInspect

Determine whether ONE candidate person is the taxpayer's § 152 dependent — qualifying child or qualifying relative, including multiple-support agreements and the divorced-parents release — as a proof-backed yes/no with citations. Feed the result into calculate_tax's credits group (qualifyingChildren / otherDependents).

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNoREQUIRED for computation: the law-in-force date — use the intended tax year's year-end (e.g. "2025-12-31" for TY2025). Omitting it is an error, never a default.
depAgeNoCandidate's age at the end of the year (§ 152(c)(3)).
targetNorule to derive (default: us.federal.dependent.is_dependent). Other targets: us.federal.dependent.qualifying_child, us.federal.dependent.qualifying_relative
depGrossIncomeNoCandidate's gross income for the year (§ 152(d)(1)(B) limit: $5,200 TY2025 / $5,300 TY2026). In dollars.
depFilesJointReturnNoCandidate files a joint return with a spouse (other than a refund-only claim) (§ 152(c)(1)(E)).
depIsFullTimeStudentNoCandidate was a full-time student for at least 5 months (§ 152(f)(2)).
depRelationshipChildNoCandidate is the taxpayer's child, stepchild, foster child, sibling, step-sibling, or a descendant of any of them (§ 152(c)(2)).
depDivorcedParentsRuleNo§ 152(e) applies to the candidate child: the parents are divorced, separated, or lived apart the last 6 months of the year; the child received over half their support from the parents and was in their custody over half the year.
depPermanentlyDisabledNoCandidate is permanently and totally disabled (§ 152(c)(3)(B)).
depYoungerThanTaxpayerNoCandidate is younger than the taxpayer (§ 152(c)(3)(A)).
depRelationshipRelativeNoCandidate bears a § 152(d)(2) relationship to the taxpayer (parent, grandparent, sibling, in-law, etc.) or lived in the household all year.
taxpayerIsCustodialParentNoThe taxpayer is the custodial parent (the parent with whom the child resided the greater number of nights, § 152(e)(4)(A)).
hasMultipleSupportAgreementNoA § 152(d)(3) multiple-support agreement is in place for the candidate: the group together provided over half the support, no one person provided over half, each member could otherwise claim the candidate, and every other over-10% contributor signed a Form 2120 waiver.
custodialParentReleasedClaimNoThe custodial parent signed a written declaration (Form 8332) releasing the claim to the child for this year (§ 152(e)(2)).
depIsQualifyingChildOfAnotherNoCandidate is the qualifying child of the taxpayer or any other taxpayer (§ 152(d)(1)(D)).
depProvidedOwnSupportOverHalfNoCandidate provided more than half of their own support (§ 152(c)(1)(D)).
taxpayerProvidedOverHalfSupportNoThe taxpayer provided more than half of the candidate's support (§ 152(d)(1)(C)).
depLivedWithTaxpayerOverHalfYearNoCandidate had the same principal residence as the taxpayer for more than half the year (§ 152(c)(1)(B)).
taxpayerProvidedOver10PercentSupportNoThe taxpayer contributed over 10 percent of the candidate's support (§ 152(d)(3)(D) — the support test under a multiple-support agreement).
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool returns a 'proof-backed yes/no with citations' and mentions handled edge cases. It does not disclose consequences of missing required inputs (e.g., asOf is described as required but not in schema's required array) or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loading the primary purpose and adding downstream guidance without any fluff. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high complexity (19 parameters) and no output schema, the description provides the overall purpose, scope, and special cases. It does not detail the output format beyond 'yes/no with citations', but extensive schema descriptions compensate for parameter semantics. The description is adequate for an agent to understand usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The tool description does not add per-parameter meaning beyond the schema; it only provides overall context and downstream guidance. It does not compensate for any gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'determine' and the resource 'dependent status under § 152'. It specifies the scope (ONE candidate, qualifying child/relative) and includes special cases (multiple-support agreements, divorced-parents release). It also distinguishes from sibling tools like calculate_tax by noting downstream use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (to determine dependency for one candidate) and directs feeding the result into calculate_tax's credits group. However, it does not explicitly exclude scenarios or mention alternatives among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_ruleAInspect

Get a tax rule's statutory citation, verbatim excerpt, validity window, parameters, and dependencies. Use to quote the actual law behind an answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
ruleIdYese.g. "us.federal.standard_deduction" — list via calculate_tax proof or corpus
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the return contents (citation, excerpt, validity, parameters, dependencies) and implies a read-only operation. It does not mention side effects or authorization, but for a lookup tool, this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, no wasted words. The first sentence defines the function; the second gives usage advice. Information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter lookup tool with no output schema, the description adequately covers purpose, output, and usage. It could mention that it is read-only, but that is not essential given the context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already describes the parameter ruleId with an example and a note on how to list values. The description does not add further semantic details beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Get' and clearly identifies the resource: 'a tax rule's statutory citation, verbatim excerpt, validity window, parameters, and dependencies.' It also states the intended use case: 'quote the actual law behind an answer.' This distinguishes it clearly from sibling tools like calculate_tax or search_tax_rules.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear use case: 'Use to quote the actual law behind an answer.' While it doesn't explicitly state when not to use or list alternatives, the context implies it's for legal citations rather than calculations or searches, which is sufficient for guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_tax_cliffsAInspect

Find exact dollar amounts where one more cent of an input costs MORE than a cent of tax (marginal rate over 100%) — e.g. the EITC investment-income kill switch, CTC phase-out steps. Every probe is a real evaluation.

ParametersJSON Schema
NameRequiredDescriptionDefault
varyYesmoney fact to vary, e.g. "wages" or "taxableInterest"
factsYesfacts for the computation: either flat corpus fact ids (see list_input_facts) or the same group objects calculate_tax accepts (filing, income, retirement, …), plus optional target and asOf. Business/fiduciary/dependent facts are accepted flat. Unknown keys are rejected by name — nothing is ever silently dropped.
toDollarsYes
fromDollarsYes
stepDollarsNocoarse scan step, default 1000
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries burden. It states 'every probe is a real evaluation' and mentions unknown keys are rejected, but lacks details on side effects, permissions, or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences that are front-loaded with purpose and examples, no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks description of return format or error handling, which is important since no output schema exists. Could mention what the output looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%; description does not add new parameter details beyond what schema provides. The 'stepDollars' default and 'facts' rejection behavior are mentioned but most parameter meaning comes from schema names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool finds exact dollar amounts where marginal tax rate exceeds 100%, with examples like EITC and CTC phase-outs. It distinguishes itself from sibling tools that calculate taxes or compare statuses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies when to use: to identify tax cliffs. Does not explicitly state when not to use or mention alternatives, but the special-purpose nature makes usage fairly clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is_tipped_occupationAInspect

Determine whether a job is on the Treasury Tipped Occupation list (Treas. Reg. § 1.224-1, final Apr 2026) for the § 224 'no tax on tips' deduction. Fuzzy-matches the job name; returns the official listing (name, TTC code, category) or a definitive 'not listed'.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobYese.g. "bartender", "software engineer", "DJ"
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavioral traits. It mentions fuzzy matching and the return format, which adds value beyond the input schema. However, it does not state whether the operation is read-only, idempotent, or any side effects, leaving gaps in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences that front-load the core purpose and then provide key additional details (fuzzy matching, return values). Every sentence adds essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple lookup tool with one parameter and no output schema, the description covers the necessary aspects: it explains what the tool does, how it matches (fuzzy), and what it returns (listing or 'not listed'). No additional context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides a description for the 'job' parameter with examples, achieving 100% coverage. The tool's description adds meaning by explaining that the parameter is used for fuzzy matching and describing the output format, going beyond the schema's basic information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: determining whether a job is on the Treasury Tipped Occupation list for the §224 deduction. It specifies the verb 'determine', the resource (Treasury list), and the regulatory context, distinguishing it from sibling tools that perform calculations or comparisons.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (for §224 deduction) but does not explicitly state when to use this tool versus alternatives, nor does it provide exclusions or indicate when not to use it. No comparison to sibling tools is given, which limits guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_input_factsAInspect

Discover every input the tax corpus understands: id, type, whether required, and its documented default. Call this if unsure what information to collect from the user.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description carries full burden. It explains that the tool returns a list of input details with their characteristics, implying read-only behavior. No side effects or auth info, but low risk for this type.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with output description, followed by usage hint. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 0 params and no output schema, description fully covers what the tool does and when to use it. Explains return fields sufficiently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters. Description adds context about what fields each input includes (id, type, required, default), beyond the empty schema. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it discovers all inputs with specific fields (id, type, required, default). Distinguishes from sibling tools by being the discovery tool for inputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Call this if unsure what information to collect from the user.' No when-not-to mentioned, but usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_tax_parameterAInspect

Look up the current-law dollar amounts behind a question ('standard deduction', 'CTC phase-out threshold', 'tips deduction cap') with their statutory citations and validity windows. Use this to fact-check ANY tax number before stating it — your training data likely predates the OBBBA.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNo
queryYesplain-English search, e.g. 'standard deduction'
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations present, so description must carry full burden. It discloses return contents (dollar amounts, citations, validity windows) and implies current data. Lacks explicit safety traits (read-only, side effects) but is acceptable for a lookup tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. Front-loads core purpose and immediately follows with usage motivation. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 2 params and no output schema, the description covers what the tool returns and when to use it. Could briefly mention if it supports partial matches or date ranges, but sufficient for most use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 2 params with 50% description coverage (query described as 'plain-English search'). Description adds example values for query but does not elaborate on optional asOf parameter. Partially compensates for schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'look up' and resource 'current-law dollar amounts behind a question', with examples like 'standard deduction'. Clearly distinguishes from siblings by noting it returns statutory citations and validity windows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'fact-check ANY tax number before stating it'. Adds context about training data predating OBBBA. Does not explicitly exclude misuse cases, but provides strong directional guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_tax_rulesAInspect

Full-text search over the encoded tax-law corpus ('kiddie tax', 'NIIT threshold', 'california renters credit'). Returns matching rules: id, title, statutory citation, effective window, and a verbatim excerpt of the law text. A hit means the engine computes this; zero hits means it is outside the corpus — say so rather than guessing. Follow up with explain_rule for a hit's full formula, or lookup_tax_parameter for its dollar amounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNo
limitNo
queryYesplain-English search, e.g. 'kiddie tax'
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavior: it returns matching rules with specific fields, and explains hit semantics. Lacks mention of rate limits or auth but is otherwise transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very concise: two sentences conveying purpose, return format, and usage guidance. No redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a search tool: covers input parameters implicitly, describes output, and gives behavioral context (hits vs. no hits). No output schema needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only query described). The description adds a plain-English example for query but doesn't explain asOf or limit. Partial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action (full-text search) and resource (encoded tax-law corpus). Mentions specific return fields and includes examples of search terms, distinguishing it from sibling tools like explain_rule and lookup_tax_parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use advice: 'A hit means the engine computes this; zero hits means it is outside the corpus — say so rather than guessing.' Recommends follow-up tools (explain_rule, lookup_tax_parameter) for further details.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_factAInspect

Fact-check a claimed dollar amount about tax law ('the 2026 MFJ standard deduction is $32,200', 'CTC is $2,000 per child') against the corpus. Returns verified / refuted (with the correct value and citation) / unknown. Never states a verdict it cannot ground.

ParametersJSON Schema
NameRequiredDescriptionDefault
asOfNo
queryYeswhat the amount is, e.g. 'standard deduction'
filingStatusNo
claimedAmountYesdollars, e.g. 50000 or "1234.56"
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses return types and a behavioral principle (never ungrounded verdict), but does not explain the process (e.g., how the corpus is consulted), error handling, or edge cases (e.g., ambiguous claims). Partial transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, no wasted words. It front-loads the purpose and immediately clarifies the output types and a key behavioral rule. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a fact-checking tool with 4 parameters and no output schema, the description covers the core function and outputs. It could mention that it specifically handles dollar amounts (implied by param descriptions) and what the 'corpus' is, but overall it is adequate for an AI agent to understand the tool's role.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 50% of parameters with descriptions (query, claimedAmount), but the tool description adds no additional parameter guidance. The filingStatus enum and asOf date are left unexplained. Given schema coverage is only 50%, the description should compensate but fails to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fact-checks a claimed dollar amount about tax law, using a specific verb (fact-check) and resource (tax law claims). It explains the three possible return values (verified, refuted with correct value and citation, unknown), distinguishing it from general lookup tools like lookup_tax_parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use (to verify a specific dollar amount claim) but provides no explicit guidance on when not to use it or how it compares to sibling tools like verify_tax_claim or search_tax_rules. The constraint 'Never states a verdict it cannot ground' is a quality guarantee, not usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_tax_claimAInspect

Verify a claimed tax amount (yours, a user's, or another tool's) against the law. Returns verdict 'verified' or 'refuted' with the correct value. Use this as a self-check before presenting any tax number. Put asOf (and target, if any) INSIDE the facts object — e.g. facts: {..., "asOf": "2025-12-31"} — otherwise the claim is checked under today's law.

ParametersJSON Schema
NameRequiredDescriptionDefault
factsYesfacts for the computation: either flat corpus fact ids (see list_input_facts) or the same group objects calculate_tax accepts (filing, income, retirement, …), plus optional target and asOf. Business/fiduciary/dependent facts are accepted flat. Unknown keys are rejected by name — nothing is ever silently dropped.
claimedAmountYesthe amount to verify (negative = refund)
toleranceDollarsNodefault 1
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses that unknown keys are rejected (never silently dropped), default tolerance is 1, and that without asOf, the claim is checked under today's law. Lacks explicit readOnly or destructive hints, but behavior is adequately inferred.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three focused sentences: purpose/output, usage advice, and formatting tip. No wasted words. Front-loaded with essential info.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Explains input structure (facts as flat IDs or group objects), output (verdict plus correct value), and edge cases (unknown key rejection, asOf handling). Adequate for a 3-param tool with nested objects and no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds value with guidance on placing asOf and target inside the facts object, which is not in the schema. Also clarifies that claimedAmount negative means refund. Provides additional context beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool verifies a claimed tax amount against the law and returns 'verified' or 'refuted' with the correct value. It specifies it can verify claims for the user, another user, or another tool, distinguishing it from siblings like verify_fact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this as a self-check before presenting any tax number' and gives specific instructions on how to structure the facts object (asOf and target inside facts). Lacks explicit when-not-to-use or alternatives, but provides clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Discussions

No comments yet. Be the first to start the discussion!

Try in Browser

Your Connectors

Sign in to create a connector for this server.

Resources