Taokeh MCP server
Server Details
Taokeh is accounting software for Malaysian SMEs — double-entry books, LHDN e-Invoice (MyInvois), SST, and full statutory payroll — and this connector opens a company's live books to the AI its owner already uses.
59 tools. The reads answer real questions from the ledger: P&L and balance sheet with server-computed comparisons, cash position, A/R and A/P aging, per-channel marketplace sales, an 8-week cash-flow forecast, tax position, document search with e-Invoice standing, and a one-call daily brief. The writes are drafts only — expenses, invoices, bills, quotes, purchase orders, receipts, credit and debit notes, adjusting journals, bank-statement imports and bank-row suggestions — every figure re-checked by the server, every draft waiting for a human tap in Taokeh. The AI can also work the Shoebox: staff snap paper on free phone logins, and the connector lists the pile, reads each photo, and files the draft with the original attached — the server maps its own stored copy, so the document trail stays byte-perfect.
One connection is bound to one company at consent; no tool takes a company argument. Migrating from another system? The same connector stages the chart of accounts, opening balances, contacts, products, historical documents and workspace settings onto the owner's own review screens. Bring your own AI subscription — no per-call fees.
- Status
- Healthy
- Uptime
- 98.0% over 44 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 74 tools
The create_*/update_* draft tools are cleanly split by document type, and the descriptions aggressively cross-reference each other (e.g. ar_aging tells you to use open_invoices for one customer, business_snapshot points to daily_brief). The main risk is the reporting cluster: business_snapshot is a strict subset of daily_brief, and cash_position/cash_forecast/channel_summary/sales_summary/financial_timeseries/working_capital all touch overlapping figures. Descriptions mitigate this well, but a few pairs can still be confused.
Names are overwhelmingly snake_case with predictable verb_noun action pairs (create_bill_draft / update_bill_draft, search_documents, resolve_customer, stage_document) and noun-phrase report names (ap_aging, cash_position). A few idiomatic outliers (find_in_taokeh, intake_contract, draft_bank_classification, my_work) deviate slightly but remain readable and snake_case. No camelCase/snake_case mixing.
74 tools is far beyond the healthy 3-15 band and into the 'too many' territory, even for a broad accounting domain. Much of the surface is the create_/update_ draft matrix (~12 document types x 2), which could plausibly be consolidated behind a doc_type/kind parameter. The breadth is defensible but the count is heavy enough to burden selection and context.
The surface covers an unusually full lifecycle: create drafts, update posted documents as reviewable diffs, revise_draft for corrections, search_* readers for every record type, resolve_* helpers, attachment upload/retrieval, migration staging, payroll, inventory and a deep reporting suite. Deliberate gaps (no AI void/delete, no AI create for buy-side debit notes) are explicitly routed to manual in-app steps rather than left as dead ends.
Available Tools
74 toolsap_agingA/P agingARead-onlyInspect
Whom you owe money and how overdue it is — your accounts-payable aging by supplier, aged from each bill's DUE date (its own date when it states no term). WHAT A ROW IS: the PAYABLE each document created on the accounts-payable control account, less what has been paid against it by asOf — never the bill's gross total. UNMATCHED PAYMENTS: a bank payment posted to a supplier but not yet matched to a bill is already inside that supplier's bucket figures and total as a NEGATIVE amount, aged from the date it left, and listed under the row's unmatchedPayments with a kind: payment (negative), refund (a supplier refund not yet matched — positive), or booked_outside (the bank line was matched to bills for more than it booked to accounts payable — positive; the fix is on the bank line). unmatchedPaymentsTotal is their sum and documentTotals the buckets of the bills alone — quote documentTotals for what is OVERDUE. So a row's total is what you owe NET and can be negative. With them, this report equals the control-account balance to the cent, except for a manual journal posted straight to the control account, the period-end unrealized FX revaluation on its own date, and a bank line whose allocations were saved but which is not posted yet. Look for manual entries on the control account before doubting either figure. A cash bill, or a bill settled from somewhere other than the supplier account (an owner-paid bill booked to a shareholder loan), created no payable and is absent however its payment method reads; a debit note nets its supplier down. A figure here can therefore be lower than the same bill's total in search_documents — say which you are quoting.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint/openWorldHint false; the description goes far beyond, disclosing what a row represents (payable less payments, never gross), how unmatched payments enter as negatives, exclusions (cash bills, owner-paid bills, debit notes), and reconciliation exceptions (manual journals, FX revaluation, unposted bank lines).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but deliberately structured with front-loaded purpose and capitalized headers ('WHAT A ROW IS:', 'UNMATCHED PAYMENTS:'). Dense yet nearly every clause carries interpretation value; a few caveats could be trimmed but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so thoroughly, explaining total, documentTotals, unmatchedPayments, unmatchedPaymentsTotal and their signs. Complete for a complex financial report.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers the single asOf parameter at 0%, but the description supplies the missing meaning: it is the valuation date ('less what has been paid against it by `asOf`' and aging 'from the date it left'). Format is left to the schema pattern, which is fine.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Whom you owe money and how overdue it is — your accounts-payable aging by supplier, aged from each bill's DUE date.' An agent can immediately tell this is the payables aging report versus ar_aging or search_documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives rich interpretation guidance (quote documentTotals for overdue, compare against search_documents) but never states when to use this tool vs alternatives like ar_aging or the balance sheet. Usage is implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ar_agingA/R agingARead-onlyInspect
Who owes you money and how overdue it is — accounts-receivable aging across ALL customers, bucketed by age. For the unpaid invoices of ONE named customer, use open_invoices. WHAT A ROW IS: the RECEIVABLE each document created on the accounts-receivable control account, less what has been received against it by asOf — never the document's gross total. UNMATCHED PAYMENTS: a bank receipt posted to a customer but not yet matched to an invoice is already inside that customer's bucket figures and total as a NEGATIVE amount (money you hold for them), aged from the date it arrived, and is listed under the row's unmatchedPayments (bankTransactionId, date, amount, bucket, kind). kind says what it is: payment (money in, not yet matched — negative; suggest matching it to invoices rather than treating the invoice balance as unpaid), refund (money paid out to the customer, not yet matched — positive), or booked_outside (the bank line was matched to invoices for MORE than it booked to accounts receivable — e.g. posted to an income account first — so the invoices read settled while the control account still holds that amount; positive; the fix is on the bank line). unmatchedPaymentsTotal is their sum and documentTotals the buckets of the documents alone — quote documentTotals for what is OVERDUE, totals for the balance. So a row's total is what the customer owes NET, a row can be negative (they have paid ahead), and a customer can appear with no invoice at all. With them, this report equals the control-account balance to the cent, except for entries that belong to no document or bank line: a manual journal posted straight to the control account, the period-end unrealized FX revaluation (on its own date only — it reverses the next day), and a bank line whose allocations were saved but which is not posted yet. If an owner asks why the two disagree, look for those before doubting either figure. The document universe itself differs on purpose wherever the gross was never collectable: a cash or gateway-paid sale posts no receivable at all and is absent; a marketplace order reads the NET PAYOUT the channel owes (gross less the platform fee it withheld, whether that fee was known at import or only booked later from the settlement); a credit note nets its customer down; and an invoice reversed by a channel return reads nil. So a figure here can legitimately be LOWER than the same invoice's total in search_documents or on its PDF — say which you are quoting rather than treating one as an error. Movement booked after asOf is excluded, so a back-dated aging shows what was owed on that date.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint, so the description carries the behavioral burden and does so richly: negative rows, unmatched payments, reconciliation discrepancies, back-dated exclusion of post-asOf movement, and deliberate document-universe differences. It stops short of stating return format or any performance limits, but for a read-only report the coverage is well beyond the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and sibling routing are front-loaded in the first sentence, and the dense remainder is information-bearing rather than filler. It is a single sprawling block, however, and a few asides (e.g. the FX revaluation detail) could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still enumerates the return structure (unmatchedPayments fields, unmatchedPaymentsTotal, documentTotals, totals) and explains edge cases that would otherwise confuse an agent, leaving nothing essential unstated for a one-parameter report.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single `asOf` parameter, so the description must compensate. It does: 'Movement booked after `asOf` is excluded, so a back-dated aging shows what was owed on that date' tells the agent what the parameter actually controls, though it omits the date format (which the schema pattern supplies).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('accounts-receivable aging across ALL customers, bucketed by age') and immediately distinguishes itself from the sibling open_invoices by scope (ALL customers vs ONE named customer). An agent can pick between the two without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative tool and the condition that selects it, tells the agent which figure to quote for which question (documentTotals for overdue, totals for balance), and explains when this report should and shouldn't match the control account.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
balance_sheetBalance sheetARead-onlyInspect
Your balance sheet as of a date: assets, liabilities and equity, with the balanced check. ONE EQUITY LINE IS NOT AN ACCOUNT: Current Earnings, whose code is an em-dash rather than an account number, is the profit to date — revenue minus expenses over every entry up to asOf — folded into equity so the sheet balances. Nothing posts to it, it has no ledger of its own, and there is no year-end close in these books, so it runs back to the company's first entry rather than restarting each January; it equals the income_statement net income over the same end date. Never describe it as an account, never offer to open its ledger, and never treat it as a missing or mis-coded row. COMPARISON (optional): a balance sheet is a POSITION AT A DATE, not a period — so unlike income_statement its comparison is ONE DATE against ANOTHER DATE, never two ranges. Set compare to 'previous_month_end', 'previous_year_end' or 'same_date_last_year', OR give an explicit earlier date as compareAsOf. You get both positions in full plus, per account line and per total, deltaCents and deltaPct computed server-side in integer cents — the movement in that balance BETWEEN the two dates. deltaPct is NULL whenever the earlier balance is zero (a percentage change from zero is undefined — say "no comparable base", never ∞% or 100%), and percentages are measured against the ABSOLUTE earlier balance. points gives both dates, how many days apart they are, and whether either is today or later (still open, so the balances can move); notes states the basis in plain words, including a warning if either sheet does not balance. Never describe the gap between the two dates as a 'period' figure: these are two positions, and the delta is the movement between them.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| compare | No | Named comparison POINT (a single earlier date, not a range). Omit for no comparison. | |
| compareAsOf | No | An explicit earlier date to compare the position against (overrides `compare`). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavior: deltas are computed server-side in integer cents, deltaPct is NULL when earlier balance is zero, percentages use absolute earlier balance, current earnings is not a real account, nothing posts to it, and warnings are included if the sheet does not balance. This is detailed behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is quite long and repeats the position-versus-period warning multiple times, but the repetition is purposeful to prevent common misinterpretations. It is well structured with clear sections for comparison semantics, delta behavior, and output notes, so the length is mostly justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though there is no output schema, the description fully explains output concepts: deltaCents, deltaPct, points, notes, null handling, and warning behavior. It also covers edge cases like zero earlier balances and imbalanced sheets, making it complete for an agent to invoke and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds strong semantic value for compare and compareAsOf, including override behavior and the meaning of named comparison points. However, asOf has no schema description and the description does not state what happens when asOf is omitted, even though it is not required in the schema. It mostly compensates for schema gaps but leaves that minor ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool returns a balance sheet as of a date: assets, liabilities, equity, and the balanced check. It also distinguishes itself from income_statement by emphasizing position vs period, so an agent can tell it apart from key siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when the comparison mode is appropriate, how to choose named comparison points versus an explicit earlier date, and warns against framing the delta as a period figure. It names income_statement as the contrasting tool and gives clear 'never' instructions, acting as practical usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_reconciliation_reportBank reconciliation reportARead-onlyInspect
The AUDIT bank reconciliation for one bank account, exactly as Banking → Reconciliation report and Clear transactions show it — so you can work the reconciliation WITH the owner instead of sending them to the app. Returns (1) report: balance per bank statement (or "Not stated" when the statement carries no closing balance — pass balance from the paper statement to work it through; used for this answer only, never stored), less unpresented payments, add deposits not yet credited, adjust for statement lines not yet in the books, = balance per books worked from the statement, beside the ledger balance and the UNEXPLAINED difference (never forced to zero), plus the Clear Transaction listing for the statement period; (2) clearStep: the book entries not yet cleared up to the statement date (each with its entryId, date, reference, description, signed amount, and any suggested statement line of the same amount), and the statement lines a tick may pair with (statementLineId, date, amount, remaining); (3) hints: the three patterns the page explains, in the page's own words (explanation) — a pair of entries that cancel out (clearDate is never before the statement starts), a cash sale with no deposit on the statement (never counter-day or POS takings), and an opening balance the books do not hold: none on the account (kind: missing) or one dated after the first statement starts (kind: late), never given when the report ties; when actionable is false, no fix is safe to name, so name none. Pick the account with bankAccountId (default: the first bank account), and the statement with statementId (default: its latest) or a date with asOf. Clear transactions lists 300 rows at a time, oldest first; offset reads the next batch. Report lists over 200 rows are truncated, their totals never. Ringgit bank accounts only — a credit card or a foreign-currency account is refused with the page's own sentence. READ-ONLY: to tick entries, file create_bank_clearing_draft; to settle a posted payment against invoices or bills, file create_payment_match_draft. The owner approves both in Taokeh.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | Reconcile to this date by hand instead of a statement (YYYY-MM-DD). | |
| offset | No | Clear transactions paging: skip this many uncleared rows (the page shows 300 at a time). | |
| balance | No | The closing balance printed on the paper statement — used only when the statement carries none ("Not stated"). Never stored. | |
| statementId | No | An imported statement — from `statements`. Omit for the latest (or give asOf instead). | |
| bankAccountId | No | The bank account — from `bankAccounts` in an earlier answer. Omit for the first bank account. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations, disclosing read-only behavior, the 'never forced to zero' difference, handling of 'Not stated' balances, pagination limits, truncation behavior, and refusal of non-ringgit accounts. It aligns with the readOnlyHint annotation and adds substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence is purposeful: it front-loads the core purpose, then systematically details return objects, edge cases, defaults, restrictions, and related tools. No fluff or redundancy; the structure aids comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return values, which it does comprehensively (report, clearStep, hints) including their subfields and special conditions. It also covers pagination, defaults, and restrictions, making the tool fully usable without external reference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though the schema describes all 5 parameters (100% coverage), the description enriches them: explains defaults (first bank account, latest statement), the purpose of balance (only when statement lacks closing balance, never stored), offset's paging (300 rows), and asOf's hand-reconciliation role. This adds significant meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: performing an audit bank reconciliation for one bank account, matching the app's behavior. It specifies the resource (bank account) and distinguishes from siblings like bank_reconciliation_status and create_bank_clearing_draft, leaving no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly explains when to use this tool (to work the reconciliation with the owner) and names alternative tools for actions (create_bank_clearing_draft for ticking entries, create_payment_match_draft for settling payments). It also details parameter defaults and selection logic, leaving no guesswork.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_reconciliation_statusBank reconciliation statusBRead-onlyInspect
Where each bank account stands on reconciliation: GL balance vs the latest statement's closing (with the difference, or an honest "unknown" when no balance is stated); rowCounts by stored status (SUGGESTED = proposed, awaiting the owner; UNCATEGORIZED = never categorised, incl. hand-typed rows; CONFIRMED = categorised, not in the ledger; POSTED; IGNORED = set aside); the unreviewed backlog (SUGGESTED + UNCATEGORIZED not paired on Clear transactions) with pairedInBooks; reconcilingItems — each item's signed effect on the difference, summing to it exactly (NOT_EXPLAINED is the honest remainder, STATEMENTS_DONT_JOIN a missing or duplicated statement); notOnStatementNotInBooks — unposted lines on no statement (hand-typed, or a recorded receipt/payment not yet posted), which do not affect the difference; and per recent statement its balances, tie-out and counts. A CREDIT CARD's figures are a liability: negative = money OWED, never cash. Same figures as the in-app Bank Reconciliation page. READ-ONLY.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, and 'READ-ONLY' merely restates that. However, the description adds substantive behavioral context the annotations cannot: the meaning of each status bucket, that NOT_EXPLAINED is an 'honest remainder', that STATEMENTS_DONT_JOIN signals a missing/duplicated statement, and crucially the CREDIT CARD liability sign convention (negative = money owed). That domain semantics is genuinely valuable beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is an extremely dense single block, essentially one sprawling sentence with only a final 'READ-ONLY.' tacked on. The core purpose is front-loaded in the opening clause, but the intervening enumeration of every status, bucket, and edge case is hard to scan and much of it could be structured or trimmed. Length is not justified by what it helps an agent decide.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the full burden of describing the return shape, and it does so thoroughly: balances and differences per account, rowCounts by status, backlog with pairedInBooks, reconcilingItems with signed effects, notOnStatementNotInBooks, and per-statement tie-out. An agent can anticipate the response structure. What it lacks is any pointer on when this tool is the right choice relative to sibling reconciliation tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is no parameter syntax for the description to document and it adds nothing spurious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (per-bank-account reconciliation status) and enumerates what the status comprises: GL vs statement balance, rowCounts by status, backlog, reconcilingItems, and per-statement tie-out. It is specific about the returned content, but it never distinguishes itself from the sibling bank_reconciliation_report or bank_review_queue, both of which sound like overlapping reconciliation views.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to call this versus bank_reconciliation_report, bank_review_queue, or the in-app page it references. The only guidance is the implicit framing 'Where each bank account stands on reconciliation', which conveys context but no selection criteria or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_review_queueBank review queueARead-onlyInspect
The IMPORTED BANK REGISTER — the rows that came in on a bank or card statement, in ANY status, defaulting to the ones still awaiting review. By default (no status) it is exactly the unreviewed queue it has always been: the unposted rows of one statement, or of the latest imported statement if you don't name one. Pass status to read the settled history instead — 'confirmed' (categorised, not yet posted), 'posted' (in the ledger), 'ignored' (the owner set them aside) or 'all'. Pass from/to WITHOUT a statementId to read a DATE WINDOW across every statement this company has imported, rather than one statement; the payload then says scope:'window' and statement is null. A window holding more than 1000 rows still awaiting review is refused with the count — narrow the dates or name a statementId; settled history has no such ceiling. ⛔ HAND-RECORDED settlements are NEVER here whatever you pass: a receipt or supplier payment typed into Taokeh by hand has no statement behind it, and this tool is the imported register only. To find those, use search_documents with docType:'payment'. Each row still awaiting review carries Taokeh's LIVE suggested category + contact, a confidence score (0–1) and a plain-language reason, computed the SAME way the in-app Banking → Review screen computes them. A suggestion does NOT come only from the tenant's transaction history: the owner can write a standing RULE for a payer (What Taokeh has learned → Bank rules), and such a rule beats what Taokeh worked out on its own — those rows' reasons read "You set this rule on " rather than "Learned: … (×N)". If a suggestion looks unlike the payer's history, the rule is usually why; say so instead of calling the history wrong. A known supplier's money-out line is suggested as SUPPLIER_PAYMENT only when that supplier has open bills the amount could pay; otherwise it is suggested as EXPENSE with the supplier kept (a payment with no bill would leave payables negative). Sorted lowest-confidence first (the ones that need a human eye), capped at 200 with an honest truncation note, plus per-confidence bucket counts so you can say e.g. "46 are mechanical, 13 need eyes". A row can also carry possibleDuplicateOf: a receipt or supplier payment the owner already recorded BY HAND that looks like the same money as this imported line (same account, same amount, within three days) — it is only a resemblance, Taokeh decides nothing from it, and you should say so rather than propose a category for money that may already be in the books. What you CAN do with a row: review it, explain it, and PROPOSE a category with draft_bank_classification — and, for the three categories whose account the owner picks by hand (EXPENSE, OTHER and INTERNAL_TRANSFER), the LEDGER ACCOUNT too. EXPENSE is a plain expense paid by card or bank (money out, no party — never A/P); an unrecognised card purchase is suggested as EXPENSE. Your proposal is written to the row as an "AI suggestion" the owner sees on Banking → Review, with your proposed account pre-selected in their dropdown and marked as yours; rows already carrying a proposal from you come back here as yourProposal, so you can see what you told them. What you can NEVER do: CONFIRM or POST a row — and proposing the account changes nothing about that: it fills a dropdown, the owner still clicks. Accepting a category and posting each line is always a human click in Taokeh; hand the owner the review link. (Since 2026-08-19 the owner can also accept Taokeh's confident suggestions, and post confirmed lines, from the /go command bar's approval deck — still one human tap either way. YOUR OWN proposals are accepted only on Banking → Review, where your reasons are shown.) ⚠ SETTLED ROWS CARRY NO SUGGESTION. A row that has already been confirmed, posted or ignored comes back with suggestedCategory, suggestedContact, confidence, reason and possibleDuplicateOf all NULL — the classifier is not re-run over decided history, because a fresh guess printed beside a decision the owner already made reads as a second opinion on settled books. Those rows carry what was actually DECIDED instead: category, glAccount ({code,name}, null if the row has none), contact (the customer or supplier on the row), journalEntryId (null unless posted), splits (where the amount was divided across accounts) and allocations (the invoices or bills the money settled). statusCounts always reports how the rows in scope break down by status. A statement line the owner PAIRED with a book entry on Clear transactions is already in the books through that entry: it is never in the unreviewed queue or the 'confirmed' list; 'all' lists it with pairedWith (the entry's reference and date), and its statusCounts count it as PAIRED — never propose a category for it or ask the owner to post it. A row still awaiting review may carry possibleTwin {entryId, ref, date, kind}. kind 'bookEntry': a book entry on no statement that is the same money (same account, same amount, within three days) — the owner's review row warns that it looks already in the books. For such a row, propose a clearing with create_bank_clearing_draft (that entryId + this row as statementLineId) instead of a category. kind 'cardBill': a money-in line on a CREDIT CARD statement with a payment into the card of the same amount already in the books (usually the bank line that paid the card bill) — it is the card bill itself, so posting it would count the same money twice; a card has no clearing, so tell the owner to exclude the line on Banking → Review instead of proposing a category. kind 'cardLinePosted': a money-out bank line paying into a credit card (a credit card payment, or any line booked to the card's own ledger account) whose card-statement settlement line (ref = that line's description) is already POSTED — posting both counts the bill twice, so the owner should unpost that card line, or post this one anyway if it is a different payment. Each is only a resemblance (a same-amount refund can look alike), so say so.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Latest transaction date (YYYY-MM-DD). | |
| from | No | Earliest transaction date (YYYY-MM-DD). WITHOUT a statementId, from/to widen the read to a DATE WINDOW across every imported statement — 'what did the bank register show in July' rather than 'what is on this one statement'. | |
| status | No | Which rows: 'unreviewed' (DEFAULT — everything not yet posted or ignored, i.e. the review queue), 'confirmed', 'posted', 'ignored', or 'all'. Only unreviewed rows carry suggestions; the rest carry what was decided. | |
| statementId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint/openWorldHint, so the description carries the real behavioral load and does so richly: the 1000-row ceiling on unreviewed windows (with the refusal behavior), the 200-row cap with truncation, lowest-confidence sort, per-bucket counts, the standing-rule override, and the hard limit that it can never confirm or post. It also documents how decided rows come back with NULL suggestions versus decided fields — far beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical scoping constraint and default behavior are front-loaded, which is good. However the body is extremely long and packs many edge cases (three possibleTwin kinds, duplicate semantics, rule explanations, /go command bar history) into one block, which is more than an agent needs to invoke the tool and risks burying the actionable instructions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, no-output-schema tool of high domain complexity, the description is complete: it explains returned fields, suggestion provenance, duplicate/twin signals, and the boundary of what can and cannot be done. Nothing an agent needs to interpret the payload or act on a row is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, but the description adds meaning the schema alone doesn't: the interactive effect of `from`/`to` *without* a `statementId` widening scope across all statements (payload `scope:'window'`), the ceiling that only applies to unreviewed windows, and the exact `status` vocabulary interpretation. This goes beyond restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a specific resource ('the IMPORTED BANK REGISTER — the rows that came in on a bank or card statement') and its default scope ('defaulting to the ones still awaiting review'). It explicitly distinguishes itself from the sibling that covers hand-recorded settlements (search_documents with docType:'payment'), so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when/when-not rules: omit `status` for the unreviewed queue, pass `status` for settled history, pass `from`/`to` without `statementId` for a date window across statements. It also names the alternative tool for the excluded case (hand-recorded settlements via search_documents), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
business_snapshotBusiness snapshotARead-onlyInspect
A one-glance health check of your business: today's and this-month's sales and expenses, so far. Sales and expenses ONLY — if you want the fuller morning read (cash position, who owes you, pending approvals, tax position, low stock) call daily_brief instead; it includes everything here.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint=true, openWorldHint=false), and the description adds meaningful behavioral context beyond them: the temporal framing ('so far', today's and this-month's), the hard exclusion scope ('Sales and expenses ONLY'), and the point-in-time nature of a snapshot. It stops short of a 5 only because it doesn't describe the return format or whether figures are aggregates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero wasted words: the core purpose is front-loaded ('one-glance health check'), the scope is stated precisely with 'ONLY', and the alternative routing follows immediately. The em-dash separation and caps signal emphasis without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only snapshot tool with no output schema, the description fully covers what an agent needs: what data is returned (sales and expenses, today and this month), what is excluded, and what to call instead for broader coverage. Nothing required for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty and the parameter count is 0, so the baseline of 4 applies — there are no parameters whose semantics need clarification. Schema description coverage is trivially 100% for an empty schema, and the description correctly communicates that this tool requires no arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: a 'one-glance health check' of 'today's and this-month's sales and expenses'. It sharply delineates scope with 'Sales and expenses ONLY' and explicitly differentiates itself from daily_brief, so an agent can distinguish it from the 58 siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-to-use-this vs. when-to-use-that rule: call daily_brief instead when wanting the fuller morning read (cash position, receivables, approvals, tax, low stock), noting daily_brief 'includes everything here'. This is a direct alternative with a named sibling and a specific selection condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cash_forecastCash-flow forecastAInspect
A forward 8-week cash-flow forecast — honest arithmetic from your books, not a model. Opening position is your LIQUID balances right now (cash + bank accounts + gateway clearing); money the owner has advanced to the business is a liability, not cash, so it is excluded — the assumptions list says so. Inflows are your OPEN customer invoices scheduled at their due dates; already-overdue receivables are NOT assumed collected — they are reported separately as overdueArCents (upside, stated honestly). If a sales channel is connected and reports money earned but not yet paid out (TikTok Shop's unsettled orders), that pool comes back SEPARATELY as channelUnsettled (amount, count, currency, and the date it should land, 14 days out — the late end of the channel's own settlement window). It is the channel's own ESTIMATE, so it is deliberately NOT counted into any week, any closing balance or the dip: a warning about cash running low must never be weakened by money that is not in the books. Treat it as stated upside, exactly like overdueArCents. If no channel is connected, the channel add-on has lapsed, or its API cannot be reached, the field is absent and the forecast is unchanged. Outflows are your open supplier bills (at bill date — Taokeh does not store supplier terms yet), payroll (estimated from your last finalized run: net pay + employer EPF/SOCSO/EIS on payday), PCB/EPF/SOCSO/EIS remittances (15th of the following month), loan installments (estimated from the last recorded installment), and SST payable at its return due date if you're SST-registered. Returns week-by-week inflows/outflows/closing balance, the dip (the lowest projected closing — when and how deep cash runs lowest), and an assumptions list spelling out every estimate. All amounts are integer CENTS (RM = cents ÷ 100).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the minimal annotations. It discloses that overdue receivables are not assumed collected, channel unsettled balances are deliberately excluded from the dip, missing channel data leaves the field absent, and all amounts are integer cents. These behavioral caveats are exactly what an agent needs to interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but each sentence carries decision-relevant detail: what is included, what is excluded, how estimates are surfaced, and how edge cases are handled. It is front-loaded with the core purpose and then systematically covers inflows, outflows, outputs, and units.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return values, and it does: week-by-week inflows/outflows/closing balance, the dip, assumptions, overdueArCents, channelUnsettled, and units. It also covers failure modes such as missing channel connections or unreachable APIs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty input schema, so the baseline is 4. The description cannot add parameter-level meaning, but it compensates by explaining the internal semantics of the computed fields and assumptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'A forward 8-week cash-flow forecast — honest arithmetic from your books, not a model.' It clearly defines what the tool computes and distinguishes it from a current-position snapshot tool like cash_position by emphasizing the forward-looking 8-week window.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: it is a forward-looking forecast built from books, inflows, outflows, and projected low points. It does not explicitly name sibling alternatives or state 'use this when...' versus other cash tools, but the 8-week forecast scope makes the selection context unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cash_positionCash positionARead-onlyInspect
How much cash you have right now: every bank account and cash pocket with its current GL balance, plus a grand total. total is LIQUID money only — physical cash + bank accounts + the payment-gateway clearing float — and accounts lists exactly those accounts. Money the OWNER has advanced to the business (the Shareholder Loan) is returned SEPARATELY as ownerFunding: it is a LIABILITY the business owes back, not cash it holds, so it is never inside total and must never be added to it or described as cash. CREDIT CARDS are returned separately for exactly the same reason, as cards (one entry per card, with cardsOwing for the total): what sits on a card is money OWED to the issuer, not cash the business holds, so it is never inside total — report it as card debt beside the cash figure, never as cash and never as a negative bank balance. Each card carries its own direction ('owed_to_issuer', or 'in_credit' when the card is overpaid) — follow it rather than assuming a debt. Balances are MYR-booked (the books are kept in ringgit); a foreign bank account shows its currency label for context, but the balance figure is still the MYR-booked GL balance.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by explaining key behavioral details: ownerFunding is a liability excluded from total, credit cards are returned separately as card debt, each card has a direction, and balances are MYR-booked. This is rich contextual transparency for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but almost every sentence adds an important semantic distinction. It is front-loaded with the core answer and then explains exclusions. Minor redundancy exists around 'accounts' and total components, but the structure is still effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the full burden of explaining the response shape and semantics. It covers the total, account breakdown, shareholder loan, credit cards, direction fields, and currency treatment—making the tool fully understandable to an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema imposes no burden. The description instead focuses on output meaning, which is appropriate. No parameter-level explanation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly defines the tool as reporting current cash across bank accounts and cash pockets, including a grand total. It is specific about what the tool computes, but it does not explicitly differentiate itself from sibling tools like cash_forecast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes a clear context: the tool answers 'How much cash you have right now.' It also provides strong exclusions about what should never be treated as cash, but it does not explicitly state when to choose this tool over alternatives such as cash_forecast.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channel_summarySales-channel summaryARead-onlyInspect
Your sales broken down by CHANNEL (Shopee, TikTok, Shopify, your storefront, …) for a date range. Per channel: the order count, gross goods revenue, seller discounts (e.g. Shopee vouchers), shipping income (buyer-paid shipping), freight-out (what shipping cost you), marketplace fees, and netProceeds — the settlement booked on each sale's own entry: for credit marketplaces (Shopee/TikTok) that is the net A/R payout the channel owes you; for the storefront it is what settled into cash / the gateway clearing account at sale time. RETURNS: returns is the revenue reversed by that channel's credit notes (returns/refunds) in the range — the Sales Returns & Allowances booked on them; netGoodsRevenue = grossGoodsRevenue − returns. The other six figures (grossGoodsRevenue, discounts, shipping, freight, marketplaceFees, netProceeds) are GROSS of returns — they come from the INVOICE side only — so a channel with heavy returns no longer looks inflated. A row appears for every channel with sales OR returns in the range — returns lag sales, so a channel whose only in-range activity is returns still shows (saleCount 0, gross figures 0, its returns, and a negative netGoodsRevenue). (A Shopee-wallet FEE credit note books marketplace fees, not a goods return, so it adds 0 to returns.) A marketplace SETTLEMENT CORRECTION — the payout came in short because the buyer returned or cancelled after the invoice was already booked — is also a credit note approved by the owner against the original invoice, and the invoice side stays as it was booked. Only the GOODS portion of such a correction shows in returns: its fee and shipping components are reclassed to marketplace fees / freight on the credit note's OWN entry, while the per-channel marketplaceFees and freight figures here are read from the INVOICE side only — so a fee-only correction (the common, batched kind) moves neither returns nor marketplaceFees and is not visible in this tool's output at all. Marketplace invoices are therefore not final at import: a channel's figures for a past range can still move as corrections are approved. NOTE: marketplaceFees covers marketplace commission/service fees booked on the sale's own entry; storefront GATEWAY processing fees are booked on separate payout entries and are NOT included per-channel here (so storefront netProceeds is gross of gateway fees). TIKTOK is the one exception, and it is INCLUDED: TikTok reports no per-order fee at import, so its fees (and the freight split out of its settlement) are booked LATER on a separate settlement entry by "Book settlement fees" (journal reference TIKTOK-FEE-). Those are folded into TikTok's marketplaceFees and freightOut here, and the labelled revenue correction on them adjusts its grossGoodsRevenue. ⚠ THE DATE BASIS DIFFERS on that side: settlement entries are matched on WHEN THE SETTLEMENT WAS BOOKED (the journal entry's own date), while sales are matched on sale date — because that is when the cost hits the P&L. So a June TikTok sale whose payout is settled in July puts its goods revenue in June and its fees in July, and a July-only summary can show a tiktok row with saleCount 0 that is purely settled fees. Say which month you are reading, and never describe a channel's fees as final until its payouts are settled. Manual / non-channel sales are excluded. Give from and to as YYYY-MM-DD.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only, and the description adds substantial behavioral detail: returns are matched to credit notes, settlement corrections do not move invoice-side figures, TikTok fees are booked later and folded in, and reported figures may not be final until payouts settle. This goes far beyond the structured annotations and clarifies real-world accounting behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but it is well-structured with labeled sections (RETURNS, NOTE, TIKTOK) and front-loads the core channel breakdown before going into exceptions. Some repetition and parenthetical detail could be trimmed, but the complexity of the accounting semantics justifies much of the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully documents the returned fields, row presence behavior, returns treatment, settlement corrections, channel exceptions, and the warning that past-range figures can move. For a complex reporting tool, this is exceptionally complete and leaves little ambiguity about what the agent will receive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden for parameter meaning. It explains that from and to define a date range, gives the YYYY-MM-DD format, and—crucially—documents that sale-date and settlement-date matching differ for TikTok. It does not explicitly clarify inclusivity or timezone handling, but it compensates well for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Your sales broken down by CHANNEL' for a date range, and it names the relevant channels. This clearly differentiates it from generic sales tools like sales_summary by establishing the channel dimension as the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong usage context: it explains the date range requirement, excludes manual/non-channel sales, and warns that settlement-based fees and returns can make a channel appear differently depending on the month read. It does not explicitly name sibling alternatives or say when to prefer another tool, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
counter_dayCounter dayARead-onlyInspect
The Counter (Kaunter) till's day, for one date (default today, Malaysia time). Returns: the receipt count and voided count, the cash and DuitNow QR totals, the 5-sen cash-rounding total, whether the day is still OPEN or has been CLOSED into the books, the receipt-number range with any voided numbers listed, and separately the sales BILLED ON CREDIT at the till. IMPORTANT — a counter day is NOT in the ledger until it is closed: while status is 'open' these figures are the till's own receipts and NOTHING has posted, so they will not appear in sales_summary, cash_position or income_statement yet, and stock/COGS have not moved either (they move once, at close). Once status is 'closed', saleId names the ONE consolidated cash sale the whole day posted as, and the figures do reconcile to the books. THE ONE EXCEPTION, and it runs the other way: creditSales are the till's "Invoice it" sales, and they POST IMMEDIATELY as ordinary credit invoices (Dr accounts receivable to a named customer, stock moved at the moment of sale). So unlike the drawer takings they ARE already in sales_summary, income_statement, open_invoices and ar_aging even while the day is open — and they are NOT in cashTotal, qrTotal or the close, because no money changed hands. Never add creditSales.total to the cash and QR totals: the drawer holds cash + QR only, and it ties to the sen. cashTotal is the sum of ROUNDED cash receipt totals (what the drawer holds); qrTotal is never rounded — Malaysia's 5-sen rounding applies to cash only. Receipts of RM10,000 or more cannot join a consolidated e-invoice, so each posts as its own sale: consolidatedCount/consolidatedTotal cover only the receipts that joined the day sale, individualCount/individualTotal the rest, and awaitingBuyerCount is how many of those still have no buyer particulars captured. Counter is a paid add-on (RM49/mo); a company without it gets status 'not_available'.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | The business date, YYYY-MM-DD in Malaysia time. Omit for today. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the readOnlyHint annotation, explaining critical behavioral nuances: an open day is not in the ledger, stock/COGS only move at close, closed days post as one consolidated sale via saleId, creditSales post immediately as credit invoices, cash is rounded but QR is not, and large receipts are handled individually. This prevents serious misinterpretation of the returned figures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but almost every sentence addresses a distinct misinterpretation risk or clarifies an important return value, which is justified given there is no output schema. It is front-loaded with the return summary before the caveats, though tighter phrasing could reduce its overall length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries full responsibility for explaining return semantics, and it does so thoroughly: open/closed behavior, saleId, creditSales, rounding rules, consolidated vs individual receipts, and the not_available state for companies without the paid add-on. An agent has enough information to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter is 'date', and the schema already covers 100% of its meaning with 'YYYY-MM-DD in Malaysia time. Omit for today.' The description repeats this but does not add meaningful parameter-level detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly defines the resource ('Counter till's day'), scope (one date, default today, Malaysia time), and the specific outputs returned (receipt count, voided count, cash/QR totals, open/closed status, receipt ranges, credit sales). It also distinguishes this tool from ledger-level sibling reports by explicitly stating that an open counter day has not posted to sales_summary, cash_position, income_statement, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong context for when this tool is relevant: use it to inspect till-level day figures before they reach the ledger, and it explains how those figures relate to sales_summary, cash_position, income_statement, open_invoices, and ar_aging. It does not explicitly state 'use this instead of X', but the ledger-vs-till distinction gives clear practical guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_bank_clearing_draftPending draft: tick bank entries as cleared (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose the ticks for Banking → Clear transactions: which book entries have appeared on the bank statement, each paired with its statement line or given the date it cleared. An admin or bookkeeper reviews the list and approves; only that tap ticks them, exactly as the page's Save ticks does (a tick posts nothing to the books and the owner can un-tick it). Read bank_reconciliation_report first and take every id from it: bankAccountId, statementId (the statement you are reconciling to — or asOf, a date, when there is none), and per item entryId (from clearStep.uncleared) plus EITHER statementLineId (from clearStep.candidateLines — the entry then clears on the later of the line's date and its own) OR clearedDate (YYYY-MM-DD; omitted, it defaults to the reconciliation date, like the page). At most 300 items per draft. A cancelling pair from hints.cancellingPairs is filed as two items with no line and the pair's clearDate. REFUSED BY NAME when you file (Taokeh runs the real tick and rolls it back) and again at approval: an entry that is not uncleared on this account up to the date (posted from a statement line, the opening balance, already ticked, dated later), a line that is posted, fully paired or dated after the reconciliation date, a line without room left for the entry, a line moving money the other way, a cleared date before the entry was booked or after the reconciliation date, a date inside a locked period, a card or foreign-currency account, an entry already in another pending draft. If anything it was filed against changes before approval (a tick made, a line posted or paired, an entry edited), approval refuses and nothing is ticked — read the report again and file afresh. note: one short line for the owner saying why.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | Reconcile to a date by hand (no statement). Give this OR statementId. | |
| note | No | A SHORT note for the owner: why these are cleared. | |
| items | Yes | The ticks. An unknown key in an item is refused by name. | |
| notes | No | NOT a filter on this tool. Pass the owner note as `note`. | |
| ticks | No | NOT a filter on this tool. Pass the ticks as `items`: [{entryId, statementLineId?, clearedDate?}]. | |
| bankId | No | NOT a filter on this tool. Pass the bank account as `bankAccountId`. | |
| entries | No | NOT a filter on this tool. Pass the ticks as `items`: [{entryId, statementLineId?, clearedDate?}]. | |
| needsReview | No | Set true when something gave you pause. | |
| statementId | No | The statement you reconcile to, from bank_reconciliation_report `statement.id`. Give this OR asOf. | |
| bankAccountId | Yes | REQUIRED — the bank account, from bank_reconciliation_report. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond annotations: it prominently states 'FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH', explains that 'a tick posts nothing', details refusal conditions by name, notes re-approval failure if underlying data changes, and clarifies defaults such as clearedDate. This is rich behavioral context that no structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries essential information for a safety-sensitive, multi-parameter operation. It is front-loaded with the most critical warning. It could be improved with bullet points or shorter paragraph breaks, but the density is justified by the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 parameters, intricate refusal logic, and an implicit human approval workflow, the description covers all operational aspects: required inputs and their sources, pairing rules, defaults, limitations, failure conditions, and post-submission behavior. Nothing essential is left to inference, making it highly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 100% schema description coverage, the description adds critical semantic value: source of each id (bank_reconciliation_report, clearStep.uncleared, clearStep.candidateLines), exclusive usage of statementId vs asOf, default behavior of clearedDate, the 300-item limit, handling of cancelling pairs, and explicit warnings about unknown keys. This is far more than the schema descriptions offer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('file a pending draft'), a concrete resource ('bank clearing draft'), and the exact domain ('Banking → Clear transactions'). It clearly differentiates itself from the other create_*_draft siblings by focusing on bank clearing ticks and emphasizing that nothing changes until human approval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong workflow guidance: 'Read bank_reconciliation_report first and take every id from it' and 'read the report again and file afresh' if something changes. It does not explicitly list alternative tools or when not to use it, but the prerequisite and the detailed steps sufficiently orient an agent. One small gap is not naming a sibling to avoid (e.g., bank_reconciliation_status), so a 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_bill_draftFile bill draftAInspect
File a supplier-bill DRAFT into Taokeh from a bill you've read. INVENTORY-ONLY — every line must be a product you stock; for services / non-stock / mixed bills use the manual bill form, or create_expense_draft if already paid. This does NOT post to the books — it creates a pending draft the user reviews and approves in Taokeh; only then does it post (receive stock, blend moving-average cost, book input SST as a non-recoverable cost). Shape the fields with intake_contract(doc_type:'bill') + resolve_vendor + resolve_product first. The server re-computes every line quantity from the tally and the grand total — so present your working, but the server's figures are authoritative. READ THE PAYMENT TERM OFF THE BILL and send it as term — a supplier's credit period is what decides when this money actually leaves, so it drives the payables aging buckets, the OVERDUE badge and the 8-week cash forecast; for one of the five preset terms Taokeh fills the due date on the draft itself and shows it on the approval screen. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check — not lengthy reasoning) for any doubt. If you have the ORIGINAL bill image/PDF, attach it — it rides the draft and lands on the posted bill automatically on approval, so the user never has to re-upload it. Small files: attachmentBase64 + attachmentMediaType inline. Send attachmentBytes (the original file’s decoded size) with it so a truncated base64 is rejected instead of filed. Anything bigger: request_attachment_upload → PUT the bytes → pass the returned attachmentToken — unless your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist), in which case inline it anyway, with attachmentBytes; never both. ⛔ LINE KEYS ARE STRICT (2026-09-09): a key this schema does not list is REFUSED BY NAME — with the key it probably meant — and NOTHING is filed. Unknown keys used to be dropped in silence, which let a line through with its price or its tax code missing.
| Name | Required | Description | Default |
|---|---|---|---|
| term | No | The SUPPLIER's payment term, exactly as the bill prints it ('Due on Receipt', 'Net 7', 'Net 14', 'Net 30', 'Net 60', or whatever this supplier actually stated — 'COD 7 days', 'Net 45', '30 days EOM'). It is stored as written, so do NOT round a real term to the nearest familiar one. Leave it out when the bill states no term: an unstated term is not 'Due on Receipt', and the bill then simply falls due on its own date. | |
| lines | Yes | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before approving — a smudged or ambiguous total, a vendor/customer you weren't sure of, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. This is NOT a place to dump lengthy reasoning, your working, or boilerplate — just the one thing to check. Leave it empty when there is nothing to flag. | |
| vendor | No | ||
| dueDate | No | The due date, YYYY-MM-DD, when the bill states one outright. You usually do not need it: for one of the five preset terms above, Taokeh fills the due date on the draft itself (bill date + 0/7/14/30/60 days) and shows it on the approval screen for the owner to check. State dueDate only to override that, or when the paper gives a due date that does not follow the term. A term Taokeh does not recognise fills NOTHING — if that bill has a due date, say it here, because otherwise the bill falls due on its own date and shows up in the payables aging and the cash forecast a whole credit period early. | |
| vendorId | No | ||
| reference | No | ||
| needsReview | No | ||
| printedTotal | No | ||
| purchaseDate | Yes | ||
| paymentMethod | Yes | ||
| shoeboxItemId | No | The id of a SHOEBOX photo (from shoebox_items) that this draft was read from — the paper someone in the business snapped and sent in. Taokeh attaches ITS OWN stored copy of that photo to the draft, so DO NOT also send the image: no re-upload, no base64, nothing to truncate. Mutually exclusive with attachmentBase64 and attachmentToken. The item must still be waiting: one already booked, set aside, or already carrying a pending draft is refused with the reason. Approving the draft closes that shoebox item against the document it became. | |
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Use this instead of attachmentBase64 for any real photo/PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The uploaded file rides the draft and lands on the posted bill on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The ORIGINAL bill as base64 — SMALL files only (a few KB). For a real photo or multi-page PDF use request_attachment_upload instead (attachmentToken). Rides the draft and lands on the posted bill on approval — no re-upload. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| confirmNewVendor | No | Set to true ONLY after the user has EXPLICITLY confirmed this vendor is genuinely new. It skips the near-miss guard (which rejects a name that looks like shorthand for an existing vendor). Never set it to push a near-miss through — ask the user first. | |
| attachmentFilename | No | Optional original filename for the attachment, e.g. bill.pdf. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'image/jpeg' or 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no helpful annotations (all hints false), the description carries the full burden and does so thoroughly. It discloses that the draft does NOT post to the books until approved, explains the posting consequences (stock, moving-average cost, input SST), details the strict line-key rejection behavior, and describes attachment handling edge cases (truncated base64, sandbox restrictions). This level of transparency exceeds what annotations would provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section contributes essential operational detail. It is front-loaded with the most critical constraints (inventory-only, non-posting, line-key strictness) and branches into attachment handling and term semantics. While not concise in absolute terms, it is efficiently structured for a tool with 19 parameters and stringent rules; the density is justified, though a tighter organization could improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (19 params, strict validation, attachment workflows, server-side recomputation), the description covers nearly every scenario an agent might encounter: when to use which attachment method, how to handle measurement-profile companies, the need for reviewer notes, the handling of unknown terms, and even the sandboxed-client fallback. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 58%, but the description compensates richly. It explains the meaning of `term` (credit period driving payables aging), the derivation of quantity from thickness/width/tally on measurement-profile companies, the strict line-key rule, attachment token vs base64 trade-offs, and the `confirmNewVendor` guard. These insights go far beyond the schema's field descriptions, giving the agent deep understanding of each parameter's purpose and interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'File a supplier-bill DRAFT into Taokeh from a bill you've read.' It clearly identifies the tool's purpose and immediately distinguishes it from siblings by stating 'INVENTORY-ONLY' and pointing to alternatives (manual bill form, create_expense_draft). This leaves no ambiguity about what the tool does or how it differs from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool (inventory-only bills) and when not to (services/non-stock/mixed bills use manual form or create_expense_draft). It also prescribes the prerequisite workflow: 'Shape the fields with intake_contract(doc_type:'bill') + resolve_vendor + resolve_product first.' This provides clear, actionable guidance on selection and sequencing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_contact_draftFile contact draft from a name cardAInspect
File a pending CONTACT draft from a business card YOU have read. Read the card image yourself and pass the fields — personName (required), company, role, phones, emails, address — AND the card image itself, which is REQUIRED: the owner checks your fields against the picture before approving, so a draft without the image cannot be verified and is refused. This does NOT add anyone to the address book: it creates a draft the owner reviews and approves in Taokeh, and only that tap files the contact person (under the company, matched to an existing customer/vendor or created as a new one) with the card image kept on the record. BE HONEST about what you could not read: leave a field EMPTY and say so in notes — never guess a phone digit, an email spelling or a company name. If the photo is blurry, is not a business card, or lists two people, say that in notes and set needsReview. Attach the image with attachmentBase64 + attachmentMediaType (a card photo is usually small enough to inline; send attachmentBytes with it so a truncated base64 is rejected instead of filed), or for a large photo or a PDF use request_attachment_upload → PUT the bytes → pass the returned attachmentToken (never both). Say which side of the book it is with partyKind ('customer' for someone you sell to, 'vendor' for someone you buy from) — default customer — and pass partyName if the user tells you the company is really an existing one under a different spelling. Report the fields back to the user in chat so they can spot a misread before they tap.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | The person's job title / role as printed, e.g. 'Sales Manager'. | |
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check against the image before approving — a phone number you could not read cleanly, two people on one card, uncertainty about which line is the company. The reviewer reads this on a small approval card, so keep it brief and human. Leave it empty when there is nothing to flag. | |
| emails | No | The email addresses printed on the card (up to 3). Never guess a spelling; leave it out and flag it in notes instead. | |
| phones | No | The phone numbers printed on the card, as printed (up to 3). Leave out any digit-group you cannot read cleanly and say so in notes — a half-guessed phone number is worse than none. | |
| address | No | The address printed on the card, as one string. | |
| company | No | The company name printed on the card. Leave it out if the card does not show one. | |
| partyKind | No | Which side of the book the card belongs to: 'customer' (someone you sell to) or 'vendor' (someone you buy from). Defaults to customer. | |
| partyName | No | The company to file the contact under, when the user tells you it differs from what is printed on the card (e.g. the card shows a brand but the books use the registered name). Omit to use the company printed on the card — the reviewer can still change it before approving. | |
| personName | Yes | REQUIRED — the person's name exactly as printed on the card. If the name is genuinely illegible, do NOT guess and do NOT file: tell the user the card is unreadable. | |
| needsReview | No | Set true when something about the card gave you pause — a blurry photo, two people on one card, a field you could not read. It flags the draft for the reviewer. | |
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the image bytes to its uploadUrl. Use this instead of attachmentBase64 for a large photo or a PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The card lands on the created contact on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The business-card IMAGE as base64 — REQUIRED (unless you pass attachmentToken). A card photo is normally small enough to inline. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | Optional original filename for the card image, e.g. card.jpg. | |
| attachmentMediaType | No | The card image's MIME type, e.g. 'image/jpeg' or 'image/png'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false, so the description carries the full burden. It discloses that the image is REQUIRED or the draft is refused, that approval is by the owner, that truncated base64 is rejected via attachmentBytes, that bad type/oversize files are rejected with nothing filed, and that attachmentToken and attachmentBase64 are mutually exclusive. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded with the core purpose and workflow. Some redundancy (e.g., honesty warning appears in both description and schema) slightly bloats it, but the density of operational guidance justifies the length for a 16-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-param write tool with no output schema and minimal annotations, the description covers the full invocation workflow: reading the card, field extraction, honesty rules, attachment paths, partyKind/partyName, and post-call reporting. Nothing essential for correct invocation appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even with 100% schema coverage, the description adds critical meaning: the image is required despite not being in the schema's 'required' list, the token/base64 exclusivity, and why attachmentBytes should accompany base64. It also clarifies partyName and needsReview semantics beyond their schema entries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('File a pending CONTACT draft') and resource ('from a business card YOU have read'), and explicitly distinguishes itself from adding to the address book. It names the draft role and the sibling domain (create_*_draft, update_contact_draft) implicitly via the workflow, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use ('after reading a card image yourself'), what not to do ('does NOT add anyone to the address book'), and alternative routing for large photos/PDFs via request_attachment_upload → PUT → attachmentToken instead of inline base64. Also instructs to report fields back to the user so misreads are caught before approval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_credit_note_draftFile credit-note draftAInspect
File a CREDIT NOTE draft into Taokeh — a sales return, short delivery, price correction or allowance that reduces what a customer owes. This does NOT post to the books: it creates a pending draft the user reviews and approves in Taokeh, and only then does it reduce the customer's balance, reverse the revenue and SST, and (if the user says the goods came back) put stock back — and then only for the lines the ORIGINAL invoice actually shipped, and never more than it shipped: a line that moved no goods is money only, and claiming otherwise is refused at approval rather than posted. Shape the fields with intake_contract(doc_type:'credit_note') first. RESOLVE THE ORIGINAL INVOICE FIRST: a credit note is issued AGAINST an invoice — find it with search_documents and pass its id as originalDocId with kind:'against_invoice'. Only a credit with NO source invoice is kind:'allowance', and an allowance is never linked to an invoice, never capped by one, inherits no SST codes and never restocks — so do not file a real return as one; ask the user which invoice it is against. Send all amounts POSITIVE (Taokeh stores the credit note negative itself). A line can only credit back as MANY units as that invoice actually sold, net of earlier credit notes — an over-quantity line is refused naming what is left, so read the invoice's own unit of measure rather than converting it (1 carton is not 100 pieces). Whether goods physically came back into stock is the USER's decision at approval — goodsReturned is only a hint that pre-ticks their checkbox on the full review page, and a one-tap approval always posts money-only. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check) for any doubt. Filed it wrong? Use revise_draft (kind: 'credit_note') rather than filing a second one. IS A CREDIT NOTE EVEN THE RIGHT DOCUMENT? Only when something CHANGED AFTER the sale — goods came back, a price was renegotiated, a discount was agreed later. If the sale simply never happened (keyed twice, wrong company, cancelled before delivery), the owner VOIDS the invoice instead: Sales → the invoice → Void, in Taokeh. There is deliberately no AI lane for voiding. And if the sale did happen but the invoice was typed wrong — wrong line, wrong quantity or price, wrong date — that is a correction, not a credit: use update_invoice_draft. And if a credit note ALREADY POSTED was itself typed wrong (a wrong line, quantity, price or date), do not file a second one on top — correct it with update_credit_note_draft. ⛔ LINE KEYS ARE STRICT (2026-09-09): a key this schema does not list is REFUSED BY NAME — with the key it probably meant — and NOTHING is filed. Unknown keys used to be dropped in silence, which let a line through with its price or its tax code missing.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | REQUIRED. 'against_invoice' = a return / short delivery / price correction on a SPECIFIC invoice (the usual case — give originalDocId or originalDocNumber). 'allowance' = a standalone money credit with NO source invoice (give the customer instead). They are different documents; when it isn't clear, ask the user which invoice it is against. | |
| lines | Yes | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, in the reviewer's language, naming what the human should double-check before approving — whether the goods actually came back, an ambiguous quantity, which invoice you matched it to. Not lengthy reasoning. | |
| cnDate | Yes | ||
| reason | No | Why the credit is given — 'damaged goods returned', 'short delivery', 'agreed price adjustment'. It prints on the credit note. | |
| customer | No | allowance ONLY: the customer name, resolved with resolve_customer first. A credit note reduces an EXISTING customer's balance and can never create a customer. For an against_invoice credit note the customer comes from the invoice — leave this out. | |
| reference | No | The credit-note number already printed on paper, if any. Leave it out and Taokeh numbers it (CN…). | |
| customerId | No | allowance: the resolved customer id instead of the name. | |
| needsReview | No | ||
| printedTotal | No | The total printed on the paper, if any. Cross-check only — the server computes the real total and flags a mismatch on the approval screen. | |
| goodsReturned | No | An ADVISORY hint only: it pre-ticks the reviewer's 'the goods came back into stock' checkbox. It does NOT decide the stock movement — the human does, at approval, and even then only the lines the ORIGINAL invoice actually shipped can come back, up to the quantity it shipped; claiming more is refused rather than posted. Ignored for an allowance (always money-only). | |
| originalDocId | No | against_invoice: the ORIGINAL invoice's id, from search_documents. The server verifies it exists, is an INVOICE and whose it is — it never guesses. | |
| originalDocNumber | No | against_invoice: the printed invoice number, when you have no id. Refused if more than one document carries it — find the right one with search_documents. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint false, etc.), so the description carries full burden. It discloses many behaviors: non-posting nature ('creates a pending draft the user reviews and approves'), refusal behaviors ('over-quantity line is refused naming what is left'), strict line-key enforcement ('a key this schema does not list is REFUSED BY NAME'), and the advisory nature of goodsReturned. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very long but well-structured with clear sections (purpose, alternatives, line keys, parameter guidance). Every sentence carries operational significance; however, some redundancy exists (e.g., repeated emphasis on money-only vs stock return). Given the tool's complexity (13 params, many edge cases), the length is justified, but a bit of trimming would improve conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 13 parameters, 77% schema coverage, and no output schema, the description is exceptionally complete. It covers all parameter semantics, error scenarios, alternative tool routing, and approval flow details. Nothing essential for an agent to call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 77%, so description must add value. It does: explains positive amounts ('Send all amounts POSITIVE'), advises leaving taxCode out to inherit original invoice code, clarifies clientAmount is 'never used to compute anything', defines description as 'TEXT, not a field', and distinguishes originalDocId vs originalDocNumber with refusal conditions. Adds substantial meaning beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool's purpose: 'File a CREDIT NOTE draft into Taokeh — a sales return, short delivery, price correction or allowance that reduces what a customer owes.' It clearly distinguishes from siblings by naming alternatives like update_credit_note_draft, revise_draft, and update_invoice_draft, and explains what it does NOT do ('does NOT post to the books').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use/when-not-to-use guidance: 'IS A CREDIT NOTE EVEN THE RIGHT DOCUMENT? Only when something CHANGED AFTER the sale...' and then lists alternatives for voiding (no AI lane), corrections (update_invoice_draft), and fixing posted errors (update_credit_note_draft). Also instructs to use revise_draft for mistakes. This is comprehensive and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_debit_note_draftFile debit-note draftAInspect
File a SELL-SIDE DEBIT NOTE draft into Taokeh — an ADDITIONAL CHARGE on an invoice this business already issued (an undercharge, a price revised upward after delivery, a surcharge that was missed). This does NOT post to the books: it creates a pending draft the user reviews and approves in Taokeh, and only then does it increase what the customer owes and the output SST. Shape the fields with intake_contract(doc_type:'sales_debit_note') first. THE ORIGINAL INVOICE IS REQUIRED: find it with search_documents and pass its id as originalDocId. There is no allowance or standalone shape here — unlike a credit note, a debit note with no invoice behind it is not a document; if there is nothing to correct upward, the right document is a NEW INVOICE (create_invoice_draft), so say that rather than filing this. CHARGE THE SHORTFALL, NOT THE NEW PRICE: unitAmount is the amount being ADDED per unit, and quantity defaults to 1 because an additional charge is usually one line of money. Every line must name a product that actually appears on that invoice; anything else is refused. MONEY ONLY, ALWAYS — a debit note moves no stock, touches no COGS, and has no restock option anywhere in Taokeh (there is no goodsReturned field on this tool and no checkbox on its review screen), so a one-tap approval is money-only by construction and not by fence. If EXTRA GOODS were delivered, raise a new invoice instead. There is NO cap on the charge, so the approver is shown the original invoice's own total beside yours — keep the figure defensible. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check) for any doubt. Filed it wrong? Use revise_draft (kind: 'sales_debit_note') rather than filing a second one. This is the SELL side; Taokeh's /debit-notes page is the separate BUY side (a purchase return against a supplier bill), which has no AI lane for CREATING one — a buy-side note already posted can only be corrected, with update_supplier_debit_note_draft. A sell-side debit note already posted has no correction lane at all: the owner voids it in Taokeh and a new one is filed. ⛔ LINE KEYS ARE STRICT (2026-09-09): a key this schema does not list is REFUSED BY NAME — with the key it probably meant — and NOTHING is filed. Unknown keys used to be dropped in silence, which let a line through with its price or its tax code missing.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, in the reviewer's language, naming what the human should double-check before approving — which invoice you matched it to, where the figure came from. Not lengthy reasoning. | |
| dnDate | Yes | ||
| reason | No | Why more is charged — 'price revised upward after delivery', 'surcharge missed'. It prints on the note. | |
| customer | No | OPTIONAL CROSS-CHECK: the customer you believe the invoice belongs to. The customer comes FROM the invoice; sending a name that does not match is refused, which is exactly what it is for. | |
| reference | No | The debit-note number already printed on paper, if any. Leave it out and Taokeh numbers it (DN…). | |
| customerId | No | The resolved customer id, as the same cross-check. | |
| needsReview | No | ||
| printedTotal | No | The total printed on the paper, if any. Cross-check only — the server computes the real total and flags a mismatch on the approval screen. | |
| originalDocId | No | REQUIRED (this or originalDocNumber): the ORIGINAL invoice's id, from search_documents. The server verifies it exists, is an INVOICE and whose it is — it never guesses, and it will not file a debit note without one. | |
| originalDocNumber | No | The printed invoice number, when you have no id. Refused if more than one document carries it — find the right one with search_documents. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide basic hints (readOnly=false, etc.), so the description carries the full burden. It discloses that this does NOT post to the books, requires an original invoice, has no stock/COGS impact, no cap on charge, strict line-key rejection, and money-only semantics. No contradiction with annotations; it adds substantial behavioral context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place, covering purpose, usage, pitfalls, and parameter nuances. It is front-loaded with the core action and critical requirements, then layers alternatives and warnings. No fluff or repetition; the density is justified by the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (11 params, no output schema, multiple sibling tools), the description is remarkably complete. It covers all prerequisites, required fields, error conditions, alternatives, and the exact behavioral flow. An agent could call this tool correctly with no additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 73% schema coverage, the description adds critical semantic meaning: unitAmount is the additional shortfall, not the full price; quantity defaults to 1; taxCode inherits from the original invoice; productRef must appear on the invoice; customer is a cross-check. This goes far beyond the schema's own descriptions and prevents misuse.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('File') and resource ('SELL-SIDE DEBIT NOTE draft'), and precisely defines what it does: an ADDITIONAL CHARGE on an already-issued invoice. It distinguishes from credit notes and new invoices, making the purpose unmistakable and differentiating it from sibling draft tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use and when-not-to-use guidance: names alternatives like create_invoice_draft for new charges, revise_draft for corrections, and update_supplier_debit_note_draft for the buy side. It also states the prerequisite (original invoice required) and the condition for refusing the tool (no upward correction needed).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_expense_draftFile expense draftAInspect
File a paid-expense DRAFT into Taokeh from a receipt you've read. This does NOT post to the books — it creates a pending draft the user reviews and approves in Taokeh; only then does it hit the ledger. Shape the fields with intake_contract + expense_accounts first. Set amountUncertain when any digit of the printed total is uncertain; set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what the human should double-check — not lengthy reasoning) for any other doubt. If you have the ORIGINAL receipt image/PDF, attach it — it rides the draft and lands on the posted entry automatically on approval, so the user never has to re-upload it. Small files: pass attachmentBase64 + attachmentMediaType inline. Send attachmentBytes (the original file’s decoded size) with it so a truncated base64 is rejected instead of filed. Anything bigger: request_attachment_upload → PUT the bytes → pass the returned attachmentToken — unless your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist), in which case inline it anyway, with attachmentBytes; never both.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| memo | No | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before approving — a smudged or ambiguous total, a vendor/customer you weren't sure of, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. This is NOT a place to dump lengthy reasoning, your working, or boilerplate — just the one thing to check. Leave it empty when there is nothing to flag. | |
| amount | Yes | ||
| currency | No | ||
| inputTax | No | ||
| reference | No | ||
| needsReview | No | ||
| shoeboxItemId | No | The id of a SHOEBOX photo (from shoebox_items) that this draft was read from — the paper someone in the business snapped and sent in. Taokeh attaches ITS OWN stored copy of that photo to the draft, so DO NOT also send the image: no re-upload, no base64, nothing to truncate. Mutually exclusive with attachmentBase64 and attachmentToken. The item must still be waiting: one already booked, set aside, or already carrying a pending draft is refused with the reason. Approving the draft closes that shoebox item against the document it became. | |
| amountUncertain | No | ||
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Use this instead of attachmentBase64 for any real photo/PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The uploaded file rides the draft and lands on the posted entry on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The ORIGINAL receipt as base64 — SMALL files only (a few KB). Base64 inside a tool call is costly, so for a real receipt photo or a multi-page PDF use request_attachment_upload instead (attachmentToken). Rides the draft and lands on the posted entry on approval — no re-upload. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| debitAccountCode | Yes | ||
| creditAccountCode | Yes | ||
| attachmentFilename | No | Optional original filename for the attachment, e.g. receipt.jpg. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'image/jpeg' or 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the minimal annotations, the description discloses key side effects and constraints: no ledger posting until approval, automatic attachment propagation, shoebox item refusal if already booked/pending, and rejection of truncated or bad-type uploads. It also explains the sandboxed network fallback for inline uploads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The information is front-loaded with purpose and primary behavior, then flows from field shaping to flags to attachment modes. It is dense and occasionally repeats the 'lands on the posted entry' benefit, but every major decision rule earns its place for an 18-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex create tool with no output schema and sparse annotations, the description covers the full life cycle: draft creation, review flags, shoebox linkage, attachment handling, and network fallback. An agent has enough context to select it and invoke the correct attachment path without further tool exploration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 44%, but the description compensates for the tricky parameters: amountUncertain, needsReview/notes, attachmentBase64/attachmentBytes/attachmentToken, and shoeboxItemId mutual exclusivity. A few required fields (date, amount, account codes) are left to the referenced intake_contract/expense_accounts workflow rather than explained inline, so it is strong but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb ('File'), target (Taokeh), scope (a paid-expense DRAFT), and source ('from a receipt you've read'). It immediately distinguishes the tool from posting to the books by explaining it creates a pending user-reviewed draft, which differentiates it from ledger-posting and other draft siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit directions: use intake_contract + expense_accounts first, set amountUncertain only when digits are uncertain, set needsReview for other doubts, and choose attachments by size/network conditions ('Small files... inline', 'Anything bigger: request_attachment_upload'). The shoeboxItemId branch and 'never both' exclusion also state exactly when not to send an image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_invoice_draftFile invoice draftAInspect
File a sales-invoice DRAFT into Taokeh from a sales note you've read. This does NOT post to the books — it creates a pending draft the user reviews and approves in Taokeh; only then does it post (and move stock). Shape the fields with intake_contract(doc_type:'invoice') + resolve_customer + resolve_product first. The server re-computes every line quantity from the tally and the grand total (with SST) — so present your working, but the server's figures are authoritative. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check — not lengthy reasoning) for any doubt. If you have the ORIGINAL sales note image/PDF, attach the original document you extracted from — the owner sees it beside the draft at review (s.82 record-keeping) and it lands on the posted invoice automatically on approval, so the user never has to re-upload it. Small files: attachmentBase64 + attachmentMediaType inline. Send attachmentBytes (the original file’s decoded size) with it so a truncated base64 is rejected instead of filed. Anything bigger: request_attachment_upload → PUT the bytes → pass the returned attachmentToken — unless your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist), in which case inline it anyway, with attachmentBytes; never both. ⛔ LINE KEYS ARE STRICT (2026-09-09): a key this schema does not list is REFUSED BY NAME — with the key it probably meant — and NOTHING is filed. Unknown keys used to be dropped in silence, which let a line through with its price or its tax code missing.
| Name | Required | Description | Default |
|---|---|---|---|
| term | No | The payment term exactly as the paper states it ('Due on Receipt', 'Net 7', 'Net 14', 'Net 30', 'Net 60'). Anything else the note actually says is kept verbatim and shown to the reviewer as-is — do not round it to the nearest familiar term. Leave it out when the paper states none. | |
| lines | Yes | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before approving — a smudged or ambiguous total, a vendor/customer you weren't sure of, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. This is NOT a place to dump lengthy reasoning, your working, or boilerplate — just the one thing to check. Leave it empty when there is nothing to flag. | |
| refNo | No | The CUSTOMER's own document number as printed on the note — their delivery-order number or purchase-order number ("DO 1074", "PO 88231"). NOT the invoice number: that is `reference`. Leave it out when the paper shows none; never copy the invoice number into it. | |
| dueDate | No | The due date when the paper states one outright. For one of the five preset terms you can leave it out — Taokeh fills the due date on the draft (invoice date + 0/7/14/30/60 days) and the owner sees it on the approval screen. Send it to override that, or when the paper's due date does not follow the term. A term Taokeh does not recognise derives NOTHING, so state the due date yourself on those or the invoice posts due on its own issue date. | |
| customer | No | ||
| saleDate | Yes | ||
| reference | No | ||
| customerId | No | ||
| needsReview | No | ||
| printedTotal | No | ||
| paymentMethod | Yes | ||
| shoeboxItemId | No | The id of a SHOEBOX photo (from shoebox_items) that this draft was read from — the paper someone in the business snapped and sent in. Taokeh attaches ITS OWN stored copy of that photo to the draft, so DO NOT also send the image: no re-upload, no base64, nothing to truncate. Mutually exclusive with attachmentBase64 and attachmentToken. The item must still be waiting: one already booked, set aside, or already carrying a pending draft is refused with the reason. Approving the draft closes that shoebox item against the document it became. | |
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Use this instead of attachmentBase64 for any real photo/PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The uploaded file rides the draft and lands on the posted invoice on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The ORIGINAL sales note as base64 — SMALL files only (a few KB). For a real photo or multi-page PDF use request_attachment_upload instead (attachmentToken). Rides the draft and lands on the posted invoice on approval — no re-upload. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | Optional original filename for the attachment, e.g. sales-note.jpg. | |
| confirmNewCustomer | No | Set to true ONLY after the user has EXPLICITLY confirmed this customer is genuinely new. It skips the near-miss guard (which rejects a name that looks like shorthand for an existing customer). Never set it to push a near-miss through — ask the user first. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'image/jpeg' or 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnly=false, openWorld=false, idempotent=false, destructive=false—no behavioral detail. The description carries the full burden and delivers: server re-computes quantities and totals, strict line key refusal (by name), attachment truncation rejection, mutual exclusivity of attachment methods, and the shoebox workflow. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with critical information. It front-loads the core concept and then organizes warnings and nuances in a logical order. Every sentence earns its place, though the length could intimidate an agent. It is not redundant, but slightly overlong for the platform's typical expectations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 20 parameters, 3 required, no output schema, and high complexity, the description covers prerequisites, server behavior, attachment handling, strict validation, and the review workflow. It tells the agent exactly what to expect (a pending draft for user review) and how to handle edge cases (truncated base64, sandboxed network). Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, but the description adds substantial meaning beyond the schema: it explains how quantity is derived from tally on measurement-profile companies, clarifies that unitPrice is required per line, describes the strict line keys, and elaborates on attachment fields (attachmentBytes prevents truncation, attachmentSha256 as optional check). It also clarifies refNo vs reference. This goes well beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action ('File a sales-invoice DRAFT') and the resource (into Taokeh), and immediately differentiates from posting to the books. Explicitly distinguishes from siblings by naming prerequisite tools (intake_contract, resolve_customer, resolve_product) and the attachment upload alternative (request_attachment_upload).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit sequencing: shape fields with intake_contract + resolve_customer + resolve_product first. Gives conditional instructions for attachments (inline for small, request_attachment_upload for large, never both). Also tells when to omit quantity and how to handle strict line keys, leaving no ambiguity about when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_journal_draftFile adjusting journal draftAInspect
File an ADJUSTING JOURNAL ENTRY as a draft in Taokeh — the accountant's entry, not a document: accruals, prepayments, depreciation, corrections, reclassifications, year-end adjustments, and REVERSING entries. This does NOT post to the books: it creates a pending draft with the debits and the credits laid out for the user to see, and only their approval writes it to the ledger. Shape the fields with intake_contract(doc_type:'journal') first — it returns this company's REAL chart of accounts, which is what every accountRef must match. IT MUST BALANCE TO THE SEN: debits equal credits or the draft is refused, with the difference named. Each line carries exactly one side (debit OR credit), and at least two lines are required. Accounts are matched by CODE (best) or by their exact name; an unknown ref is refused and an ambiguous name reports the candidates rather than picking. THIS TOOL NEVER CREATES AN ACCOUNT — if the adjustment needs one that does not exist, tell the user to add it in Taokeh and file again. FIXING A WRONG ENTRY THAT IS ALREADY POSTED: file a REVERSING entry here. Taokeh's connector never edits and never deletes posted history — that is deliberate and it is the point: a ledger you can rewrite is not a ledger, and every entry must stay auditable. So find the wrong entry with search_journal, file a reversing journal dated today (its debits become credits and its credits become debits, same accounts, same amounts), then file the correct entry. Tell the user plainly that this is what you are doing and why, rather than reporting that you 'cannot' fix it. CONTROL ACCOUNTS ARE ALLOWED BUT NEVER SILENT: a line on accounts receivable, accounts payable, inventory, the SST control or opening-balance equity is reconciled to a subledger, so a journal there moves the control with no document behind it and the aging-vs-control tie-out will show the difference. Such a draft is filed and flagged — it can be approved ONLY on its full review page in Taokeh, never by a one-tap email approval or a deck swipe. Say in notes why the control line is right. WHAT NOT TO USE THIS FOR: anything that has a real document. A supplier bill is create_bill_draft, a sale is create_invoice_draft, a paid expense is create_expense_draft, a customer payment is create_receipt_draft. Those carry party, SST and stock consequences a raw journal silently skips. Attach the WORKING PAPER you read the adjustment off — the depreciation schedule, the accrual computation, the bank letter. It rides the draft and the approver sees it beside your figures. Prefer request_attachment_upload → attachmentToken; if your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist) send attachmentBase64 + attachmentMediaType inline instead — correct even for a full PDF — always with attachmentBytes, the file’s decoded size on disk, so a truncated paste is rejected instead of filed. Never both. Set needsReview and add a SHORT reviewer note in notes for any doubt. Filed it wrong? Use revise_draft (kind: 'journal') rather than filing a second one.
| Name | Required | Description | Default |
|---|---|---|---|
| memo | No | What the entry is FOR, in one line — "accrue December electricity", "reverse the duplicated August rent". It prints on the entry and is the first thing the approver reads. | |
| lines | Yes | At least two lines. The debits must equal the credits TO THE SEN — an unbalanced entry is refused with the difference named. ⛔ LINE KEYS ARE STRICT: a key this schema does not list is REFUSED BY NAME and NOTHING is filed. | |
| notes | No | A SHORT reviewer note: one or two plain sentences, in the reviewer's language, naming what the human should double-check before approving — which schedule the figure came from, which entry this reverses, why a control-account line is right. Not lengthy reasoning. | |
| entryDate | Yes | ||
| reference | No | Your own reference for the adjustment, if there is one (a schedule number, a working-paper ref). | |
| clientTotal | No | Your own arithmetic for the entry total — carried onto the review screen for the human to compare against, never used to compute anything. | |
| needsReview | No | ||
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Mutually exclusive with attachmentBase64. | |
| attachmentBase64 | No | The supporting WORKING PAPER as base64 — SMALL files only. For a real schedule or a multi-page PDF use request_attachment_upload instead (attachmentToken). A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | The original filename, for the reviewer. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are sparse (all false), so the description carries the full burden — and it delivers richly. It discloses that the tool does NOT post to the books, that drafts must balance to the sen or be refused, that accounts are matched by code/exact name, that unknown refs are refused, that ambiguous names are reported rather than picked, that the tool never creates an account, and that posted history is never edited or deleted by design. This goes far beyond what annotations express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but it is organized into clearly scoped sections and every sentence carries operational meaning for a complex, high-stakes financial action. The core semantics are front-loaded in the first sentence, and the rest is structured around balancing rules, control accounts, exclusions, attachment handling, and correction workflow. It is dense rather than padded, though a slightly tighter organization would improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 13 parameters, no output schema, and zero clue-giving annotations, the description is unusually complete. It covers entry behavior, balancing constraints, account matching rules, control-account approval consequences, attachment handling with fallbacks, revision via revise_draft, and explicit alternatives. An agent has enough context to invoke the tool correctly in both normal and edge-case scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is high at 85%, but the description adds significant meaning beyond the schema: accountRef should be a chart-of-accounts code or exact name, never a guess; lines must carry exactly one side and at least two lines are required; attachmentToken and attachmentBase64 paths are mutually exclusive with conditions on reachability of taokeh.my; attachmentBytes should be the decoded on-disk size to reject truncated uploads; notes should be a short reviewer note, not lengthy reasoning. This is exactly the kind of parameter-level context an agent needs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'File an ADJUSTING JOURNAL ENTRY as a draft in Taokeh — the accountant's entry, not a document.' It immediately distinguishes the tool from document-based entry tools and lists the exact use cases (accruals, prepayments, depreciation, corrections, reclassifications, year-end adjustments, reversing entries). An agent cannot mistake this for create_bill_draft or create_invoice_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance and names sibling alternatives with conditions: 'WHAT NOT TO USE THIS FOR: anything that has a real document. A supplier bill is create_bill_draft, a sale is create_invoice_draft...' It also prescribes the reversing-entry workflow with search_journal, and tells the agent to say plainly what it is doing rather than reporting it 'cannot' fix an entry.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_payment_draftFile a pending supplier-payment draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. File a SUPPLIER-PAYMENT draft ("we paid Ah Seng RM5,000 from Maybank on Tuesday against bills 12 and 14"). This does NOT post: only the owner's approval writes a real bank payment and settles the bills, and when it does, accounts payable and the bank move exactly as they would if the owner had matched the line in Banking. A payment settles an EXISTING supplier's open bills: resolve the supplier (resolve_vendor) and see what is owed (ap_aging, search_documents) first, then give EITHER a lump total (auto-allocated oldest-first) OR explicit per-bill allocations. PREFER purchaseId over reference in an allocation: a supplier's own bill number is NOT unique in Taokeh, so a reference that matches two open bills is refused naming both rather than guessed. MYR only — a foreign-currency bill is refused and routes to Banking → Match, which handles the exchange difference. Amounts are RINGGIT, the same figures ap_aging and search_documents show you — never sen, never a foreign figure converted by you. A CREDIT CARD can't be the paying account here: paying a supplier on a card is a spend that arrives on the card's own statement — import that statement instead. IF THIS COMPANY IMPORTS BANK STATEMENTS and the payment is already sitting on an imported line, do NOT file here — propose the allocation on that line with draft_bank_classification. Filing it here would put the same money in Taokeh twice, and the server refuses when it can see the imported line. The server re-derives the allocation against live outstanding at create AND at approval, so its figures are authoritative and a bill settled in the meantime makes the draft refuse rather than over-pay. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check — not lengthy reasoning) for any doubt. Filed it wrong? Use revise_draft (kind: 'payment') rather than filing a second one. If you have the payment proof (transfer slip, remittance advice), attach the original document you extracted from — the owner sees it beside the draft at review (s.82 record-keeping) and it lands on the settlement entry automatically on approval. Small files: attachmentBase64 + attachmentMediaType inline. Send attachmentBytes (the original file’s decoded size) with it so a truncated base64 is rejected instead of filed. Anything bigger: request_attachment_upload → PUT the bytes → pass the returned attachmentToken — unless your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist), in which case inline it anyway, with attachmentBytes; never both.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before approving — a smudged or ambiguous total, a supplier you weren't sure of, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. This is NOT a place to dump lengthy reasoning, your working, or boilerplate — just the one thing to check. Leave it empty when there is nothing to flag. | |
| total | No | A lump ringgit amount to auto-allocate oldest-first across the supplier’s open bills. Give this OR explicit allocations — never both. | |
| vendor | No | The supplier name, resolved with resolve_vendor first. A payment settles an EXISTING supplier's bills and can never create a supplier. | |
| vendorId | No | The resolved vendor id, instead of the name. | |
| allocations | No | The explicit per-bill split, at most 50 lines. Each amount is capped at that bill’s live outstanding by the server. | |
| needsReview | No | ||
| paymentDate | Yes | ||
| printedTotal | No | The total printed on the payment advice, if any. Cross-check only — the server computes the real total and flags a mismatch on the approval screen. | |
| bankAccountId | No | The PAYING bank/cash account (from intake_contract(doc_type:'payment') → payingAccounts). Optional — omit it and the reviewer picks at approval. A credit-card account is refused by name. | |
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Use this instead of attachmentBase64 for any real photo/PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The uploaded file rides the draft and lands on the settlement entry on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The ORIGINAL payment proof (transfer slip / remittance advice) as base64 — SMALL files only (a few KB). For a real photo or multi-page PDF use request_attachment_upload instead (attachmentToken). Rides the draft and lands on the settlement entry on approval — no re-upload. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | Optional original filename for the attachment, e.g. transfer-slip.jpg. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'image/jpeg' or 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations by explaining that the draft does NOT post until the owner approves, that the server re-derives allocations at create and approval time, that duplicate imported-line filings are refused, and that truncated attachments are rejected. These are critical behavioral details that the sparse annotations do not convey, and there is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the most important safety constraint and nearly every sentence carries functional value. However, it is a dense wall of text with heavy all-caps and long parenthetical asides, which makes parsing harder than necessary; it could be tightened or structured without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter mutation tool with no output schema, the description is remarkably complete. It covers prerequisites, allocation rules, currency constraints, credit-card exclusions, the imported-statement duplicate risk, attachment mechanics, review notes, error handling, and what happens at approval. Nothing essential for selecting or invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is already high at 87%, the description adds substantial meaning to parameters: purchaseId should be preferred over reference because supplier bill numbers are not unique, total and allocations are mutually exclusive, attachmentBytes detects truncated base64, and notes should carry only a short reviewer flag. This materially improves correct invocation beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'FILES A PENDING DRAFT ONLY' and 'File a SUPPLIER-PAYMENT draft'. It clearly distinguishes this tool from the many create_* siblings and from draft_bank_classification by stating exactly what kind of draft it files and that human approval is required before anything posts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage guidance is explicit and thorough. It names concrete alternatives and their trigger conditions: use draft_bank_classification if the payment is already on an imported bank line, use revise_draft if filed wrong, and route foreign-currency or credit-card cases elsewhere. It also prescribes prerequisites like resolve_vendor, ap_aging, and search_documents before filing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_payment_match_draftPending draft: match a posted payment to open invoices or bills (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose how ONE bank line already posted as a customer payment (money in) or a supplier payment (money out) settles open invoices or bills — the owner's "this deposit paid these three invoices". An admin or bookkeeper reviews it and approves; only that tap allocates, exactly as Banking → Match → Save does, and the documents then read as paid. Pass bankTransactionId (a POSTED line from bank_review_queue with status "posted") and allocations: [{saleId, amount}] for money in, [{purchaseId, amount}] for money out, ringgit, adding up to the WHOLE line to the sen (the in-app pane's ±0.10 rounding allowance does not apply to a proposal). At most 50 documents. WHICH DOCUMENTS: only the ones Banking → Match offers this line — the open invoices of the customer the line is posted against (or, when that customer has none open, the invoices of the company's walk-in customer — the one the owner marked on its customer page, else a customer named 'CASH SALES'), or the supplier's open bills. A line with no customer at all is offered nothing. A document of ANY other party is refused, because no screen in Taokeh offers it: if the payment was posted against the wrong customer, or no customer, the owner unposts the line in Banking, sets the right customer, and posts it again — then file this. A refusal lists the documents the line can take, with their ids. ALSO REFUSED BY NAME, at filing and again at approval: a line not posted, not a customer/supplier payment, already matched (a match someone saved is never replaced), or on a foreign-currency account (the owner settles those in the app); a document not open; an amount over what a document still owes; a document dated AFTER the payment (an advance is allocated by hand); the same document twice; a pending match already filed for the line. If the line or a document changes before approval (paid, credited, re-tagged, matched), approval refuses and nothing is allocated. note: one short line for the owner saying how you know (the remittance, the customer's message).
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | A SHORT note for the owner: how you know which documents it paid. | |
| lines | No | NOT a filter on this tool. Pass the documents as `allocations`: [{saleId or purchaseId, amount}]. | |
| notes | No | NOT a filter on this tool. Pass the owner note as `note`. | |
| bankLineId | No | NOT a filter on this tool. Pass the bank line as `bankTransactionId`. | |
| allocations | Yes | Where the money goes. An unknown key is refused by name. | |
| needsReview | No | Set true when something gave you pause. | |
| transactionId | No | NOT a filter on this tool. Pass the bank line as `bankTransactionId`. | |
| bankTransactionId | Yes | REQUIRED — the posted bank line, from bank_review_queue (status "posted"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations all false, the description carries the full burden and does so thoroughly. It discloses that nothing changes until human approval, that approval allocates exactly like Banking → Match → Save, that refusals occur both at filing and approval, and that any change to the line or document before approval causes refusal. This is rich behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place. It is front-loaded with the most critical behavioral caveat, then logically proceeds through purpose, parameters, eligible documents, refusal conditions, and the note. The structure is clear and the density is justified by the complex matching rules.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description comprehensively covers input requirements, eligibility, refusal conditions, and the approval workflow. However, it does not explicitly state what a successful response looks like (e.g., a draft identifier or confirmation), which is a minor gap given there is no output schema. The failure response is partially described (refusal lists documents), but success return is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds substantial meaning beyond the schema: it specifies that bankTransactionId must be a POSTED line from bank_review_queue with status 'posted', explains the allocations format with saleId vs purchaseId and the requirement to sum to the whole line, and explicitly warns that 'lines', 'notes', 'bankLineId', and 'transactionId' are NOT filters and should be passed under the correct parameter names. This is far beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, emphatic statement of what the tool does: 'FILES A PENDING DRAFT ONLY' and then specifies the exact operation: proposing how one posted bank line settles open invoices or bills. It distinguishes itself from sibling create_*_draft tools by emphasizing the human-approval step and the Banking → Match → Save equivalence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool (for a posted bank line from bank_review_queue), enumerates refusal conditions (not posted, already matched, foreign-currency, etc.), and provides a concrete alternative workflow (unpost, set the right customer, repost). This is explicit when/when-not guidance with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_purchase_order_draftFile purchase-order draftAInspect
File a PURCHASE-ORDER draft into Taokeh — an order you intend to PLACE with a supplier, typically read off their quotation. It POSTS NOTHING and MOVES NOTHING: a purchase order is an intent to buy, so approving it creates no journal entry, no stock movement and no amount owing. The books move later, when the goods arrive and the user converts the order into a bill on its page in Taokeh — a step this tool never takes for them. Approval also saves the order as a DRAFT; issuing it to the supplier is the user's own tap, and Taokeh emails nobody. INVENTORY-ONLY — every line must be a product the company stocks; a service or one-off non-stock charge belongs on the manual purchase-order form. NO SST on this document: a purchase order in Taokeh carries no tax code and no tax line, and input SST is settled on the bill against the supplier's tax invoice. Shape the fields with intake_contract(doc_type:'purchase_order') + resolve_vendor + resolve_product first. The server re-computes every line quantity from the tally and the order total — so present your working, but the server's figures are authoritative. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check — not lengthy reasoning) for any doubt. Filed it wrong? Use revise_draft (kind: 'purchase_order') rather than filing a second one. Goods that have ALREADY arrived are not this door — use create_bill_draft. ⛔ LINE KEYS ARE STRICT (2026-09-09): a key this schema does not list is REFUSED BY NAME — with the key it probably meant — and NOTHING is filed. Unknown keys used to be dropped in silence, which let a line through with its price or its tax code missing.
| Name | Required | Description | Default |
|---|---|---|---|
| term | No | The payment/delivery term line, e.g. 'Net 30'. | |
| lines | Yes | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before they order — a quantity you weren't sure of, a supplier name you guessed at, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. Leave it empty when there is nothing to flag. | |
| poDate | Yes | ||
| vendor | No | ||
| vendorId | No | ||
| reference | No | The BUYER'S OWN purchase-order number, if they already have one — it becomes this order's document number. NOT the supplier's quotation number (put that in `notes`). Usually omit it: Taokeh numbers the order itself (POxxxx). | |
| needsReview | No | ||
| expectedDate | No | When the goods are expected, YYYY-MM-DD. Optional — omit it rather than guess. | |
| printedTotal | No | ||
| confirmNewVendor | No | Set to true ONLY after the user has EXPLICITLY confirmed this supplier is genuinely new. It skips the near-miss guard (which rejects a name that looks like shorthand for an existing vendor). Never set it to push a near-miss through — ask the user first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false and generic, so the description carries the full burden — and meets it richly. It states that the tool 'POSTS NOTHING and MOVES NOTHING' (no journal entry, stock movement, or amount owing), that approval saves a DRAFT and emails nobody, that no SST is carried, and that the server recomputes authoritative quantities/totals. The strict line-key refusal behavior is also disclosed. No contradiction with annotations; filing a draft is a write, consistent with readOnlyHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but nearly every passage earns its place given the domain complexity: tax treatment, measurement-profile quantity derivation, workflow ordering, and error recovery. It is front-loaded with the core purpose and behavioral guarantees before the constraints. The heavy use of ALL-CAPS emphasis is stylistically loud but flags genuinely critical constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description covers the critical operational facts: result semantics ('saves the order as a DRAFT', 'the server's figures are authoritative'), failure behavior ('NOTHING is filed' on bad keys), prerequisites, alternatives, and the review workflow. Nothing an agent needs in order to call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 45% (below the 50% threshold), so the description needed to compensate — and it does. It adds workflow meaning around several parameters: the quantity/tally/thickness/width derivation rule, the SHORT reviewer note requirement for `notes`, the strict line-key behavior applying to `lines`, and the confirmNewVendor guard context. A few parameters (printedTotal, unit, working, vendorId) remain lightly documented, but the description compensates for coverage gaps better than the schema alone does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('File'), resource ('PURCHASE-ORDER draft'), and system ('Taokeh'), with the intent ('intend to PLACE with a supplier, typically read off their quotation'). It distinguishes itself from close siblings by name ('use create_bill_draft' for arrived goods, 'use revise_draft' for corrections), so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-not conditions and alternatives: 'Goods that have ALREADY arrived are not this door — use create_bill_draft', 'Filed it wrong? Use revise_draft (kind: \'purchase_order\') rather than filing a second one', and the inventory-only exclusion that routes services to the manual form. It also names the prerequisite shaping tools (intake_contract, resolve_vendor, resolve_product) and when to set needsReview.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_quote_draftFile quote draftAInspect
File a quotation DRAFT into Taokeh from a request you've read or been told. A quote is an ESTIMATE — it does NOT post to the books or move stock; it creates a pending draft the user reviews and approves in Taokeh, which posts a real quotation they can then convert to an invoice or delivery order. Shape the fields with resolve_customer + resolve_product first. The server re-computes every line quantity from the tally and the grand total (with SST) — so present your working, but the server's figures are authoritative. A quote may have no customer (a walk-in estimate). Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check — not lengthy reasoning) for any doubt. If you read the request off a document, attach the original document you extracted from — the owner sees it beside the draft at review (s.82 record-keeping) and it lands on the posted quote automatically on approval. Small files: attachmentBase64 + attachmentMediaType inline. Send attachmentBytes (the original file’s decoded size) with it so a truncated base64 is rejected instead of filed. Anything bigger: request_attachment_upload → PUT the bytes → pass the returned attachmentToken — unless your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist), in which case inline it anyway, with attachmentBytes; never both. ⛔ LINE KEYS ARE STRICT (2026-09-09): a key this schema does not list is REFUSED BY NAME — with the key it probably meant — and NOTHING is filed. Unknown keys used to be dropped in silence, which let a line through with its price or its tax code missing.
| Name | Required | Description | Default |
|---|---|---|---|
| term | No | Payment term carried onto the quote, e.g. 'Net 30'. Defaults to 'Due on Receipt'. | |
| lines | Yes | ||
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before approving — a smudged or ambiguous total, a vendor/customer you weren't sure of, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. This is NOT a place to dump lengthy reasoning, your working, or boilerplate — just the one thing to check. Leave it empty when there is nothing to flag. | |
| customer | No | ||
| quoteDate | Yes | ||
| reference | No | ||
| customerId | No | ||
| validUntil | No | Optional "valid until" date for the quotation (YYYY-MM-DD). | |
| needsReview | No | ||
| printedTotal | No | ||
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Use this instead of attachmentBase64 for any real photo/PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The uploaded file rides the draft and lands on the posted quote on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The ORIGINAL request/quote document as base64 — SMALL files only (a few KB). For a real photo or multi-page PDF use request_attachment_upload instead (attachmentToken). Rides the draft and lands on the posted quote on approval — no re-upload. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | Optional original filename for the attachment, e.g. rfq.pdf. | |
| confirmNewCustomer | No | Set to true ONLY after the user has EXPLICITLY confirmed this customer is genuinely new. It skips the near-miss guard (which rejects a name that looks like shorthand for an existing customer). Never set it to push a near-miss through — ask the user first. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'image/jpeg' or 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Very explicit about side effects: a draft does NOT post to books or move stock; the server recomputes quantities and totals (authoritative); strict field validation refuses unknown keys and files nothing; truncated base64 is rejected via attachmentBytes. This is beyond the annotations (readOnlyHint=false, destructiveHint=false) and adds critical expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but organized: starts with the core action, then field-shaping, then behavior, then attachment workflow. Some sentences are long, but the information is high-value and the order is logical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For 17 params with no output schema, it covers the main judgment areas: request extraction, reviewer notes, measurement profiles, attachment handling, and strictness. It doesn't describe the response payload or failure modes, but the behavioral notes suggest the server will reject with errors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 17 params, description clarifies the tricky ones: attachmentBytes' purpose (truncation check), needsReview guidance with example note length, measurement-profile handling of quantity vs tally/width/thickness, and attachmentBase64 vs attachmentToken mutual exclusivity. That's substantial added meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete verb ('File') and a clear resource ('a quotation DRAFT') and immediately distinguishes it from posting a real quotation ('does NOT post to the books or move stock'). It also names the downstream flow (review → approval → conversion to invoice/delivery order), so an agent knows exactly what this tool is for and what it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives practical guidance: resolve customer/product before shaping fields, attach original documents, set needsReview for doubts, inline small files vs use attachmentToken for real PDFs. It doesn't explicitly name sibling alternatives (e.g., 'use update_quote_draft when editing'), but the draft-vs-posted contrast and the attachment alternatives cover the main routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_receipt_draftFile receipt draftAInspect
File a customer-payment DRAFT into Taokeh ("customer paid X"). This does NOT post — it creates a pending draft the user reviews and approves in Taokeh; only then does it write a real bank receipt and settle the invoices. A receipt settles an EXISTING customer's open invoices: resolve the customer (resolve_customer) and see what they owe (open_invoices) first, then give EITHER a lump total (auto-allocated oldest-first) OR explicit per-invoice allocations. MYR only — a foreign-currency invoice is refused and routes to Banking → Payments. The server re-derives the allocation against live outstanding, so its figures are authoritative. Set needsReview and add a SHORT reviewer note in notes (one or two sentences naming what to double-check — not lengthy reasoning) for any doubt. If you have the payment proof (bank-in slip / remittance advice), attach the original document you extracted from — the owner sees it beside the draft at review (s.82 record-keeping) and it lands on the settlement entry automatically on approval. Small files: attachmentBase64 + attachmentMediaType inline. Send attachmentBytes (the original file’s decoded size) with it so a truncated base64 is rejected instead of filed. Anything bigger: request_attachment_upload → PUT the bytes → pass the returned attachmentToken — unless your shell cannot reach taokeh.my (a sandboxed client behind a network allowlist), in which case inline it anyway, with attachmentBytes; never both.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note: one or two plain sentences, written in the reviewer's language, flagging what the human should double-check before approving — a smudged or ambiguous total, a vendor/customer you weren't sure of, a judgment call you made. The reviewer reads this on a small approval card, so keep it brief and human. This is NOT a place to dump lengthy reasoning, your working, or boilerplate — just the one thing to check. Leave it empty when there is nothing to flag. | |
| total | No | ||
| customer | No | ||
| customerId | No | ||
| allocations | No | ||
| needsReview | No | ||
| receiptDate | Yes | ||
| printedTotal | No | ||
| bankAccountId | No | ||
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the file bytes to its uploadUrl. Use this instead of attachmentBase64 for any real photo/PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The uploaded file rides the draft and lands on the settlement entry on approval, exactly as an inline one does. | |
| attachmentBase64 | No | The ORIGINAL payment proof (bank-in slip / remittance advice) as base64 — SMALL files only (a few KB). For a real photo or multi-page PDF use request_attachment_upload instead (attachmentToken). Rides the draft and lands on the settlement entry on approval — no re-upload. A bad type/oversize file is rejected and NOTHING is filed. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | Optional original filename for the attachment, e.g. bank-in-slip.jpg. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'image/jpeg' or 'application/pdf'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false hints, so the description carries the full burden — and it delivers: it discloses that no posting or settlement happens until human approval, that the server re-derives allocations against live outstanding and its figures are authoritative, and that truncated base64 or bad-type files are rejected rather than filed. It also discloses the network-dependent behavior (inline anyway when taokeh.my is unreachable) and the 'never both' mutual exclusion. Nothing contradicts the annotations; creating a pending draft is consistent with readOnlyHint=false and destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (~350 words) but nearly every sentence maps to a distinct operational decision an agent must make, and it front-loads the core draft-not-post semantic before each concern. Minor tightening is possible — the 's.82 record-keeping' parenthetical and some attachment-outcome phrasing are repeated in the schema — but the density is warranted for a 15-parameter tool with two attachment paths and a fallback rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a very complex tool with uninformative annotations and no output schema, the description covers the full call workflow: prerequisites, currency exclusion, allocation modes, review obligations, both attachment routes, and the downstream record-keeping outcome. The remaining gaps are the unstated return value (e.g., whether a draft ID comes back for use with revise_draft) and the two undocumented parameters noted above — minor against the otherwise thorough coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 47% schema coverage, the description compensates for the decision-critical parameters: the EITHER lump `total` (auto-allocated oldest-first) OR explicit `allocations` mutual exclusion, the paired `needsReview`+`notes` workflow with quality guidance, and the attachmentBase64/attachmentToken either/or with `attachmentBytes` as a truncation guard. However, `bankAccountId` and `printedTotal` receive no meaning from either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb phrase — 'File a customer-payment DRAFT into Taokeh ("customer paid X")' — that names the exact resource and the direction of money movement, cleanly separating it from create_bill_draft and create_invoice_draft. The explicit 'This does NOT post' adds the crucial draft-vs-post distinction that the title alone cannot convey.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions and sequencing: resolve the customer via resolve_customer and check open_invoices first. States hard exclusions (MYR only; foreign-currency invoices are refused and route to Banking → Payments) and names the alternative tool for large attachments (request_attachment_upload), plus a concrete network-fallback rule for sandboxed shells.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_recurring_invoice_draftFile a pending recurring-invoice draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING IS CREATED UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a NEW recurring invoice — a billing schedule that issues the same invoice to the same customer every month, quarter or year (a retainer, a subscription, a maintenance contract, rent). This does NOT create anything: it files a pending DRAFT the owner reviews and approves; only that tap saves the schedule. Give the customer (customerId from resolve_customer, or customer as an EXACT existing customer name — this tool never creates a customer), the lines (each a real catalogue product by productId or exact sku, with the quantity and the agreed unit price that will be billed EVERY cycle), the cadence, and the startDate — whose day of the month becomes the billing anchor (a 31st anchor lands on the 28th/29th/30th in shorter months). Optionally endDate, maxOccurrences, title, notes and paymentMethod (CREDIT = the customer owes it and pays later, CASH = settled at issue; these book differently, so ask rather than guess). ⛔ WHAT IT CANNOT DO: it cannot set issueMode: "auto" — a schedule you propose files an invoice DRAFT each cycle for a person to approve, and nothing proposed through this connector may buy itself the right to post invoices unattended; the owner switches that on themselves if they want it. It cannot set MFRS 15 revenue spreading (recognition), which is an accounting-policy decision on the schedule's own page. It cannot set the next issue date directly — that is derived from the start date and the cadence, and the review page tells the owner the exact date the FIRST invoice would be issued before they approve. BE HONEST: never guess a contracted price or a start date — leave the schedule unproposed and ask.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | What gets billed every cycle. Every line must be a real catalogue product. ⛔ LINE KEYS ARE STRICT: a key this schema does not list is REFUSED BY NAME and NOTHING is filed. | |
| notes | No | A SHORT reviewer note, in the reviewer's language: what this schedule bills and what they should double-check. | |
| title | No | A short label for the schedule, e.g. 'Monthly retainer — Acme'. Optional; it never prints on the invoice. | |
| cadence | Yes | ||
| endDate | No | Stop issuing after this date. Optional. | |
| customer | No | An EXACT existing customer name. A name that matches nobody is refused — this tool never creates a customer. | |
| startDate | Yes | The first billing date. Its DAY OF THE MONTH becomes the anchor for every later cycle. A date in the past does not backfill — the first invoice lands on the first cycle on or after today. | |
| customerId | No | The customer id from resolve_customer. Give this OR `customer`. | |
| needsReview | No | Set true when something gave you pause — a price the user was unsure about, a start date you inferred. | |
| paymentMethod | No | CREDIT (default) = the customer owes it. CASH = settled at issue. Ask the user; these book differently. | |
| maxOccurrences | No | Stop after this many invoices. Optional. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Extremely transparent. States this files a draft only, that nothing is posted until a human reviews, that it never buys itself auto-issue, that a customer must already existarked, that no customer will be created. It also flags that wrong guesses should be avoided and user should be asked. That's exactly what an agent needs to know about side effects and prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but dense. The leading ALL-CAPS warning is an excellent front-loaded structure that prevents misuse. A few redundancies (e.g. repeating that nothing is created until human approval) could be trimmed, but everything present carries operational weight. Not quite a 5 because it could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers what the tool does, prerequisites (existing customer, real products), constraints (cannot auto-issue, no revenue spreading), field semantics, and behavioral guardrails. Sibling context is implied through repeated contrast with single-invoice drafts. Nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While the schema already describes each parameter, the description adds critical context: `customer` must be an existing resolved customer, lines may be productId or sku, prices bill every cycle, day-of-month anchoring for recurring dates, paymentMethod semantics (CREDIT vs CASH), and behavior guards (never guess). This is far beyond the schema's field-level descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise action (file a recurring invoice draft), a specific resource (recurring schedule), and its key distinguishing constraint — nothing is created until a human approves in Taokeh. This is clearly differentiated from create_invoice_draft and other draft tools by the 'recurring' focus and the explicit workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use it (propose a recurring invoice schedule) and when not to — it explicitly states it cannot set issueMode to auto or set revenue recognition, and warns against guessing prices/dates. It also says to ask rather than guess for paymentMethod. It doesn't explicitly name sibling tools like create_invoice_draft as alternatives, hence 4 not 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
daily_briefDaily briefAInspect
Your morning brief in one call: today's and this-month's sales & expenses, cash position (cash.total is LIQUID money only — cash + banks + gateway float; any owner advance comes back separately as cash.ownerFunding, and any credit-card debt as cash.cards / cash.cardsOwing — both are liabilities that must never be added to, or netted against, the cash total), who owes you (top debtors + overdue buckets — the buckets count invoices only; receivables.unmatchedPayments, when present, is customer money received but not yet matched to an invoice, already netted inside receivables.total), what's waiting for your approval (approvals.total counts DECISIONS, not rows: a real approval contributes its item count, while an operational to-do like "listings to link" counts as 1 however long its backlog — read each item's own count for the size. A STAGED SET (staged_documents, staged_products, …) is one decision too, so its count is always 1: read rows for how many are actually waiting, flagged for how many need a human look, and fenced for how many cannot be entered while the accounting start date stands where it is (they are dated on or before it, so the opening balances already carry them). A fenced row is not work the owner can finish on this queue: the remedies are to leave it out, or to move the accounting start date — after which Taokeh re-checks the fence and the row becomes enterable), tax position, and low stock. Assembles the same figures as the individual tools (business_snapshot, cash_position, ar_aging, tax_position, low_stock) plus your pending-approval queue, so "brief me" is a single read.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description repeatedly calls this a 'single read' and presents it as a read-only aggregation, but annotations set readOnlyHint=false, contradicting that safety profile. Although it adds rich domain semantics about cash, approvals, and staged rows, the annotation contradiction mandates a score of 1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It front-loads the summary, but the rest is a single dense, run-on paragraph with numerous parenthetical caveats. The field-level semantics are valuable for a no-output-schema tool, though the structure is not scannable and the prose is overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a complex aggregate return, the description explains the key output fields and their boundaries in detail: cash.total is liquid only, owner funding and cards are separate liabilities, receivables.unmatchedPayments are already netted, approvals count decisions rather than rows, and staged/fenced rows have special semantics. This is enough for correct interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there are no parameter semantics to clarify; baseline 4 applies. The description correctly does not attempt to add parameter detail beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource set: it is the one-call morning brief covering sales, expenses, cash, receivables, approvals, tax, and stock. It distinguishes itself from siblings by naming the individual tools it assembles (business_snapshot, cash_position, ar_aging, tax_position, low_stock).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly positions the tool for 'brief me' as a single read that replaces aggregating the individual figures, and names those alternatives. It does not explicitly state when not to use it, such as when only one metric is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_bank_classificationSuggest bank-row categories (human reviews in Taokeh)AInspect
Propose a category (and optionally a contact) for up to 50 bank rows still waiting for review — imported statement lines AND rows the owner typed in by hand, which land as UNCATEGORIZED — get their ids from bank_review_queue. PROPOSALS ONLY: this writes a suggestion the owner sees as an 'AI suggestion' badge on the Banking → Review screen; it never confirms, never posts, and nothing you send here moves money — the owner reviews and posts every line in /banking. Only SUGGESTED/UNCATEGORIZED rows accept a proposal (already-decided CONFIRMED/POSTED/IGNORED rows are rejected, and so is a statement line the owner paired with a book entry on Clear transactions — it is already in the books). A row that looks like a book entry still on no statement is accepted but comes back with possibleTwin (kind 'bookEntry') and a warning: propose a clearing (create_bank_clearing_draft) for it instead. A credit-card line that looks like the card bill already paid into the card comes back with possibleTwin (kind 'cardBill') and a warning: the owner should exclude it — there is no clearing for a card. A card payment whose card-statement line is already posted comes back with kind 'cardLinePosted': posting both would count the bill twice. category must be one of the review screen's own set. contactName applies only to CUSTOMER_PAYMENT (a customer) or SUPPLIER_PAYMENT (a vendor) and must match an EXISTING contact exactly — an unresolvable name rejects that row (create the contact first, or omit it). Send contactName whenever you know the payer: a CUSTOMER_PAYMENT or SUPPLIER_PAYMENT proposed without one is accepted, but Taokeh will not POST it until the owner picks the customer or supplier, because a payment with nobody on it would sit in receivables or payables under nobody. accountCode proposes the LEDGER ACCOUNT for the row (e.g. a hosting bill to 6500) and is accepted ONLY for the three categories whose account the owner picks by hand — EXPENSE, OTHER and INTERNAL_TRANSFER. EXPENSE is a plain expense paid by card or bank: money OUT only, no contact (it is not a supplier payment and never touches payables), and its accountCode must be one of the expense accounts expense_accounts lists. OTHER and INTERNAL_TRANSFER never take a receivables/payables control, a contra or clearing account (1016 Contra Clearing, Accumulated Depreciation, the gateway clearing account), Inventory, Opening Balance Equity or Retained Earnings — the owner's picker does not offer them either; OTHER never takes a bank, card or cash account (that is an INTERNAL_TRANSFER), and INTERNAL_TRANSFER takes ONLY one of them (a sales, expense or equity account is OTHER, or EXPENSE for money spent). INTERNAL_TRANSFER never takes a credit-card account for money going out of a bank line — paying a card off is CARD_PAYMENT with that card's code — and on a card's own statement, money in from a bank account is the card bill being paid: it is never posted from the card side, so the owner excludes that line (the bank line that paid it posts as CARD_PAYMENT). A balance transfer between two cards is posted from the CHARGE line (money out, on the card that paid); the paid-off card's money-in line is never posted either, so the owner excludes it. Every other category already books against a fixed account (a customer payment to receivables, a bank charge to bank charges, and so on), so sending accountCode with one is rejected and the reply names the account that category already carries. The code must exist in this company's chart (expense_accounts lists the expense side); the line's own bank or card account is refused, because that would post the money against itself (under OTHER every bank, card and cash account is refused — that move is an INTERNAL_TRANSFER). CARD_PAYMENT is the one category that REQUIRES accountCode, and takes only a credit-card account: it means "this bill settles that card", so propose it with the card's code or not at all — a card payment is never an expense, and naming an expense account (or naming a card account under OTHER) is refused, because the card's spending was already booked line by line from the card statement. The account you propose arrives PRE-SELECTED in the owner's account dropdown on Banking → Review, marked as proposed by you — it saves them hunting the chart, it does not decide anything: if they have already picked an account themselves, theirs stands. allocations proposes WHICH invoices or bills the money settles — the question Taokeh's own matcher hands to a human whenever the payment is partial, spans several documents, or could clear more than one combination, and the one you may already know from a remittance advice or the payer's own message. Send it ONLY with CUSTOMER_PAYMENT on a money-IN row or SUPPLIER_PAYMENT on a money-OUT row, and ONLY together with contactName, because every line is checked against THAT contact's live open documents: up to 20 entries of { docNumber, amount }, where docNumber is the document number as Taokeh shows it (open_invoices / ap_aging list them). EVERY amount is in RINGGIT, and specifically the outstanding balance as Taokeh reports it on those tools — never the document's own foreign-currency figure, and never its gross total where the two differ; a foreign-currency invoice is settled in Banking, not here. The amounts must add up to the WHOLE line within 0.10 — a proposal that leaves part of the deposit unexplained is refused with the exact shortfall, never trimmed to fit — no line may exceed what its document still owes, and an unknown number, a duplicate, or a document dated AFTER the payment is refused too; the refusal names that contact's own open documents with their outstanding amounts, so you can correct it in one more turn. What it does is PRE-FILL: the owner's Match pane opens with those documents ticked and those amounts entered, under a banner saying you proposed it and quoting your note. It settles nothing — the owner reads it and clicks Save — and if a document has been paid or part-paid since you proposed it, the pane says your proposal no longer fits and fills in nothing rather than allocating a stale figure. It never reaches the one-tap Accept on the payments list either; that button only ever applies Taokeh's own exact match. note is your stated reason (≤300 chars), shown to the owner — write it well, because the owner reads it when deciding. Your proposals are also collected into a 'Proposals from your AI' panel on Banking → Review, where the owner can accept a whole batch of them in one tap after reading them grouped by category with your reasons; that acceptance sets the category and contact only and is still the owner's decision — it never posts to the ledger. If a row's existing suggestion says "You set this rule on ", the OWNER has written a standing rule for that payer (What Taokeh has learned → Bank rules) — proposing something different is allowed and is sometimes right, but say in your note that you are contradicting their own rule, and expect them to keep the rule. Per-row outcomes are reported — nothing is silently skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| suggestions | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare non-readOnly/non-destructive/non-idempotent; the description goes far beyond, disclosing that these are proposals only, never posted, never move money, that accountCode arrives pre-selected but doesn't decide, that allocations pre-fill the Match pane and settle nothing, and that batch acceptance sets category/contact only. This is exactly the behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the PROPOSALS ONLY safety statement, which is good, but the body is an extremely dense run-on wall of prose mixing categories, accounts, twins, allocations and rules with no headings or list structure. Much of it earns its place given the tool's complexity, but the layout makes it hard to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, nested, high-stakes proposal tool with no output schema, the description covers rejection cases, per-row outcome reporting, possibleTwin handling, contact/account/allocation rules, and owner-facing side effects. Nothing an agent needs to call it correctly appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description fully compensates: category constraints, contactName matching rules and which categories accept it, accountCode allowed categories and refusals, allocations format ({docNumber, amount}, up to 20, RINGGIT, must sum within 0.10 of the whole line), and note (≤300 chars, shown to owner). Far exceeds the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Propose a category (and optionally a contact) for up to 50 bank rows still waiting for review.' It names the exact row states handled (SUGGESTED/UNCATEGORIZED) and distinguishes itself from siblings by pointing to bank_review_queue for ids and create_bank_clearing_draft for the twin case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when/when-not guidance throughout: which rows accept a proposal, which are rejected (CONFIRMED/POSTED/IGNORED, Clear-transaction pairs), when to route to create_bank_clearing_draft instead, when the owner should exclude a card line, and the rule that CARD_PAYMENT requires accountCode while other categories reject it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
employer_cost_estimateEmployer cost estimateARead-onlyInspect
Cost a HYPOTHETICAL hire: give a monthly wage and get back what that employee would really cost the employer each month — the wage plus employer EPF, SOCSO, EIS and the HRD Corp levy at this company's own rate — and what the employee would actually take home after employee EPF, SOCSO, EIS, SKBBK and PCB/MTD. Use it for "if I hire someone at RM4,000, what does it really cost me?" and "what would they take home?". Optional: residency (RESIDENT, the default, or NON_RESIDENT — a non-resident is taxed at a flat 30%), citizenship (MALAYSIAN or FOREIGNER — this drives the HRD Corp levy and makes SKBBK mandatory; it does NOT change the EPF, SOCSO or EIS rates applied, and it does not change tax residency, so set residency too), ageBand (UNDER_60 or 60_AND_OVER — reduced EPF and SOCSO Category 2), and epfEmployeeRate if the employee elects a reduced EPF rate. The PCB comes from Taokeh's implementation of LHDN's official computerised MTD method — the calculation IRBM confirmed in writing on 13 August 2026 (letter ref 2026-256) — not from a simplified formula. The defaults are RESIDENT, single, no children and NO TP1 reliefs, because reliefs are per-employee paperwork nobody has filled in for a person who does not exist yet; every assumption is spelled out in the reply and MUST be repeated to the user. Present it as an estimate on stated assumptions, never as a quote or as tax advice. READ-ONLY — it creates nothing, hires nobody and files nothing. ADMIN AND BOOKKEEPER + PAYROLL ONLY, like every payroll tool here.
| Name | Required | Description | Default |
|---|---|---|---|
| ageBand | No | ||
| residency | No | ||
| citizenship | No | ||
| monthlyWage | Yes | ||
| epfEmployeeRate | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds substantial behavioral context: it is READ-ONLY and 'creates nothing, hires nobody and files nothing.' It discloses the calculation source (LHDN's official MTD method, with a written confirmation reference), spells out all default assumptions, mandates repeating assumptions to the user, and positions the result as an estimate rather than a quote or tax advice. This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: use case, parameter semantics, tax computation provenance, default assumptions, compliance guardrails, and access restrictions. It is front-loaded with the core purpose and then moves systematically through parameter details. There is no filler or tautology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description clearly states what the tool returns: employer cost and employee take-home. It fully covers the five parameters, the applicable tax method, the default assumptions, and the mandatory communication requirement to repeat assumptions to the user. This is more than enough for an agent to select, invoke, and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it delivers. It explains monthlyWage through the RM4,000 example, residency via RESIDENT/NON_RESIDENT and the flat 30% tax rate, citizenship via HRD Corp levy and SKBBK effects, ageBand via reduced EPF and SOCSO Category 2, and epfEmployeeRate via the employee's elected reduced EPF rate. Every parameter receives meaning beyond its raw schema definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Cost a HYPOTHETICAL hire' and concretely defines the outputs: employer cost (wage plus EPF, SOCSO, EIS, HRD Corp levy) and employee take-home pay. It gives a representative question ('if I hire someone at RM4,000...') that makes the tool's scope unmistakable and distinguishes it from generic payroll or reporting siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use framing with example questions and explains the default assumptions users should expect. It also provides role restrictions ('ADMIN AND BOOKKEEPER + PAYROLL ONLY'). However, it does not name any alternative sibling tool or explicitly state when NOT to use this tool, so the guidance stops short of a full when-versus-alternatives treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_accountsExpense accountsARead-onlyInspect
The expense accounts you can post to in this company — pick a category (what the money was for) and a paid-from account (where it came from) by code, for shaping a paid-expense draft. expenseCategories is a CURATED list: expense and asset accounts only, with cash/bank, the A/R and Inventory controls and contra accounts left out, because a cost never lands in one of those — and without the accounts only Taokeh's own engines book to (depreciation, foreign-exchange losses, loss on disposal, cash rounding, and salaries/employer statutory once the company runs payroll in Taokeh). It is the everyday set, not the whole rule. When the draft is filed, a CATEGORY code is rejected only if it is not an expense or asset account, or is a cash/bank, control or contra account; an engine account left off the list (e.g. depreciation) is still accepted by code, so use one only when the receipt really is that cost. The paid-from code has its own rule: any asset, liability or equity account except the A/R and A/P control accounts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description goes far beyond that, explaining that expenseCategories is a CURATED subset, listing exactly what is excluded (cash/bank, A/R and Inventory controls, contra accounts, engine-only accounts), and detailing the distinct validation rule for category vs. paid-from codes. This is rich operational context an agent needs to avoid picking a rejected code.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is substantive, but it is delivered as one dense run-on paragraph with nested asides that make the core point hard to extract. It front-loads the resource name but buries the practical routing guidance mid-sentence, hurting scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the semantic burden and mostly does: it explains what the list contains, what it excludes, and the filing-time validation for each code type. It stops short of describing the return shape (e.g. code-plus-name entries) or ordering, which would fully close the gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document; baseline for a param-less schema is 4. The description instead explains the meaning of the values the tool returns, which is a reasonable substitute.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource — the expense accounts postable in this company, split into category and paid-from codes. It even ties the purpose to a workflow ("for shaping a paid-expense draft"), which is clearer than a bare name restatement. It does not explicitly distinguish itself from sibling lookups like search_expenses or find_in_taokeh, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: use it to pick a category and paid-from code when building a paid-expense draft, plus the acceptance rule at filing time for each code type. It effectively tells the agent what codes are valid where, but never names a sibling alternative or states when NOT to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expenses_summaryExpenses summaryARead-onlyInspect
How much you have spent: your paid expenses today and this month so far.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, and the description adds useful scope information: only paid expenses are counted, and the time window is today and month-to-date. However, it does not describe the return format, whether the result is a single total or separate numbers, or behavior when there are no expenses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the core question 'How much you have spent', and contains no filler or redundant wording. Every word contributes to meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only summary tool, the description provides sufficient detail to understand the tool's purpose and invoke it correctly. The lack of an output schema means the description could ideally clarify the return shape (e.g., two totals vs. one), but this is a minor gap for such a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema description coverage is 100%, so there is no parameter burden for the description to carry. The baseline of 4 applies because the schema fully covers an empty parameter set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a summary of paid expenses, with a specific scope (today and this month so far). It uses a clear 'have spent' action on the 'expenses' resource and is distinguishable from siblings like create_expense_draft or search_expenses, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking how much has been spent today and this month so far, but it does not state when to use this tool versus alternatives, nor does it provide any exclusion criteria. The context is clear enough but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
financial_timeseriesFinancial time series (monthly)ARead-onlyInspect
A MONTH-BY-MONTH series of this company's figures across a window you name — the one call to make when you are asked to analyse a year or three years of trading. NEVER call income_statement 36 times and stitch the months together yourself: ask here once. You get, per month: revenue, cogs, grossProfit, opex, netIncome, cashClose, arClose, apClose and netCashMovement, in MYR. Give from and to as YYYY-MM (inclusive). The window is capped at 36 months — ask for a narrower window, or several windows, rather than expecting more. metrics narrows what is computed: any subset of pnl, cash, ar, ap, cashflow (default all); a metric you leave out has its fields OMITTED from every month, never returned as zero. THREE THINGS YOU MUST READ BEFORE YOU INTERPRET IT. First, a month that has not finished yet comes back with partial: true and a through date: its figures cover only the days up to and including that date — TODAY — and not the rest of the month, so they are month-to-date. Say so, and do not let the last point bend a trend. Second, months that ended BEFORE this company's accounting start date are left OUT OF THE SERIES, not zeroed: that history is still in the old system, and a note names the cutover date — this is a rule about the series, not about the ledger, because back-dating is allowed here, so a report run over an earlier window on /reports may still show stray entries before that date — and the month CONTAINING that date is partial: true with startsAt, because its trading only begins there; it is incomplete on the left the way a running month is on the right, never a weak month. Third, months that have not started yet are ABSENT too. An absent month is missing or unlived history, never a month with no business — never read one as a collapse in trading. EQUAL BASIS: every figure is produced by the same engine the web report uses for the exact dates stated (the month, or the month up to through when it is partial) — the P&L from /reports, cashClose from the cash-position liquid set (cash, banks, gateway clearing; owner advances and credit-card debt are not cash and are not in it), AR/AP from the aging reports' totals as of month end, netCashMovement from the cash-flow report's net. One caveat to carry: netCashMovement and the month-on-month change in cashClose are measured on DIFFERENT account sets, so they can legitimately differ — neither is a check on the other. ONE CALL ALSO HAS A TIME BUDGET: a very wide window with every metric can exceed it, in which case the call is REFUSED (it says how many months it managed) rather than answered with a short series — ask again for fewer months or fewer metrics. This tool returns FIGURES ONLY. It computes no ratios, no growth rates, no trend labels and no commentary — the reading is yours to write, from these numbers and the notes.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes | ||
| metrics | No | Which families of figures to compute. OMIT for all five; do not send an empty list. A family you leave out is omitted from every month, not zeroed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is exceptionally transparent: it explains partial months with `partial: true` and `through`, absent months and the cutover rule, the equal-basis engine behind the figures, the netCashMovement vs cashClose discrepancy, the time-budget refusal behavior, and the fact that it returns figures only with no interpretation. The readOnlyHint annotation is consistent with this being a read-only computation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the length is justified by the tool's many edge cases. It is well-organized into labeled sections (THREE THINGS YOU MUST READ, EQUAL BASIS, TIME BUDGET, FIGURES ONLY), which helps an agent parse the critical caveats. A few phrases could be tightened, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description names every field returned, explains how each is computed, clarifies partial/absent/cutover behavior, states currency, and documents failure modes. An agent has everything needed to interpret the response correctly and avoid common misinterpretations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description carries the burden for `from`/`to` and it does: YYYY-MM format, inclusive, 36-month cap, and MYR currency. For `metrics`, it adds that leaving a metric out omits those fields entirely and never zeroes them, going beyond the schema's basic enum description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, specific statement: a month-by-month series of company figures across a user-named window. It also differentiates itself from income_statement by explicitly saying to request this once instead of stitching 36 months together, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives direct when-to-use guidance: use this for analyzing a year or three years of trading, and never call income_statement 36 times. It also advises asking for narrower windows or fewer metrics when the time budget is exceeded, covering the main alternative and constraint behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_in_taokehFind in TaokehARead-onlyInspect
A FEATURE FINDER, NOT A RECORD SEARCH. It searches Taokeh's own SCREENS, FEATURES and SETTINGS — never the books. Use it to tell the user WHERE in Taokeh to do something or change a setting: it returns the exact page(s) and what each one does, grounded in Taokeh's LIVE feature catalogue, so you never guess or invent navigation. ⛔ It does NOT look up documents, invoices, customers, products or any other RECORD, and it cannot tell you it failed to — a document number, a customer's DO/PO number, a person's name or an amount typed in here will still rank some page (the closest-worded screen), and that page will be a wrong answer wearing a right one's clothes. Records live elsewhere: search_documents (invoices, credit notes, quotes, delivery orders, bills, POs — and a customer's own Ref No), search_expenses (posted paid expenses), search_journal (ledger entries), resolve_customer / resolve_vendor / resolve_product (a named party or item), open_invoices (what one customer owes). Give a natural-language query about a CAPABILITY ("where do I turn on payment reminders", "how do I connect Shopee", "where's the SST setting") and it ranks the real features and hands back the top matches, each with the page's absolute deep-link URL(s), a one-line description of what the page does, an optional configHint (Taokeh keeps a feature's options on its OWN page, not in a global Settings menu), and an addOn flag (a paid add-on this company may not have). It GUIDES to the UI only — it changes nothing: no setting is ever toggled by the connector, that stays a human action in-app. If nothing matches it says so honestly — then tell the user to browse the left sidebar or contact support; never make up a path.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | What the user wants to DO or the setting they want to CHANGE, in words — not a document number, a customer name or an amount (those are records; use search_documents / resolve_customer instead). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true annotation, the description discloses the tool's failure mode (it will rank some page even for record-like input and return a wrong answer), its output shape (deep-link URLs, one-line page description, configHint, addOn flag), and guarantees that no setting is ever toggled. This is exactly the contextual behavior an agent needs to trust results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core distinction and each sentence carries substantive guidance, but it is long and repeats the 'not a record' warning more than once. The density and warning are justified by the tool's deceptive failure mode, though it could be tightened slightly without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description fully compensates by enumerating the return elements (page URLs, descriptions, configHint, addOn flag), covers edge behavior (nothing matches), and provides alternative routing. Nothing an agent needs to use it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the query parameter at 100% coverage, giving the baseline 3. The description adds value by clarifying it must be a natural-language CAPABILITY query, providing concrete examples, and reinforcing what it must NOT be (document/customer/amount), so the parameter semantics are further strengthened.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb phrase ('A FEATURE FINDER, NOT A RECORD SEARCH') and names its exact scope: Taokeh's screens, features, and settings. It explicitly distinguishes itself from record-search siblings by stating what it is not, so an agent can immediately tell find_in_taokeh apart from search_documents, search_expenses, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance ('tell the user WHERE in Taokeh to do something or change a setting') and equally explicit when-not-to-use with a detailed list of alternative tools for records, including what each alternative covers. It even instructs the fallback behavior when nothing matches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_attachmentGet original document linkARead-onlyInspect
Get a short-lived link to the ORIGINAL DOCUMENT filed against something in the books — the receipt behind a posted expense, or the receipt/sales note/supplier bill/bank-in slip/name card riding a PENDING draft. Use it to CHECK a document is on file, or to re-read one you filed earlier. Find what to ask for first: search_expenses reports hasAttachment and an attachments list (with an attachmentId) on every posted expense it returns, and each create_*_draft / revise_draft result reports whether an original rides that draft. Give owner plus the id of the thing that owns it — for a posted expense that is the expenseId from search_expenses, plus the attachmentId when the entry carries more than one file; for a draft it is the draftId. Returns a signed URL (about 15 minutes, works for anyone holding it — so treat it as you would the document itself), the filename, the media type and the size in bytes; the bytes themselves are NEVER inlined here, because a base64 blob in a tool result costs about one token per character. Download it and read it yourself: Taokeh does not read, OCR or interpret the file for you — your own AI does that, on your own subscription. Read-only; it changes nothing, emails nobody, and no link it hands out can reach another company. An APPROVED draft honestly reports no original: on approval the file moves onto the posted document, so ask for it there instead.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The owner's id — the expenseId from search_expenses for 'expense', otherwise the draftId a create_*_draft or revise_draft call returned. | |
| owner | Yes | What the original hangs off: 'expense' = a POSTED paid expense (use the expenseId from search_expenses); 'expense_draft' / 'invoice_draft' / 'bill_draft' / 'quote_draft' / 'receipt_draft' / 'payment_draft' / 'contact_draft' = a PENDING draft (use its draftId); 'shoebox' = a photo someone in the business sent in from their phone and nobody has read yet (use the id from shoebox_items). | |
| attachmentId | No | Which file, when a posted expense carries more than one — the `attachmentId` from that expense's `attachments` list in search_expenses. Omit for a draft (a draft carries at most one original), and omit for an expense with exactly one file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses that the returned link is short-lived (about 15 minutes) and that anyone holding it can access the document, adding a security caution beyond the readOnlyHint annotation. It also identifies the return items (signed URL and filename). No contradiction with the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is verbose and ends with a nonsensical repeated phrase 'the filename, the filename, ...' which wastes tokens and confuses the return-value statement. While the main content is ordered logically, the malformed tail and excessive examples hurt conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, prerequisites, parameter sourcing, and a security caveat, which is fairly complete for a read-only retrieval tool. However, the corrupted tail leaves the exact return payload unclear (only 'the filename' is restated), and error behavior (e.g., no original on file) is not addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes all three parameters at 100% coverage, and the description adds cross-tool guidance on where to find the correct ids (expenseId from search_expenses, draftId from drafts, attachmentId from attachments list). This enriches the schema beyond generic field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence explicitly states the tool's action: 'Get a short-lived link to the ORIGINAL DOCUMENT filed against something in the books.' It also distinguishes between posted expenses and pending drafts, making the target resource unambiguous. Even with the malformed tail, the central purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells the agent when to invoke it ('to CHECK a document is on file, or to re-read one you filed earlier') and how to prepare inputs using search_expenses and create_*_draft results. It lacks an explicit 'when not to use' clause or named alternatives, but the usage context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_document_pdfGet document PDF linkARead-onlyInspect
Get a shareable link to the ACTUAL PDF of one invoice, quote or credit note — so you can hand the user (or their customer) the real document, e.g. "here's your invoice". Identify the document by its docId (from search_documents) OR its printed docNumber (e.g. the invoice number) — give one; if a number matches several documents I'll list them so you can pick the id. Returns a short-lived signed URL (works for ~15 minutes, for anyone who has it — so share it deliberately), the document number, and when it expires. Covers invoices, quotes and credit notes (they share one document layout). A CANCELLED document still has a PDF, stamped: a voided quote or a voided CREDIT NOTE comes back with a large diagonal VOID across every page, so it can be handed over as the record of what was cancelled but never passed off as live — say so when you share it, and never quote its amount as money owed or credited. On a voided credit or debit note the PDF also states it in WORDS, not only the watermark (2026-09-14): its header box reads Status VOID with no payment term or due date, the total is labelled what it WAS ('Total credited' / 'Total charged') in a muted bar rather than the usual Grand Total, and a line beneath reads 'Nothing is credited now — this credit note was voided on '. So the document says the thing you must say; quote THAT, not the figure. search_documents finds a voided credit note too and reports its status as "Credit note (VOID)"; read that before you describe the document. (A voided INVOICE is the exception — voiding one removes it from the books entirely, so there is no link left to give and search_documents will not find it either.) LANGUAGE: the PDF's printed words (Bill to, due date, column headings, totals) follow the company's Document language setting — the same setting covers the company's invoice, quote, delivery order, credit and debit note and purchase order PDFs — — English unless the owner chose Bahasa Malaysia on Settings → Preferences — so a voided note's sentence may read in Malay; names and descriptions always print as typed. It is not the app language, and you cannot change it per link. This does not email anyone — it hands YOU a link to pass on.
| Name | Required | Description | Default |
|---|---|---|---|
| docId | No | The document id (from search_documents). Give this or docNumber. | |
| docType | Yes | Which document: 'invoice', 'quote' or 'credit_note'. | |
| docNumber | No | The printed document number (e.g. the invoice/quote reference). Give this or docId; resolved case-insensitively and must match exactly one document. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds substantial behavior the agent could not otherwise know: the URL is signed, short-lived (~15 min), and shareable by anyone holding it. It discloses that voided documents still produce stamped PDFs (diagonal VOID, 'Status VOID' header, 'Total credited/charged' labels, an explicit sentence in the PDF), that voided invoices have no link at all, and that printed language follows a company setting. This far exceeds the annotation bar.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and no section is entirely off-topic, but the mass is disproportionate: the voided-document behavior is restated across several long clauses, there is a dangling dated parenthetical ('2026-09-14') and a doubled em-dash typo. The same facts could be conveyed in roughly a third of the words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the return payload (signed URL, document number, expiry) and covers the real edge cases: ambiguous numbers, voided invoices being unfindable, voided credit notes still resolving, and the language setting. Nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description matches and slightly extends the schema: docId comes 'from search_documents', docNumber is the printed number, and it adds the ambiguity behavior — 'if a number matches several documents I'll list them so you can pick the id' — which the schema does not state. DocType is implicit in the description's enumeration of invoice/quote/credit note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Get a shareable link to the ACTUAL PDF of one invoice, quote or credit note.' It distinguishes itself from search_documents (the identifier source) and states exactly what it does not do ('This does not email anyone — it hands YOU a link'). An agent can pick it apart from siblings immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use — hand the real document to a user or their customer — and two explicit identification paths (docId from search_documents or docNumber). It also routes the agent to search_documents to read a document's VOID status before describing it. It does not contrast against other link/attachment siblings (get_attachment, request_attachment_upload), so it stops short of full alternatives coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_bank_statementImport bank statement into the review queue (nothing posts)AInspect
Import a bank statement into Taokeh for reconciliation, from a statement you've read (PDF/CSV/image). This does NOT post to the books — every row lands in Banking → Review for the user to categorize and post line by line. A statement is the bank's own list of transactions for ONE account over a period: use intake_contract(doc_type:'statement') first to see the tenant's bank accounts + the arithmetic law. Give the account id, the printed opening + closing balances, the printed statement date (statementDate) and beginning-balance date (beginningBalanceDate) when the statement shows them, and every transaction row (signed amount: + money in, − money out). The server checks opening + the rows tie to the closing balance and rejects with the exact delta if they don't — never invent or omit rows to force it to balance. CREDIT CARDS: a card statement can be filed here too, but ONLY with signConvention declared, because a card prints the opposite convention (a purchase INCREASES what you owe). Pass 'PRINTED_CARD' when the figures are exactly as the card statement prints them (the usual case) or 'STORED' if you deliberately converted them to Taokeh's convention (− = a spend). Never guess: declaring the wrong one books the whole month backwards and the arithmetic check cannot catch it, because negating every figure still ties out — the server runs a separate directional check and refuses a declaration that reads as a card which never owed anything. signConvention is refused on a bank account. Each card line is then expensed on its OWN statement date against the card, and the "payment received" line that settles the card bill is excluded automatically (the bank statement's lump owns that movement). Single receipts, invoices or bills go through their own draft doors, not here. ARCHIVING THE ORIGINAL: this import archives the statement file WHEN you send one — call request_attachment_upload FIRST, PUT the statement's raw bytes to its uploadUrl, then pass the returned attachmentToken here; the original is then kept on the statement (s.82 record-keeping) and the user can download it from the statement page, exactly as the web upload at Banking → Import does. Without a token nothing is archived and the books carry rows with no source document behind them. Already filed one bare? Call this tool AGAIN with the IDENTICAL account, balances and rows plus attachmentToken — the identical reading matches the same statement, so the file is ADOPTED onto it and no rows are staged twice. Filed one before without statementDate / beginningBalanceDate? The same identical call with the two dates added records them on that statement, again staging nothing twice.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| bankAccountId | Yes | ||
| statementDate | No | The statement date printed on the statement (the ENDING BALANCE date), YYYY-MM-DD — it is the end of the period the statement covers, even when the last transaction is earlier. Omit it only if the statement prints none. | |
| closingBalance | Yes | ||
| confirmOverlap | No | Set to true ONLY after the user has confirmed this is a genuinely new, separate statement whose period overlaps one already imported. It bypasses the overlap guard that stops the same transactions being counted twice. Never set it to push an overlap through — ask the user first. | |
| openingBalance | Yes | ||
| signConvention | No | REQUIRED when bankAccountId is a CREDIT CARD, and refused on a bank account. Which convention you signed the rows AND the opening/closing balances in. 'PRINTED_CARD' = exactly as the card statement prints them: a purchase is POSITIVE (what you owe goes up), a payment or refund is NEGATIVE, and the balances are the printed balances owed — this is the right choice when you read the card's own statement. 'STORED' = already converted to Taokeh's stored convention: a spend is NEGATIVE, a payment or refund POSITIVE, and the balances negative while money is owed. Do not guess between them: the wrong declaration books the entire month backwards, and the arithmetic check cannot catch it because negating every figure still ties out. Ask the user if unsure. | |
| attachmentBytes | No | The decoded byte size of the ORIGINAL file on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing a corrupt file. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the statement bytes to its uploadUrl. This is the normal way to archive a statement PDF — it carries the file out-of-band (no base64 in this call). Mutually exclusive with attachmentBase64. The file is archived on the statement and the user downloads it from the statement page in Taokeh. | |
| attachmentBase64 | No | The ORIGINAL statement as base64 — SMALL files only (a few KB). A real bank statement PDF is never that small, so in practice use request_attachment_upload + attachmentToken instead. Mutually exclusive with attachmentToken. A bad type/oversize file is rejected and NOTHING is imported. | |
| attachmentSha256 | No | The SHA-256 of the ORIGINAL file as 64 hex chars — optional second check alongside attachmentBase64, so corrupted bytes are rejected instead of filed. | |
| attachmentFilename | No | The original filename, for the user browsing their statements. | |
| attachmentMediaType | No | The attachment's MIME type, e.g. 'application/pdf'. Required when attachmentBase64 is given. | |
| beginningBalanceDate | No | The date on the BEGINNING (opening) BALANCE line, YYYY-MM-DD — the start of the period. Every row must fall between it and statementDate. Omit it only if the statement prints none. | |
| confirmSignConvention | No | Set true ONLY after the user confirms that the card genuinely sat in CREDIT for the whole period. It bypasses the directional check that refuses a statement reading as a card which never owed anything — the usual cause of which is a printed statement declared 'STORED' (or the reverse). Never set it to push a refusal through: ask the user first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all neutral false values (readOnlyHint, idempotentHint, destructiveHint, openWorldHint), so the description carries the full disclosure burden and discharges it thoroughly: the non-posting review-queue landing, the server's balance check that 'rejects with the exact delta,' the card directional check, automatic exclusion of the payment-settlement line, and the archiving/adoption behavior of repeated calls. It even discloses consequences ('declaring the wrong one books the whole month backwards') and what happens without a token ('nothing is archived').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long, but front-loaded: the first sentence states the purpose and the second the non-posting guarantee, with subsequent paragraphs organized topically (core rules, credit cards, archiving, re-calls) and no filler. Some repetition with the schema's own parameter descriptions (e.g., signConvention) exists, but the density is justified by the tool's complexity and hazard profile.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, 4-required import tool with no output schema, the description covers prerequisites, sequencing, edge cases (cards, overlap), failure modes (balance delta, directional refusal), and idempotent re-call behavior. The only notable gap is that it never describes the response/return value, which matters more given the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 73%, above the high bar, so the baseline is 3; the description still adds meaning the schema lacks: the signed-amount convention (+ money in, − money out), the invariant that opening + rows must tie to closing, and the effect of signConvention on the directional check. It also explains that attachmentToken must come from request_attachment_upload and that identical re-calls adopt the file, enriching several parameters at once.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Import a bank statement into Taokeh for reconciliation' — and immediately scopes what it is not: 'This does NOT post to the books — every row lands in Banking → Review.' It differentiates from siblings by routing single receipts/invoices/bills to 'their own draft doors,' and the title's '(nothing posts)' qualifier is reinforced throughout.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequential prerequisites: 'use intake_contract(doc_type:'statement') first' to see the tenant's bank accounts and the arithmetic law, and 'call request_attachment_upload FIRST' before archiving. It names exclusions with alternatives ('Single receipts, invoices or bills go through their own draft doors, not here') and specifies exact re-call conditions where identical parameters adopt a file without double-staging.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
income_statementIncome statementARead-onlyInspect
Your profit and loss for a period: revenue, cost of goods sold, gross profit, expenses and net income. COMPARISON (optional): to answer "how does this month compare with last?" ask for it here — NEVER call this tool twice and subtract the figures yourself. Set compare to 'previous_period' (the window of EQUAL LENGTH immediately before this one) or 'same_period_last_year' (the same dates one year earlier), OR give an explicit earlier window with BOTH compareFrom and compareEnd. You get both periods' full statements plus, for every account line and every total, deltaCents and deltaPct computed server-side in integer cents. deltaPct is NULL whenever the earlier figure is zero — a percentage change from zero is undefined, so report it as "no comparable base", never as ∞% or 100%. Percentages are measured against the ABSOLUTE earlier figure, so a movement out of a loss keeps the sign of the movement. windows states both windows, each one's day count, whether they are the same length, and whether either is still open (running to today or beyond, so its figures can still move); notes says all of this in plain words. READ notes BEFORE quoting a variance: if equalLength is false you are looking at a partial period against a full one and must say so. PERIODIC STOCK: when the company is on periodic stock, sold goods that have a cost in the window, has cost of goods sold of RM 0 and has not counted its stock this financial year, the answer carries a notes line saying so (with or without a comparison) — its gross profit is then overstated, so never quote that margin as the real one; tell the owner a stock take (or switching to recording cost on every sale in Settings → Accounting) fixes it.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes | ||
| compare | No | Named comparison window. 'previous_period' = the equal-length window immediately before from…to; 'same_period_last_year' = the same dates one year earlier. Omit for no comparison. | |
| compareEnd | No | End of an explicit comparison window (use WITH compareFrom). | |
| compareFrom | No | Start of an explicit comparison window (use WITH compareEnd; overrides `compare`). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although annotations declare readOnlyHint=true and openWorldHint=false, the description adds rich behavioral context beyond safety: it details server-side delta computation in integer cents, the null handling of deltaPct when the earlier figure is zero, sign behavior when moving out of a loss, the presence of `windows` and `notes` fields, and periodic-stock edge-case warnings. These go well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core purpose then moves to comparison guidance and edge cases. It is dense and information-rich, with every sentence serving a purpose, but the long periodic-stock paragraph and nested warnings make it slightly sprawling for an at-a-glance read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description fully compensates by explaining the structure of returned fields (`deltaCents`, `deltaPct`, `windows`, `notes`) and how to interpret them. It also covers edge cases like zero-base percentages and periodic stock, making it complete for correct invocation and reporting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, so the description compensates by explaining the semantics of `compare` (with definitions of 'previous_period' and 'same_period_last_year') and how `compareFrom`/`compareEnd` form an explicit earlier window. It adds meaningful interpretation of comparison parameters, though `from`/`to` are not described beyond the schema patterns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Your profit and loss for a period' and then enumerates the exact line items returned (revenue, COGS, gross profit, expenses, net income). This clearly distinguishes it from siblings like balance_sheet or cash_forecast, which cover different financial statements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance for comparisons — 'to answer "how does this month compare with last?" ask for it here — NEVER call this tool twice and subtract the figures yourself.' It also prescribes how to set comparison parameters and warns to read `notes` before quoting a variance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
intake_contractIntake contractARead-onlyInspect
The shape a Taokeh document needs, so you can turn a paid receipt, a sales note, a supplier bill, a customer payment, a payment you made to a supplier, a sales return, a bank statement or a whole set of OPENING BALANCES into a correctly-formed submission: required fields, how to read totals / measurements / SST, and a worked example grounded in this company. Pass doc_type 'expense' (default), 'invoice', 'bill', 'purchase_order', 'statement', 'receipt', 'payment', 'credit_note', 'sales_debit_note', 'journal', or — when this company is SWITCHING from another accounting system — 'historical_document' (ONE already-issued invoice or supplier bill being brought across in bulk with stage_document), 'trial_balance', 'aged_receivables' or 'aged_payables', which are staged in one call each with stage_opening_balances, and 'accounts', 'products', 'contacts', 'settings' or 'employees' (its CHART OF ACCOUNTS, its item list, its customer/supplier book, its COMPANY SETUP and its STAFF LIST), which are staged in one call each with stage_master_data — do 'accounts' first, since the trial balance matches against the chart. The 'settings' contract lists every stageable key, its meaning and current value, plus every owner-only setting the AI may never write and the reason/page for each. The 'employees' contract is ADMIN-ONLY. Nothing posts on its own — the user reviews everything in Taokeh.
| Name | Required | Description | Default |
|---|---|---|---|
| doc_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Nothing posts on its own.' It adds extra behavioral context: admin-only for employees, owner-only settings the AI may never write, and the review flow. These go beyond the annotations, though not exhaustively.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded with the core purpose. It is dense with essential information and avoids filler. While a more structured format (e.g., bullets) could improve readability, every sentence earns its place given the tool's role as a contract reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that defines contracts across many document types, the description covers all necessary context: purpose, doc_type variants, staging relationships, precedence, and permission constraints. There is no output schema, but the tool's output (the contract) is sufficiently implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a bare string for doc_type with no enum, and schema coverage is 0%. The description fully compensates by listing every valid doc_type value and explaining its meaning and usage, making the parameter self-documenting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool's purpose: defining the shape a Taokeh document needs for various transaction types, and enumerates all doc_type values. It clearly distinguishes itself from siblings by positioning itself as a reference contract rather than a creation or staging tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: it explains when to use this tool (to understand the document shape) and when to use alternatives like stage_document, stage_opening_balances, and stage_master_data for specific doc types. It also prioritizes 'accounts' first and notes admin-only restrictions, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
low_stockLow stockARead-onlyInspect
Which products are at or below their reorder point right now — the "what do I need to reorder" list. Each row: sku, name, on-hand quantity, reorder point and unit. Capped, with the true count so you know if more exist. Services ("Service — no stock") never appear. Uses the same rule as the dashboard low-stock chip (reorder point set AND on-hand ≤ reorder point). If this company uses Counter (the till), quantities are as at the last day-close: counter sales move stock once, when the day is closed, so an open trading day shows more on hand than the shelf does.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, yet the description adds substantial context beyond them: results are capped with a true count, services ('Service — no stock') are excluded, the rule matches the dashboard chip, and Counter-use companies see day-close figures with a clear explanation of the open-trading-day skew. This is exactly the behavioral detail that prevents misreading the numbers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then the row shape, capping, exclusions, the rule, and the Counter caveat. Nearly every sentence earns its place, though the density is high and the Counter sentence is long enough to be a slight drag on scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the full burden and discharges it: it enumerates the returned columns (sku, name, on-hand, reorder point, unit), discloses capping plus true count, and explains the accounting-timing caveat. An agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero input parameters, so the baseline is 4; there is nothing to document and the description does not need to compensate. It correctly spends its words on output shape instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource outcome: the products at or below their reorder point, framed as 'the what do I need to reorder list'. It gives the exact inclusion rule, which is enough to distinguish it from stock_level and stock_movements without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The framing ('what do I need to reorder') makes the use case explicit and the disclosed rule (reorder point set AND on-hand ≤ reorder point) tells the agent exactly which products qualify. It does not explicitly name a sibling alternative or a when-not condition, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
my_workMy workAInspect
YOUR OWN QUEUE in this company, in one call — what you filed and what happened to it. Answers four questions: (1) PENDING — everything of yours still waiting on the owner's approval, across every lane (expense, invoice, bill, quote, receipt, supplier payment, contact, credit note, debit note, purchase order, adjusting journal, your bank-category proposals), each with its reference, party, amount, how many days it has waited, and the exact next step in words plus the owner's review link; (2) APPROVED — what the owner approved since since (default: the last 7 days) and WHAT IT BECAME, with the posted document's own reference, so you can say "that one is now invoice INV-0042"; (3) REJECTED — what the owner or the engine refused, with the reason VERBATIM whenever one is stored. Every reject door in Taokeh offers the owner an optional "Why?" box, and staged migration rows carry the engine's own refusal text; where a reason is present, reason holds those exact words and reasonRecorded is true — READ IT and fix precisely what it names rather than re-filing the same thing. The box is never compulsory, so a refusal with an empty box reports reason: null and reasonRecorded: false; read that as "unknown", ask the owner what was wrong, and do NOT guess why before re-filing; (4) STAGED — your migration work (historical documents, opening balances, master data) sitting at a commit door. Params: since (YYYY-MM-DD or ISO, optional), kind (optional single-lane filter), limit (default 50, max 200 — every bucket reports total, returned and an honest truncation note; nothing is silently dropped). READ-ONLY. It changes nothing, it cannot approve anything, no tool can approve on the owner's behalf, and no tool of any kind moves money — every pending item is one human tap in Taokeh.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No | ||
| since | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotation Contradiction: annotations set readOnlyHint=false while the description claims 'READ-ONLY. It changes nothing, it cannot approve anything, no tool can approve on the owner's behalf, and no tool of any kind moves money.' This is a direct contradiction with the structured annotation and undermines an agent's ability to trust side-effect expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured: it uses numbered questions, separates the four result buckets, and clearly labels parameter semantics. There is some redundancy in the repeated safety statements, but the length is largely justified by the tool's complexity and the lack of output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and sparse annotations, the description carries the full burden of explaining return behavior. It details what each bucket contains, the `reason`/`reasonRecorded` semantics, defaults, truncation behavior, and the read-only guarantee. An agent has enough information to call the tool and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description fully compensates: it defines `since` as YYYY-MM-DD or ISO with a 7-day default, `kind` as an optional single-lane filter, and `limit` as default 50/max 200 with honest truncation reporting. It also enumerates the lane categories, mapping directly to the schema's enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific, identifiable purpose: it returns the caller's own queue of submitted items across PENDING, APPROVED, REJECTED, and STAGED buckets, and names the lanes and fields included. This clearly distinguishes it from sibling reporting and draft-creation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly frames when to use the tool: to answer four questions about the user's own submitted items and their statuses. It also gives parameter-level guidance for `since`, `kind`, and `limit`. It does not explicitly name alternatives or state when not to use it, but the scope is clear enough for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_invoicesOpen invoicesARead-onlyInspect
The unpaid (or part-paid) invoices for one customer, oldest first — search by customer name. Each invoice's outstanding is what it still owes net of payments AND of the live credit notes raised against it (a refunded note no longer counts), which is the figure to quote and to propose. grossOutstanding is the invoice before those notes: a payment may be allocated up to it (a customer who paid the whole invoice), and anything above outstanding then stays on the credit note as the customer's refundable credit. An invoice its credit notes fully cover is not listed. outstandingBalance is the customer's NET balance — these invoices less any payment received but not yet matched to one — so it can be lower than the invoices add up to (see balanceNote).
| Name | Required | Description | Default |
|---|---|---|---|
| customer | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real behavioral context beyond that: results are ordered oldest-first, invoices fully covered by credit notes are excluded from the list, and outstandingBalance is a net customer figure that can be lower than the listed invoices sum. It does not mention pagination, result limits, or match semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and ordering are front-loaded in the first clause, and each subsequent sentence explains a distinct figure (outstanding, grossOutstanding, outstandingBalance). It is nonetheless a dense block of prose for a one-parameter read tool, and the payment-allocation aside in the middle is longer than needed to support correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining the returned figures and does so thoroughly, including the relationship between outstanding, grossOutstanding, and outstandingBalance, and the balanceNote caveat. Remaining gaps are match semantics and pagination/result limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter is undocumented in the schema (0% coverage). The description clarifies that it is a customer name to search by rather than an ID, which is genuinely useful, but it does not say whether matching is exact, partial, or fuzzy, nor what happens on ambiguity (resolve_customer exists as a sibling, suggesting ambiguity is possible).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource+scope: 'The unpaid (or part-paid) invoices for one customer, oldest first', which an agent can distinguish from balance-level siblings like ar_aging or ap_aging. It stops short of naming an alternative tool or stating what this is NOT, so sibling differentiation is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'search by customer name' implies the intended invocation, and 'for one customer' implies the scope versus a portfolio-level aging report. However, there is no explicit when-to-use versus when-not guidance and no alternative named (e.g. ar_aging for all customers), so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payer_historyPayer historyARead-onlyInspect
How this company has treated a given bank counterparty (a payer or payee, named as printed on the statement line) before: the historical bank rows a human already CONFIRMED or POSTED for that name — grouped by category, contact and GL account, with a count, date range and total amount each — plus the live suggestion Taokeh would now make for it. That live suggestion carries a basis: 'LEARNED' means it was worked out by tallying the rows among those that are still POSTED — a confirmed-but-unposted row, or a post the owner undid, teaches nothing — and timesSeen says how many; 'OVERRIDE' means the OWNER has written a standing rule for this payer on the memory page (What Taokeh has learned → Bank rules), which BEATS the tally — ruleSetOn is the date they set it, and timesSeen is 0 because a rule was never "seen" any number of times. Never report an override as history: the owner said so, the books did not. This is the tenant-scoped PRECEDENT behind a suggestion, so you can explain WHY a row is being categorized a certain way. READ-ONLY; the owner still confirms every categorization in Taokeh.
| Name | Required | Description | Default |
|---|---|---|---|
| counterparty | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=false, and the description goes well beyond them: it defines what 'teaches' the tally (a confirmed-but-unposted row or an undone post teaches nothing), explains that OVERRIDE beats the tally, that timesSeen is 0 for rules, and that ruleSetOn is the date set. This is exactly the kind of semantic disclosure annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and almost every sentence carries needed semantics (basis meanings, override precedence). It is dense and runs long for a one-parameter tool, with heavy emphasis capitalization (CONFIRMED, POSTED, LEARNED, OVERRIDE, BEATS) that borders on noise, keeping it off a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden itself, and it does: it enumerates the returned fields (count, date range, total amount, grouping keys) and the suggestion's basis/timesSeen/ruleSetOn. An agent has enough to call it and to interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single counterparty parameter, so the description must compensate. It does so partially by clarifying the value is 'named as printed on the statement line', which tells the agent to pass the raw statement text, but it gives no detail on matching, casing, or unmatched names. Adequate but leaves a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns the historical, human-confirmed bank rows for a named bank counterparty, grouped by category/contact/GL account, plus the live suggestion. This is clearly distinguishable from sibling resolvers (resolve_vendor, resolve_customer) and from draft_bank_classification, which draft rather than report precedent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear context for use ('the tenant-scoped PRECEDENT behind a suggestion, so you can explain WHY a row is being categorized a certain way') and a usage directive ('Never report an override as history'). It stops short of naming alternatives or stating when NOT to call it, so it does not reach a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payroll_summaryPayroll summaryARead-onlyInspect
A month's payroll from gross to net, as it was actually run: total gross (basic, allowances, overtime, public-holiday pay and any bonus), what was deducted from staff (EPF, SOCSO, EIS, SKBBK, PCB/MTD, zakat), what the employer contributed on top (EPF, SOCSO, EIS, HRD Corp levy), the total net pay, and the employer's TRUE total cost for the month — wages plus every employer contribution. Also lists who is on approved leave in the next 14 days, so "what does payroll look like this month" and "who is out next week" are one call. With no month it reads the latest FINALIZED or POSTED run; give a month as YYYY-MM for a specific one. Set perEmployee: true for the per-person breakdown — those are individual salaries, so ask for them only when the user actually wants them; each entry carries that person's real employeeId. Set staff: true for the full employee ROSTER with ids, which works even when payroll has never been run — that is where the employeeId for update_employee_draft comes from, and it is the only place an id is published: never key an employee by name. Every figure is a sum of the STORED payslips of that run — the same numbers the staff were paid on and the payroll journal was posted from — never a recalculation. If nothing has been run, it says so plainly rather than returning a month of zeros. ADMIN AND BOOKKEEPER + PAYROLL ONLY: payroll sits behind its own role in Taokeh and this tool refuses on any other connection (a plain Bookkeeper included). READ-ONLY — it cannot create a run, pay anybody, or file anything with LHDN, KWSP or PERKESO.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | ||
| staff | No | Return the full employee ROSTER (id, name, staff no., status, designation, basic salary) alongside the summary — independent of any run, so it works even before the first payroll. This is where the employeeId for update_employee_draft comes from. Individual salaries: ask for it only when the user actually needs it. | |
| perEmployee | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that it is read-only and cannot create runs or file with authorities, that it reads from stored payslips (never recalculation), that it returns leave data, and that it states plainly when no run exists instead of returning zeros. It also notes the tool refuses non-payroll connections. These are substantial behavioral disclosures that enrich the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with essential information, and it is well-structured: it opens with the core output, then covers parameters, data source, and permissions. Every sentence adds value; however, it could be trimmed slightly without losing critical detail, making it a 4 rather than a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three optional parameters, no output schema, and a read-only annotation, the description is remarkably complete. It covers what the output includes, how to handle month, the purpose and sensitivity of perEmployee, how to obtain employee IDs for update_employee_draft, the data source (stored payslips), edge cases (no run), and role-based access. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (staff has a description; month and perEmployee do not). The description compensates fully: month is explained with format and default behavior, perEmployee is explained as sensitive and carrying real employee IDs, and staff is elaborated with its purpose for obtaining employee IDs. This goes well beyond the schema to clarify each parameter's semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns a full payroll summary from gross to net, including contributions, net pay, employer cost, and also lists upcoming approved leave. It distinguishes itself from other financial tools by specifying the exact components and the one-call capability for two distinct queries, and it explicitly references the sibling update_employee_draft for employee ID sourcing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool ('what does payroll look like this month' and 'who is out next week'), how the month parameter behaves (latest run if omitted, YYYY-MM for specific), and when to set perEmployee and staff flags, including a warning that perEmployee returns sensitive individual salaries. It also specifies role restrictions (ADMIN AND BOOKKEEPER + PAYROLL ONLY) and that it refuses other connections, providing clear guidance on appropriate invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
profit_driversProfit driversARead-onlyInspect
The DIAGNOSTIC breakdown of WHY net profit moved — deterministic, not a model's arithmetic. Compares the current month-to-date against the SAME day-span of the previous month (day 1 through min(today's day, the prior month's last day)), so a partial month is never compared against a full one. Returns: the net-profit figure for both windows and the delta; a drivers bridge (revenue delta, COGS delta, then the operating-expense accounts that moved most — top 5 by absolute change plus an other roll-up); oneOffs (disposal and FX gains/losses, pulled out of the movers so a lumpy asset sale doesn't read as an operating trend — empty if you have no such accounts); the prior month's FULL-month net profit as a stated secondary reference (priorMonthFull); and an assumptions list. Each driver's direction is its effect on PROFIT ('improving'/'worsening'/'flat'). All amounts are integer CENTS (RM = cents ÷ 100). Optional month (YYYY-MM) picks a month other than the current one — a completed past month compares its whole length; omit for the current month-to-date. On periodic stock with no stock take this financial year and COGS of RM 0 for the current window, a notes line says the gross-profit story is incomplete — COGS is not a real driver there, so do not present its absence as a margin gain.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | Month to analyse, YYYY-MM. Omit for the current month-to-date. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover read-only/open-world, so the description carries the behavioral load and does so richly: exact comparison windows, the drivers bridge structure, oneOffs handling of disposal/FX so a lumpy asset sale isn't read as a trend, integer-cents units, `direction` semantics relative to profit, and the periodic-stock COGS caveat that warns against misreading a missing COGS as margin gain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a long single paragraph but front-loads the core diagnostic purpose and then layers needed detail. Given there is no output schema, enumerating the return shape earns its place, though the prose is dense enough to be slightly heavy for a one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully enumerates the return shape (both net-profit windows, delta, drivers bridge, oneOffs, priorMonthFull, assumptions, notes), and covers edge cases and units, leaving nothing an agent needs in order to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one optional param, so the schema already documents `month`'s format. The description adds genuine semantic value beyond it by clarifying that a completed past month compares its whole length while the default is current month-to-date against the same day-span.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('diagnostic breakdown of WHY net profit moved') and explicitly frames it as deterministic decomposition rather than a generative model, which distinguishes it from siblings like income_statement or financial_timeseries that report rather than explain movement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance on the one parameter (omit `month` for current month-to-date; supply it to analyse a completed past month), and explains the comparison baseline. It does not, however, explicitly name an alternative sibling tool or say when to prefer this over income_statement, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_attachment_uploadRequest attachment uploadAInspect
Get a one-time link to upload a receipt/bill/statement/document that is too large to inline as attachmentBase64 (a real photo or a multi-page PDF — base64 inside a tool call is very costly). Returns an uploadUrl and an attachmentToken: PUT the file's RAW bytes to uploadUrl within the time limit, then pass attachmentToken IN PLACE OF attachmentBase64 to create_expense_draft / create_invoice_draft / create_bill_draft / create_quote_draft / create_receipt_draft / create_contact_draft / create_journal_draft — the file rides that draft and lands on the posted document on approval, exactly as an inline attachment does — or to import_bank_statement, which archives it as the statement's original. CALL THIS FIRST, BEFORE the create/import call: filing bare and adding the file after is the harder path. If you DID already file bare, the repair on the five document lanes (expense, invoice, bill, quote, receipt) is to call the same create tool AGAIN with the IDENTICAL reference + date + amount plus attachmentToken — while the draft is still pending the file is ADOPTED onto it (the result says duplicateAttachment:'added') and no second draft is created; the same repair works on import_bank_statement with the identical account, balances and rows. The link is single-use, expires quickly, and works only for this company. Use this for anything bigger than a few KB; keep attachmentBase64 for tiny files only.
| Name | Required | Description | Default |
|---|---|---|---|
| filename | No | Optional original filename, e.g. receipt.jpg — travels with the upload and shows on the review card. | |
| mediaType | No | Optional hint of the file's MIME type, e.g. 'image/jpeg'. Advisory only — the server determines the real type from the file's own bytes (magic bytes), so a wrong hint changes nothing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the sparse annotations, the description discloses many behavioral traits: the link is single-use, expires quickly, is company-scoped, accepts only raw bytes, and requires the token in place of base64. It also explains the duplicateAttachment adoption behavior and media-type handling, giving strong transparency for a non-idempotent, write-oriented tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but nearly every sentence carries operational value: upfront purpose, workflow ordering, repair scenario, lifecycle constraints. It is front-loaded well, though the repair paragraph could be tighter without losing critical guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description names the return values (uploadUrl, attachmentToken), explains how to use them, covers alternatives, and even handles the edge case of repairing a previously filed draft. This is complete enough for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description itself adds no parameter-specific meaning beyond what the schema already documents; it focuses on the returned uploadUrl and attachmentToken rather than filename/mediaType details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific, actionable goal: obtain a one-time link to upload a large file instead of inlining base64. It clearly distinguishes this tool from siblings by naming the create/import tools that consume the attachmentToken and contrasting it with attachmentBase64.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('anything bigger than a few KB'), when not to ('keep attachmentBase64 for tiny files only'), and the ordering requirement ('CALL THIS FIRST, BEFORE the create/import call'). It even gives a repair path for cases where the agent already filed bare.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_customerFind customerARead-onlyInspect
Resolve a customer name for an invoice draft against this company's real customers. Statuses: resolved (with id + the customer's current A/R balance, outstandingBalance — the credit-check-in-chat moment. It is NET: open invoices less any payment posted to this customer but not yet matched to an invoice, so it can be lower than the invoices add up to, and zero or NEGATIVE when they have paid ahead — say so rather than calling it a debt), ambiguous (with candidates), or none. When there is no exact/prefix match but existing customers look CLOSE (nearMiss:true), the name is likely salesperson shorthand for one of them ("JJ Dungun" for "PERNIAGAAN JJ") — ask the user which one; only file as new if they confirm it is genuinely new. Never guess among candidates — ask the user.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint=false, but the description goes well beyond that: it enumerates the three possible statuses, defines outstandingBalance as NET (open invoices less unmatched payments, can be zero or negative when paid ahead), and warns against calling a negative balance a debt. That is precisely the extra context an agent needs before speaking a number to a user.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and each subsequent clause carries distinct operational detail, but the dense nesting of parentheticals (the credit-check-in-chat aside, the balance arithmetic) makes it heavier than it needs to be. Nothing is wasted, though.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe returns — and it does, covering all three statuses and the fields returned for each (id, outstandingBalance, candidates, nearMiss). Nothing needed to call or interpret it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single param, so the description must carry the load. It conveys that the input is a customer name as written/abbreviated and that prefix and near-miss matching are attempted, which adds real meaning beyond the bare string type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Resolve) and resource (a customer name for an invoice draft) plus the scope (against this company's real customers). An agent can distinguish it from resolve_vendor/resolve_product by the resource alone without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing guidance for each outcome: ask the user which candidate when ambiguous, treat nearMiss as likely shorthand, only file as new on user confirmation, and never guess. This is genuine when/when-not guidance, not inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_productFind productARead-onlyInspect
Resolve a product by name or SKU for an invoice line. Returns id, sku, name, unit, the default sell price the form would use, and isService (true = "Service — no stock": selling it never counts stock or posts a cost). Plain names or SKUs work best. Statuses: resolved / ambiguous (candidates) / none — never guess among candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this read-only and closed-world, but the description adds substantial behavior: the exact return fields, the three possible statuses (resolved/ambiguous/none), the explicit ambiguity directive, and the isService semantics ("selling it never counts stock or posts a cost"). With no output schema, this disclosure is genuinely valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then adds return fields and status behavior. Dense but each clause carries information; the parenthetical isService explanation is long but earns its space by clarifying a non-obvious flag.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter resolver with no output schema, the description covers purpose, input expectations, return shape, and the ambiguity-handling contract. An agent has everything needed to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load. It does: the single `query` param is characterized as a product name or SKU, with "plain names or SKUs work best" implying the preferred input form. It doesn't specify matching rules or format edge cases, so not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Resolve) and resource (product) plus the concrete context (for an invoice line). It is cleanly distinguishable from the sibling resolve_customer and resolve_vendor by resource, and from update_product_draft by the resolve/read semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"for an invoice line" and "Plain names or SKUs work best" give clear usage context and input guidance, and the status handling ("never guess among candidates") steers behavior. It stops short of naming alternatives or stating when NOT to use this tool, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_vendorFind vendorARead-onlyInspect
Resolve a supplier / vendor name for a bill draft against this company's real vendors. Statuses: resolved (with id), ambiguous (with candidates), or none. When there is no exact/prefix match but existing vendors look CLOSE (nearMiss:true), the name is likely shorthand for one of them — ask the user which one; only file as new if they confirm it is genuinely new. Never guess among candidates — ask the user.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses important behaviors beyond the readOnlyHint annotation: return statuses (resolved, ambiguous, none), the nearMiss flag, and the rule to ask the user rather than guess among candidates. The phrase 'only file as new if they confirm' could sound like a write operation, but readOnlyHint is not contradicted because it is guidance for downstream action, not a claim that this tool mutates data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with its purpose and uses four sentences, each carrying useful information. There is slight redundancy between 'ask the user which one' and 'Never guess among candidates — ask the user,' but it reinforces a critical safety rule without becoming bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the main resolution paths: resolved, ambiguous, and nearMiss, along with explicit user-confirmation requirements. The main gap is the lack of explicit guidance for the status 'none' when no close vendors exist, although the nearMiss rule implies the general principle. Since there is no output schema, the status descriptions are essential and mostly sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining the single 'name' parameter is a supplier/vendor name from a bill draft that should be matched against the company's real vendors. This gives the parameter meaningful context that the bare schema does not provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Resolve a supplier / vendor name for a bill draft against this company's real vendors.' It clearly distinguishes from sibling tools like resolve_customer and resolve_product by being vendor-specific. The statuses it mentions further clarify its role as a resolution rather than a creation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: use this when reading a vendor name from a bill draft and needing to match it to existing vendors. It gives strong behavioral guidance for ambiguous and nearMiss cases, including 'ask the user' and 'never guess.' It does not explicitly name alternatives or state 'do not use for customers/products,' but the vendor-specific wording makes that implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revise_draftRevise a pending draftAInspect
Correct a draft you already filed, instead of filing it again. Pass the draft's kind and id plus a patch of just the fields that were wrong — the server re-validates them exactly as it did at filing (same account rules, same date rules, same line derivation, same totals) and updates the draft in place. Use this whenever you realise a filed draft is wrong: for an expense, re-calling create_expense_draft with corrected fields is refused as a duplicate of the very draft you are trying to fix, and for every kind a second create files a CONFUSING SECOND DRAFT the human then has to notice and reject — revise the one you filed instead. ONLY a PENDING draft can be revised; a draft the user has already approved or rejected (or one that has expired) is refused, naming its status — file a fresh draft in that case. Fields you do not send are left exactly as filed. REVISING DOES NOT APPROVE ANYTHING: the draft stays pending, nothing posts, and the human still taps Approve in Taokeh — the same one-tap link from the original filing still works and shows the revised figures. The reviewer is TOLD it changed: a short 'revised by your AI' line naming the changed fields is appended to the note they read on the approval card, so add a plain-language note saying WHY you revised it.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Which kind of draft to amend — the kind the filing tool returned (each create_*_draft and update_*_draft tool files one kind; 'invoice_update' / 'bill_update' are the posted-invoice and posted-bill corrections). 'sales_debit_note' is the SELL-side debit note create_debit_note_draft files (an additional charge on an invoice you issued); 'journal' is the adjusting entry create_journal_draft files. | |
| note | No | A SHORT plain-language reason for the revision, in the reviewer's language — 'the paid-from account should be the owner's capital account, not petty cash'. It is APPENDED to the note they already have (nothing is overwritten) and shows on the approval card. Keep it to a sentence. | |
| patch | Yes | Only the fields you are correcting. Every key must belong to this draft kind — a misspelled or foreign key is refused rather than silently ignored. The card IMAGE of a contact draft and the original document riding any other draft are deliberately NOT patchable: they are the evidence the approver checks your fields against. | |
| draftId | Yes | The draftId the create_*_draft call returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false hints (readOnly, openWorld, idempotent, destructive), so the description carries the full burden. It discloses that the draft updates in place, stays pending, nothing posts, the reviewer is told via an appended note line, fields not sent are left untouched, the card image is not patchable, and invalid keys are refused. This is rich behavioral context far beyond the sparse annotations, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though long, every sentence earns its place. It fronts the purpose, then routes usage, then constraints, then side effects for the reviewer. No filler or repetition of schema details; each clause adds a distinct fact needed to invoke the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested `patch` object, 23+ kinds, no output schema, and weak annotations, the description is remarkably complete. It covers failure modes (refused with status, duplicate refusal, invalid keys), lifecycle guarantees (stays pending, no posting), side effects on the reviewer (appended note), and excludes patchable evidence. Nothing an agent needs to decide whether to call it and how is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, which sets a baseline of 3. The description adds global patch semantics beyond the schema: patch may only contain fields belonging to the draft kind, foreign/misspelled keys are refused, and the evidence image is deliberately not patchable. It also explains that `note` should be in the reviewer's language and is appended, not overwriting. This adds meaningful context on top of an already detailed schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb ('Correct'), resource ('a draft you already filed'), and the core action ('instead of filing it again'). It explicitly distinguishes itself from sibling create/update tools by naming the duplicate-refusal behavior and the confusing-second-draft risk, so an agent can tell revise_draft apart without opening sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'when to use' ('whenever you realise a filed draft is wrong'), explicit 'when not to use' ('ONLY a PENDING draft can be revised; approved/rejected/expired drafts are refused — file a fresh draft'), and names the alternative actions (re-calling create_expense_draft is refused, a second create files a confusing second draft). No room for inference left.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_summarySales summaryARead-onlyInspect
How much you have sold. With no dates: today and this month so far. With a date range: the invoiced sales for that period, broken down by month, customer and product. Everything is MYR — a foreign-currency invoice is converted at the rate frozen on that invoice — and INVOICES only: credit notes are in none of these figures, so a period with returns is not netted here. ⚠ TWO TAX BASES, and they are named in the payload rather than left for you to guess: on byMonth and byCustomer, total is tax-INCLUSIVE and exTax is before SST; byProduct is a per-LINE figure and is always EX-TAX, so it foots to exTax, never to total. Quote one basis, say which, and never present a by-product figure as the same kind of number as a by-month one.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only readOnlyHint and openWorldHint, and the description adds substantial behavioral context beyond them: currency is MYR with foreign invoices converted at the invoice-frozen rate, credit notes are excluded so returns are not netted, and the two tax bases are disclosed with exactly which payload fields are tax-inclusive vs. ex-tax. This is precisely the kind of trap-avoiding detail an agent needs and cannot get from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core meaning, then layered detail in decreasing priority. Every sentence carries real information (currency basis, credit-note exclusion, tax bases), though the closing imperative about quoting one basis is advisory framing rather than tool semantics and could be trimmed without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe the return shape — and it does, naming byMonth, byCustomer, byProduct and the total/exTax fields and their differing tax bases. Combined with the mode explanation and exclusion rules, an agent has everything needed to invoke and interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for `from` and `to`, so the description carries the semantic burden — and it does explain that omitting them yields today/month-to-date while supplying a range yields that period's invoiced sales. It does not restate the YYYY-MM-DD pattern, but that is already in the schema, so the gap is minor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('How much you have sold') and immediately defines the two operating modes — no dates yields today plus month-to-date, a date range yields invoiced sales for that period. It further specifies the breakdown dimensions (month, customer, product), so an agent knows exactly what this returns and how it differs from sibling summaries like expenses_summary or income_statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the two usage contexts (no dates vs. with a date range), which effectively tells the agent when each behavior fires. What it does not do is name an alternative sibling (e.g. income_statement or financial_timeseries) or state when NOT to use this tool, so it stops short of the explicit routing found in the best definitions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_document_linesSearch document linesARead-onlyInspect
Read the LINE ITEMS of documents IN BULK — what was actually sold or bought, line by line, across many documents in one call. search_documents answers WHICH documents; this answers WHAT WAS ON THEM. Use it for any question that needs the detail behind the totals: average price per item across a year, how a price moved between two customers, which sizes or grades actually shift, how much of something was supplied. Without it the only way to see lines was to open one PDF per document — 180 documents is 180 round trips and 180 chances to misread a layout. docType is REQUIRED and must be one of: 'invoice', 'credit_note', 'sales_debit_note' (an additional charge this business issued to a customer), 'quote', 'delivery_order', 'bill', 'debit_note' (the BUY-side one, raised against a supplier), 'purchase_order'. Narrow with any combination of number (partial doc number — on invoices this also matches the customer's own Ref No, their DO/PO number), party (partial customer or vendor name), from/to (document date range) and productText — a case-insensitive contains matched against the PRODUCT NAME or the LINE DESCRIPTION, so "chengal" or "2x4" pulls only those lines. docType on its own is a legitimate whole-book pull. Each row: docId, docNumber, docDate, party, lineNo, product {sku, name}, description, quantity, unit, unitStamped, unitPrice, unitCost, discount, lineTotal, taxCode, docStatus, currency, fxRate, isMeasureDrift — verbatim from what was stored, never re-derived. CURRENCY: unitPrice, discount and lineTotal are in the DOCUMENT's own currency exactly as stored, NOT converted to ringgit; fxRate is the rate frozen on that document (1 unit of currency = fxRate MYR; 1 on a ringgit document). Convert with it yourself where you need ringgit, and never average or total rows of different currencies together unconverted. (sales_summary and search_documents report ringgit instead — converted once per document — so a foreign document's figures differ between them and here by exactly that conversion.) ⚠ Sell-side unitCost is ALWAYS ringgit (the stock cost). MARGIN — ONE RULE, in ringgit: sell-side line margin (RM) = lineTotal × fxRate − quantity × unitCost. lineTotal is the line AFTER its discount and BEFORE SST; unitCost is already ringgit; fxRate is 1 on a ringgit document. Never use unitPrice − unitCost: it ignores the discount and, on a foreign-currency line, subtracts ringgit from foreign money. On an isMeasureDrift row count its lineTotal but NO cost (it is a money correction, not a unit of goods — the posted cost of goods leaves it out). On the BUY side (bill, buy-side debit note) unitCost is the supplier's price in the DOCUMENT currency, not ringgit — multiply by fxRate for ringgit. MEASURE ROUNDING ROWS: isMeasureDrift true marks the measure engine's rounding correction — a line of quantity 1 at a small negative price (typically described 'Discount') that brings a measured line back to the true per-foot money. It is a real line and belongs in document totals, but EXCLUDE isMeasureDrift rows when computing an average price or quantity per product, or they drag the average down and add a phantom unit. It is null on bills, buy-side debit notes and purchase orders, which store no such row. ⚠ The flag is only reliable from the date each table started storing it — invoice, credit-note and sell-side debit-note lines from 4 Sep 2026, delivery-order lines from 19 Sep 2026, quote lines from 23 Sep 2026. Older rows all read false, and nothing was backfilled, so a line written before that date CAN be an unmarked rounding row; Taokeh does not guess which (a qty-1 negative line may equally be a real discount or fee). Say so if your answer leans on older lines. ⚠ DIMENSIONS AND SPECIFICATIONS ARE TEXT, NOT FIELDS. Taokeh has no column for a timber size, a grade, a colour or a variant: they live inside the product name and the line description as the business types them ("Chengal 2x4x8"). So this hands you that text as written and YOU do the parsing and the grouping — the server will not invent a dimension it does not store, and two lines for the same real size may be spelled differently. unit is the unit of measure stamped on the line when it was written; where the line stores none it falls back to the product's CURRENT unit, and unitStamped says which: true = the line's own stamp, false = the product stand-in. A false is right for an ordinary unmeasured line, but a MEASURED line (tons, feet) written before its table could store the unit also reads false and may show the product's unit (often 'ea') — invoice/bill lines before 3 Sep 2026, delivery-order lines before 19 Sep 2026, quote lines before 23 Sep 2026, and purchase-order lines written before the deploy that added their unit column (migration 20260923120000, late September 2026). So never assume pieces, and do not treat a false-unitStamped quantity as a measured one without checking the description. Rows come back newest document first, then in the document's own line order, capped at 200 with total, shown and more — when more is true, narrow by date (from/to) and pull the periods in turn rather than accepting a partial answer as the whole. unitCost is the COST BASIS STAMPED ON THE LINE WHEN IT WAS POSTED — like unit and the tax code beside it — so a historical line reports the cost as of THAT SALE, not the product's cost today. That is what makes per-product margin answerable here, line by line — see MARGIN below. On a BILL, a buy-side debit note or a PURCHASE ORDER there is only one price column, so unitCost and unitPrice are the same figure — what the SUPPLIER charged (on a PO: what was ordered at). On a QUOTE unitCost is null (a quote stores no cost). On a DELIVERY ORDER it is the product's moving-average cost captured when the DO was saved — a dispatch-time reference, NOT the cost of goods: the invoice the DO becomes re-reads cost when it posts, and that is the figure on the invoice line. QUOTES AND DELIVERY ORDERS ARE NOT SALES. Voided ones are left out (as voided posted documents are); every other row carries docStatus — quote OPEN or CONVERTED, delivery order OPEN or INVOICED (null on posted types). ⚠ A CONVERTED quote's lines and an INVOICED delivery order's lines ALSO appear as invoice lines (and a quote converted into a delivery order appears on that DO too), so never add a quote or DO pull to an invoice pull — for what was actually sold, use 'invoice'; for what is still only quoted or delivered-not-billed, keep docStatus OPEN. PURCHASE ORDERS ARE NOT PURCHASES either: a CANCELLED one is left out, every other row carries docStatus DRAFT, SENT or CONVERTED, and a CONVERTED order's lines ALSO appear as the bill's lines — for what was actually bought, use 'bill'. Before you propose a correction to an existing quote, delivery order or purchase order (update_quote_draft / update_delivery_order_draft / update_purchase_order_draft), read its lines here first: those tools take the FULL replacement line set, and each row's lineNo is what their keep entries refer to. Header-only search: search_documents. Paid expenses (which have no product lines): search_expenses. The ledger postings behind a document: search_journal. Nothing is written, and no draft is created.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Latest DOCUMENT date (YYYY-MM-DD). | |
| ref | No | NOT a filter on this tool. Pass the customer's own Ref No (their DO / PO number) as `number` — that one parameter searches BOTH the invoice's own number and the customer's Ref No, and the row it returns reports each separately as `number` and `refNo`. | |
| sku | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). | |
| from | No | Earliest DOCUMENT date (YYYY-MM-DD). The main way to break a >200-line pull into honest slices. | |
| item | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). | |
| size | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). There is no size field to filter on. | |
| text | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). | |
| docId | No | NOT a filter on this tool. This tool reads lines across MANY documents — narrow by `number` (the document number), `party` or a date range instead. For one document, pass its number as `number`. | |
| limit | No | NOT a filter on this tool. The page size is fixed at 200 lines. Narrow with from/to, party or productText and pull the periods in turn. | |
| party | No | Partial customer or vendor name (case-insensitive contains). | |
| refNo | No | NOT a filter on this tool. Pass the customer's own Ref No (their DO / PO number) as `number` — that one parameter searches BOTH the invoice's own number and the customer's Ref No, and the row it returns reports each separately as `number` and `refNo`. | |
| number | No | Partial doc number (case-insensitive). On INVOICES this also matches the customer's own Ref No — their DO or PO number. | |
| vendor | No | NOT a filter on this tool. Pass a customer or vendor name as `party`. | |
| docType | No | REQUIRED. One of: 'invoice', 'credit_note', 'sales_debit_note', 'quote', 'delivery_order', 'bill', 'debit_note', 'purchase_order'. Note the two debit notes: 'sales_debit_note' is one this business ISSUED to a customer; 'debit_note' is the BUY side, against a supplier bill. | |
| product | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). | |
| customer | No | NOT a filter on this tool. Pass a customer or vendor name as `party`. | |
| supplier | No | NOT a filter on this tool. Pass a customer or vendor name as `party`. | |
| docNumber | No | NOT a filter on this tool. Pass the document number as `number` (partial matches are fine). | |
| productId | No | NOT a filter on this tool. This search matches product TEXT, not ids — pass the name or size words as `productText`. | |
| reference | No | NOT a filter on this tool. Pass the customer's own Ref No (their DO / PO number) as `number` — that one parameter searches BOTH the invoice's own number and the customer's Ref No, and the row it returns reports each separately as `number` and `refNo`. | |
| description | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). | |
| productName | No | NOT a filter on this tool. Pass it as `productText` — one case-insensitive contains matched against BOTH the product name and the line description, which is where sizes, grades and variants live (Taokeh stores no column for them). | |
| productText | No | Case-insensitive text matched against the PRODUCT NAME or the LINE DESCRIPTION — this is where sizes, grades and variants live, because Taokeh stores no column for them. e.g. "chengal", "2x4". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the readOnlyHint and openWorldHint annotations: fixed 200-row cap with total/shown/more, ordering rules, currency conversion guidance, margin formula, docStatus semantics, isMeasureDrift caveats, unit fallback behavior, and historical cost-basis explanation. It ends with 'Nothing is written, and no draft is created,' reinforcing the read-only annotation without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very long, but appropriately so for a high-stakes tool with many edge cases around currency, margin, measure drift, and document types. It front-loads the core purpose and use cases before diving into caveats, and uses bold labels and clear sectioning to make the length navigable, though a tighter edit would improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description carries the full burden of explaining return rows, ordering, pagination, currency behavior, margin calculation, and per-document-type semantics — and it does so comprehensively. It covers buy-side vs sell-side differences, voided/converted document statuses, and data-reliability cutoffs, leaving no critical gap for an agent deciding whether and how to call this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Though schema coverage is 100%, the description adds meaning the schema alone cannot convey: docType is REQUIRED despite the schema not enforcing it, number also matches the customer's Ref No on invoices, productText matches both product name and line description, and many schema parameters (ref, sku, item, size, docId, limit, etc.) are explicitly marked as NOT filters with instructions on what to use instead. This is far beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read the LINE ITEMS of documents IN BULK' and immediately contrasts with search_documents ('answers WHICH documents; this answers WHAT WAS ON THEM'). This makes the tool's scope unmistakable and distinguishes it from closely named siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool ('any question that needs the detail behind the totals'), what it replaces (opening one PDF per document), and names concrete alternatives: search_documents for header-only, search_expenses for paid expenses, search_journal for ledger postings. It even warns to read lines here before calling draft-update tools that take full replacement line sets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_documentsSearch documentsARead-onlyInspect
Find invoices, credit notes, sell-side debit notes, quotes, delivery orders, bills, buy-side debit notes and purchase orders by any combination of: doc number (partial), party (customer/vendor) name (partial), date range, amount range, due-date window (invoices and bills), and doc type. Returns compact rows (docId, type, number, date, party, total, status), newest first, capped — with a more flag when there are further matches to narrow down. INVOICE rows carry two more fields: refNo, the CUSTOMER's own reference for that job — their delivery-order or purchase-order number, the one they quote back at you on the phone — and dueDate, the date that invoice falls due. refNo is NOT the invoice number; the invoice's own number is number. refNo is null when the invoice was raised without one. dueDate is null on a CASH sale — paid at the counter, so it is not money anyone is waiting on and it deliberately sits outside every due-date window and aging bucket; a null there means CASH, not a missing date. Only CREDIT invoices report one. BILL rows carry dueDate too — the date a supplier bill falls due, from the payment term printed on it, and null on a CASH bill for the same reason. Only invoices carry refNo (a bill's own number IS the supplier's number, so it needs no second field), and only invoices and bills carry dueDate — every other type omits them rather than reporting null about a column its document has no such thing as. number searches BOTH, so "find INV-0031" and "find the invoice for their DO 1074" are the same call. Invoices, credit notes, bills and quotes also report whether their ORIGINAL document (a scan or upload) is on file: hasAttachment plus an attachments list you can fetch with get_attachment. Delivery orders and purchase orders omit that field because they cannot carry an original. INVOICES, CREDIT NOTES and SELL-SIDE DEBIT NOTES (sales_debit_note — additional charges this business added to an invoice it issued; not to be confused with debit_note, the buy-side one raised against a supplier bill) also carry an einvoice field — this business's MyInvois posture for that document AS RECORDED IN TAOKEH: validated (a MyInvois validation is recorded here, with uuid + validatedAt), consolidated (covered by a consolidated e-invoice, which holds the uuid — the individual document has none by design), platform (a marketplace sale: Shopee/TikTok Shop/Lazada issues the e-invoice, nothing for this business to submit), exported (put into a MyInvois batch export from Taokeh, nothing recorded back yet) or none. Taokeh CANNOT see the MyInvois portal, so say "no validation recorded in Taokeh" — never "never submitted to LHDN". Bills and buy-side debit notes omit the field. Use the docId with get_attachment (the SOURCE document someone filed) or get_document_pdf (the PDF Taokeh generates). PAYMENTS — docType:'payment' — are an OPT-IN type you must ask for BY NAME: they are never included when you omit docType, because a payment is not a document. It is a BANK LINE that invoices or bills point at, so the grain is the movement of money: one receipt that settles three invoices is ONE row carrying three allocations ({docId, docType, docNumber, party, amount, amountForeign, notes}). An allocation's docType is the real kind of the document settled — 'invoice', 'credit_note' (a refund walks a credit note back), 'sales_debit_note', 'bill' or 'debit_note' — not an assumption from which side of the books it sits on, and notes is that allocation's own note rather than one note borrowed for the whole transfer. That makes "which payment covered INV-0031" a single call — number and party here match the ALLOCATED documents, not the bank line, and on the invoice side number also matches the customer's own Ref No (their DO/PO number, or a marketplace order key), exactly as it does for invoices in the ordinary search. Each row: docId (⚠ a BANK TRANSACTION id, NOT a document id — do not hand it to get_document_pdf or get_attachment), date, direction ('in'/'out'), total (the bank line's own magnitude), status (the bank row's status), bankAccount, party (the single customer or supplier when every allocation is to the same one, else null — read the per-allocation party then), source and description (the bank line as stored). source is DERIVED from where the line came from, not stored: 'statement' (it arrived on an imported bank statement and a human matched it in Banking), 'receipt_draft' or 'payment_draft' (an AI-filed draft the owner approved), or 'manual' — recorded by hand in the app, which is EVERYTHING ELSE rather than one door: the receipt door, the /go command bar, the supplier-payment door, a refund against a hand-recorded line. Taokeh does not distinguish them, so do not tell the owner which screen was used. A NEGATIVE allocation amount is a REFUND walking that document back down, not a settlement. CONTRAS ARE NOT PAYMENTS and never appear here: a set-off moves no bank money, and Taokeh books it as two internal clearing lines which are excluded from this read on purpose. To see one, use search_journal or the contra note on the documents themselves.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| ref | No | NOT a filter on this tool. Pass the customer's own Ref No (their DO / PO number) as `number` — that one parameter searches BOTH the invoice's own number and the customer's Ref No, and the row it returns reports each separately as `number` and `refNo`. | |
| from | No | ||
| text | No | NOT a filter on this tool. This search matches document numbers and party names only — use `number` or `party`. To search memo/line text, use search_journal (`text`) or search_expenses (`text`). | |
| dueTo | No | Latest DUE date (YYYY-MM-DD). Same docType rule as dueFrom — 'invoice' or 'bill'. Pair the two for a window: dueFrom + dueTo across this week with docType:'bill' answers "what do I owe before Friday"; dueTo alone with a past date is everything already due. | |
| party | No | Partial customer or vendor name (case-insensitive contains). | |
| refNo | No | NOT a filter on this tool. Pass the customer's own Ref No (their DO / PO number) as `number` — that one parameter searches BOTH the invoice's own number and the customer's Ref No, and the row it returns reports each separately as `number` and `refNo`. | |
| number | No | Partial doc number / reference (case-insensitive). On INVOICES this also matches the customer's own Ref No — their DO or PO number, whatever they printed on the paperwork. So a customer chasing "our PO 4471" is one call: pass 4471 here. That Ref No is NOT the invoice number; the invoice's own number matches here too, and you cannot tell from a hit which of the two it was — read `number` and `refNo` on the row it comes back with. On MARKETPLACE-IMPORTED invoices the Ref No is the platform's order key (SHOPEE-<order_sn>, TIKTOK-<order id>, SHOPIFY-<order id>), not a customer's number — so a marketplace order id pasted here finds its invoice too. | |
| vendor | No | NOT a filter on this tool. Pass a customer or vendor name as `party`. | |
| docType | No | Restrict to one type: 'invoice','credit_note','sales_debit_note','quote','delivery_order','bill','debit_note','purchase_order','payment'. Omit for all — EXCEPT 'payment', which is opt-in only and is never returned unless you name it. NOTE the two debit notes: 'sales_debit_note' is one this business ISSUED to a customer (additional charges on an invoice — money owed TO this business); 'debit_note' is the BUY side, issued to a supplier against a bill. | |
| dueFrom | No | Earliest DUE date (YYYY-MM-DD) — this is the "what's due this week" filter, and it asks about when the money falls due, not when the document was raised (that's `from`/`to`). INVOICES OR BILLS, one at a time: pass docType:'invoice' (money coming in) or docType:'bill' (money going out) with it, or the call is refused. Nothing else in this search has a due date, and quietly dropping every quote and delivery order would be a wrong answer wearing a right one's clothes; netting receivables and payables into one list would be another. | |
| customer | No | NOT a filter on this tool. Pass a customer or vendor name as `party`. | |
| supplier | No | NOT a filter on this tool. Pass a customer or vendor name as `party`. | |
| docNumber | No | NOT a filter on this tool. Pass the document number as `number` (partial matches are fine). | |
| maxAmount | No | ||
| minAmount | No | ||
| reference | No | NOT a filter on this tool. Pass the customer's own Ref No (their DO / PO number) as `number` — that one parameter searches BOTH the invoice's own number and the customer's Ref No, and the row it returns reports each separately as `number` and `refNo`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Extremely rich behavioral disclosure beyond the readOnlyHint: `docId` is a bank transaction id rather than a document id, dueDate null means cash, Taokeh cannot observe the MyInvois portal, source is derived not stored, and negative allocations mean refunds. These are exactly the kind of footguns an agent needs exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a very long wall of prose covering many edge cases, some of which repeat schema-level parameter descriptions. While organized into thematic paragraphs, it lacks bullet structure and is far larger than needed for an agent to parse quickly. It earns its content but not conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully compensates by explaining row shapes, field semantics, null meanings, attachment behavior, e-invoice values, payment allocations, and excluded contra entries. An agent has enough context to call the tool correctly and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 76% and the schema already documents several parameters, but the description adds important meaning: `number` matches both invoice number and customer Ref No, `docType` has the payment opt-in rule, the two debit-note types are differentiated, and due-date windows apply only to invoices and bills. Some parameter details are duplicated from the schema, so it does not max out, but it clearly adds substantive semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb ('Find') and enumerates the exact document types and filter combinations, then describes the returned rows. This clearly distinguishes the tool from document-line, journal, and expense searches, so an agent can confidently select it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong situational context: when to use due-date filters, how payments are opt-in, and explicitly routes contra entries to search_journal. It doesn't compare to every sibling, but it gives enough exclusion and alternative guidance for safe selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_expensesSearch paid expensesARead-onlyInspect
Search PAID EXPENSES already posted to the books (the /expenses ledger) by any combination of: reference (partial), text in the memo/description (partial), date range, amount range, and category account. Returns compact rows (expenseId, date, amount, category account code + name, memo, reference), newest first, capped — with a more flag. Every row also says whether the ORIGINAL RECEIPT is on file: hasAttachment plus an attachments list (filename, media type, size) — an empty list means genuinely no receipt is attached, not "unknown". To read one, pass the row's expenseId and attachmentId to get_attachment. ALWAYS check here BEFORE filing an expense draft (create_expense_draft): if the same receipt is already booked, filing again would double-book it. search_documents does NOT cover paid expenses — this tool is the only way to see them.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| ref | No | NOT a filter on this tool. Pass the supplier/receipt reference as `reference`. | |
| from | No | ||
| memo | No | NOT a filter on this tool. Pass memo text as `text`. | |
| text | No | Partial text to match in the expense memo/description (case-insensitive contains). | |
| refNo | No | NOT a filter on this tool. Pass the supplier/receipt reference as `reference`. | |
| number | No | NOT a filter on this tool. Pass the supplier/receipt reference as `reference` — a posted expense has no document number of its own. | |
| vendor | No | NOT a filter on this tool. A posted expense carries no party field here — search the supplier name as `text` (it matches the memo/description) or as `reference`. | |
| account | No | NOT a filter on this tool. Pass the chart-of-accounts code as `accountCode` (exact). | |
| category | No | NOT a filter on this tool. Pass the category by its chart-of-accounts code as `accountCode` (exact) — get codes from expense_accounts. | |
| supplier | No | NOT a filter on this tool. A posted expense carries no party field here — search the supplier name as `text` (it matches the memo/description) or as `reference`. | |
| maxAmount | No | ||
| minAmount | No | ||
| reference | No | Partial supplier/receipt reference (case-insensitive contains). | |
| accountCode | No | Restrict to one expense category by its account code (exact). | |
| description | No | NOT a filter on this tool. Pass memo text as `text`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the read-only safety is already known. The description goes further by explaining return-row composition, newest-first ordering, capping with a `more` flag, and the exact meaning of an empty `attachments` list versus an unknown state. This is valuable behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every sentence carries operational meaning: search scope, return shape, attachment semantics, attachment retrieval, double-booking prevention, and sibling-tool distinction. Information is front-loaded and logically ordered, though it could be tightened slightly without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully compensates by specifying the returned fields, ordering, cap behavior, and attachment representation. It also covers the key workflow context (check before create_expense_draft) and the relationship to get_attachment and search_documents. Nothing essential is missing for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers 75% of parameters with useful descriptions, including NOT-filter disclaimers for ref, memo, vendor, and similar aliases. The description adds the 'any combination' semantics and maps search concepts like reference, text, date range, amount range, and category account onto the intended use. It does not enumerate every parameter name, but the schema plus description together are clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: search expenses that are already posted/paid in the /expenses ledger. It lists the exact filter dimensions and explicitly distinguishes itself from search_documents, which does not cover paid expenses. This gives an agent a precise, non-ambiguous purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: check this tool before creating an expense draft to avoid double-booking a receipt, and use get_attachment to read attachments. It also names search_documents as a non-alternative, making the usage boundary crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_journalSearch journal entriesARead-onlyInspect
Search the general ledger's JOURNAL ENTRIES — every posting, whatever door created it (a manual journal, a scanned receipt, an invoice, a bank row, a paid expense, payroll, depreciation…). Returns each entry WITH its balanced lines (account code + name, debit, credit, line description), so you can see exactly HOW something was booked, not just that it exists. Filter by any combination of: date range, text (case-insensitive, matched against the entry memo AND its line descriptions), accountCode (entries touching that account), source (the door that created it), and an amount range on the entry's total debits. Give at least one filter — this never dumps the whole ledger. Newest first, capped, with a more flag. Pairs with search_expenses for reconciling: search_expenses shows what the /expenses register holds, search_journal shows every ledger entry including journal-era postings the register never covered. Read-only — it changes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| ref | No | NOT a filter on this tool. A journal entry has no reference field — search its memo and line descriptions with `text`. | |
| from | No | ||
| memo | No | NOT a filter on this tool. Pass memo text as `text` (it matches the entry memo AND every line description). | |
| text | No | Partial text matched against the entry memo AND its line descriptions (case-insensitive contains). | |
| party | No | NOT a filter on this tool. The ledger has no party column — search the name as `text`, or find the document with search_documents. | |
| refNo | No | NOT a filter on this tool. A journal entry has no reference field — search its memo and line descriptions with `text`. | |
| number | No | NOT a filter on this tool. A journal entry has no document number — search its memo and line descriptions with `text`, or find the document itself with search_documents. | |
| source | No | The door that created the entry (exact, case-insensitive). One of: MANUAL, SALE, PURCHASE, ADJUSTMENT, BANK, CREDIT_NOTE, DEBIT_NOTE, PAYROLL, IMPORTED_SERVICE, FX_REVAL, DEPRECIATION, ASSET_DISPOSAL, EXPENSE, LOAN, REVENUE_RECOGNITION, ASSET_ACQUISITION, BANK_OPENING. Omit for all. | |
| account | No | NOT a filter on this tool. Pass the chart-of-accounts code as `accountCode` (exact). | |
| maxAmount | No | Maximum entry size, measured on the entry's total debits (MYR). | |
| minAmount | No | Minimum entry size, measured on the entry's total debits (MYR). | |
| reference | No | NOT a filter on this tool. A journal entry has no reference field — search its memo and line descriptions with `text`. | |
| accountCode | No | Restrict to entries that have a line on this account code (exact). | |
| description | No | NOT a filter on this tool. Pass the text as `text` (it matches the entry memo AND every line description). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful behavior: results are 'Newest first, capped, with a more flag', and each entry comes with balanced debit/credit lines. The only redundancy is the closing 'Read-only — it changes nothing', which merely restates the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but front-loads the core purpose, then moves through return shape, filters, constraints, and sibling relationship. The enumeration of creation doors is slightly verbose but reinforces the 'every posting' scope; no wasted sentences overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter tool with no output schema, the description covers the return payload, ordering, cap indicator, filter semantics, required filter, and sibling distinction. An agent has enough to select and invoke this tool correctly without opening the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 87%, so the baseline is solid; the description still adds value by explaining the filter contract ('any combination of'), clarifying text matches 'the entry memo AND its line descriptions', and requiring at least one filter. It also orients the amount range to 'the entry's total debits', which matches the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb-resource combination: 'Search the general ledger's JOURNAL ENTRIES — every posting, whatever door created it'. It further clarifies the output includes balanced lines, which differentiates it from a mere existence-check and from sibling search_expenses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Give at least one filter — this never dumps the whole ledger' and names the sibling alternative: 'Pairs with search_expenses for reconciling: search_expenses shows what the /expenses register holds, search_journal shows every ledger entry including journal-era postings the register never covered.' This gives clear selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shoebox_itemsShoebox items waitingARead-onlyInspect
List the photos sitting in this company's SHOEBOX — paper the people in the business snapped on their phones and sent in, which NOBODY has read yet. This is the pile to work from when the user asks you to "do the receipts" or "clear the shoebox". Only items still waiting are listed: anything already booked or set aside is settled and deliberately absent. For each item you get its id, what kind of paper it is, the file name, type and size, when it arrived (Malaysia time) and who sent it in — never an amount, because nothing has been read. To SEE one, call get_attachment with owner:'shoebox' and the item's id; you read the photo yourself, on your own subscription — Taokeh does not OCR or interpret it for you. To BOOK one, file the matching draft (create_expense_draft, create_bill_draft or create_invoice_draft) with shoeboxItemId set to that id, and DO NOT re-send the photo: the server attaches its own stored copy, so it rides the draft and lands on the posted document on approval. An item already carrying a pending draft says so (pendingDraft) — file nothing more against it; correct the existing draft with revise_draft instead. An item TAOKEH itself is already reading says so too (beingRead): someone tapped "Book it" or asked Taokeh to book the shoebox, that read is paid for and its draft is waiting for the owner at the reviewPath given — file nothing against it either, and tell the user where it is waiting. Reading this pile costs the company no AI credits.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many waiting items to return, newest first. Defaults to 50. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and the description reinforces a safe, read-only profile, adding a cost disclosure ('Reading this pile costs the company no AI credits'). It transparently discloses return-field semantics — explicitly 'never an amount' — plus the fact that Taokeh does not OCR/interpret photos and the agent must read them itself on its own subscription. The material is rich, though the corrupted run-on in the middle makes some disclosures harder to parse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is far longer than necessary for a one-parameter list tool, and the middle section is visibly corrupted — repeated phrases, missing punctuation, and a broken 'two-send the photo' fragment. Purpose and cost notes are nicely front-loaded, but the excessive, garbled prose fails the conciseness test.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema, the description covers the essential operational facts: exact fields returned, how to view an item, how to book it, and which state flags (pendingDraft, beingRead) to respect before acting. Nothing an agent needs to invoke it correctly is missing; the garbled prose hurts readability but not coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the single optional `limit` parameter is fully documented with default (50), maximum (200), and ordering ('newest first'). The tool description adds nothing about `limit`, so the baseline of 3 applies — the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource — 'List the photos sitting in this company's SHOEBOX' — and sharply scopes it to items nobody has read yet. It differentiates from siblings by naming the viewing tool (get_attachment) and booking tools (create_*_draft), so an agent knows this is the unread-inbox list, not the action tool. The garbled middle section does not obscure this core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger phrases: 'the pile to work from when the user asks you to do the receipts or clear the shoebox'. Routes next steps to named alternatives (get_attachment to view, specific create_*_draft tools to book) and states exclusion rules (don't file against items with pendingDraft or beingRead). This is about as explicit as usage guidance gets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stage_documentStage a historical documentAInspect
Bring ONE already-issued historical document (a sales invoice, or a supplier bill) across from the system this company is migrating FROM. This is the BULK migration door and it is NOT create_invoice_draft: nothing is posted, nothing is drafted for individual approval, and no approval card is raised per document. Staged documents group into monthly batches the owner reviews and approves together at /switch/documents. Use it ONLY for documents that were genuinely issued in the old system (or written in a paper book) — a NEW document belongs in create_invoice_draft / create_bill_draft. The lane must be OPEN (the owner turns it on at /switch → "Bring over your documents") and it can only be open before the books are locked; if it is closed the call is refused with instructions. IDEMPOTENCY IS ON sourceDocId: staging the same sourceDocId again UPDATES that pending row instead of filing a second one. Once the owner has ENTERED it, or deliberately LEFT IT OUT, nothing you send can change it, re-book it or bring it back — only the owner can. So re-running your whole export is always safe. THE CUTOVER: bring over only documents dated AFTER the company's accounting start date. Anything dated on or before it is already carried by the owner's opening balances, and Taokeh will flag it and refuse to enter it — staging those wastes both our time, so filter them out of your export if you can. ORDER MATTERS: stage the SUPPLIER BILLS for a period before the sales invoices for it, because Taokeh works out cost of sales from the stock that was bought. The server re-computes every quantity and the grand total from the lines (with SST); your own printed total goes in sourceTotal and is used ONLY to show the owner a tie against the server's figure. Anything the server cannot settle — a product that does not resolve, a customer name matching several contacts, a printed total that disagrees — is staged anyway, FLAGGED, and held out of bulk approve for the owner to open individually. The document keeps its ORIGINAL number (reference) and its ORIGINAL date, and posts marked as historical so Taokeh never e-invoices or chases it. PAPER: if you read this document off a photo or a scan, ATTACH IT — request_attachment_upload, PUT the bytes, pass attachmentToken here. The owner then reviews your figures beside the actual slip on that document's own screen, edits anything you misread, and approves it there; the original lands on the posted document. Without it they are approving your arithmetic on your word alone. Show your per-line working, and mark any balancing/catch-all line residual:true — a residual line carrying a material share of the total is flagged for the owner, because that is exactly where a misread hides. NOT EVERY LINE IS A CATALOGUED PRODUCT, and an old document is full of the ones that are not: a delivery or transport charge, a labour or installation line, a service fee, a rounding line. Those still go through as ordinary lines — productRef is REQUIRED and must not be blank, so put the line's OWN wording in it ("Delivery charge"), and Taokeh CREATES that wording as a non-stock item (MD, 2026-09-09) so the document derives normally — the row is flagged, the created item is named for the owner, and it carries no stock and no price of its own. That is the designed path: it is not an error, so do not drop the line, do not fold its amount into another line, and do not invent a SKU for it. Omitting productRef altogether is the one thing that fails — the call is refused at the boundary before anything is staged. ⛔ ROW KEYS ARE STRICT (MD, 2026-09-09): a key the schema does not list is REFUSED BY NAME — with the key it probably meant, e.g. sell_price → unitPrice — and NOTHING is staged. Send the contract's apiField, not the spreadsheet header. Unknown keys used to be dropped in silence, so a set could stage "successfully" with its prices missing; that is the bug this refusal closes. NOT IN THE CATALOGUE? IT IS NOW: a line whose productRef matches NOTHING is created as a NON-STOCK item named after the line's own wording (no stock, no price of its own, marked as created by the migration), so the document derives and the owner reviews figures instead of a blocked row. Tell the owner which items were added — the reply names them. A ref that matches SEVERAL products is still flagged rather than duplicated. ⚠ NOT ON A DOCUMENT THE CUTOVER FENCE REFUSES: Taokeh will not grow the catalogue for a document it can never enter, so a pre-cutover row stages with no figures and an unresolved-product note ALONGSIDE its fence. The fence is the blocker; the product note is a consequence of it. Do not call resolve_product for such a row, do not ask the owner to add a catalogue item for it, and never route it to the manual bill form or the expense door — the only answers are to leave it out or to have the owner move the accounting start date.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | ||
| notes | No | A SHORT note for the owner about this document — one or two sentences naming anything they should check. Not a place for lengthy reasoning. | |
| party | No | The customer (sales invoice) or supplier (purchase bill) name as printed. Matched against the existing contacts; a name matching several, or one that looks like shorthand for an existing contact, is flagged for the owner rather than guessed. Omit for a walk-in cash sale with no named customer. | |
| docDate | Yes | The date the document was ORIGINALLY issued, YYYY-MM-DD. Never today unless it really was today — the whole point is that these documents keep their real dates. | |
| docKind | Yes | 'sales_invoice' = an invoice this company ISSUED to a customer. 'purchase_bill' = a bill a supplier issued TO this company (inventory lines only, same rule as create_bill_draft). Other kinds are not supported yet. | |
| partyId | No | The Taokeh customer/vendor id, when you already resolved it with resolve_customer / resolve_vendor. Wins over `party`. | |
| reference | No | The document number exactly as printed in the old system (e.g. INV1603). Taokeh keeps it, so the customer's copy and yours still match. It must be unique in this company: a number already on the books is refused by name. Leave it out only if the document genuinely has no number. Before the first such invoice is entered, the owner must confirm where the real invoice series continues; Taokeh never silently starts migrated books at INV1. | |
| needsReview | No | Set true when you want the owner to look at this one specifically. It is then held out of bulk approve and gets its own door. | |
| sourceDocId | Yes | The id this document has in the system you are migrating FROM — the primary key, row id or export id. THIS IS THE DEDUPE KEY: staging the same sourceDocId twice updates one row instead of booking the document twice, which is what makes re-running your export a no-op. If the old system has no id, use the document number, but use it consistently across the whole export. | |
| sourceTotal | No | The grand total PRINTED on the source document. Advisory only — Taokeh re-derives its own figure and shows the two side by side, so the owner can see a whole batch tie to the old system's report. Send it whenever the paper states one; it is the single most useful check in the whole migration. | |
| paymentMethod | No | 'CREDIT' (default) = it was on account, so it carries into receivables/payables. 'CASH' = it was settled at the counter. Get this right: it decides whether the document leaves a balance behind. | |
| attachmentToken | No | The ORIGINAL this document was read off — the photo of the handwritten cash-sale slip, the scanned invoice page. Call request_attachment_upload FIRST, PUT the raw bytes to its uploadUrl, then pass the token here (there is no inline base64 on this door: a backlog export would blow the context). The file rides the staged row, the owner sees it BESIDE your figures on the document's own review screen, and it lands on the posted invoice or bill automatically when they approve. Re-staging the same sourceDocId WITH a token replaces the file; re-staging without one leaves the file already there untouched, so a bare filing is repairable. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With sparse annotations (all hints false), the description carries the full burden and delivers richly: idempotency keyed to sourceDocId, the 'once entered or left out, only the owner can change it' irreversibility, strict row-key refusal at the boundary, non-stock item auto-creation, and the cutover fence refusing catalogue growth. The idempotentHint=false annotation is a coarse general hint while the description precisely scopes idempotency to a key — a nuance, not a hard contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is thorough but verbose — a wall of text with ALL-CAPS emphasis, emoji section markers (⛔⚠), and repeated warnings. For a tool with 12 params and many edge cases, depth is justified, but the aggressive formatting hurts scannability and several points (e.g., the non-stock path) are restated multiple times. Not under-specified, but not tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For the most complex tool in the set, nothing material is missing: when to use vs. siblings, dedupe semantics, cutover fence, ordering, flagging/held-out-of-approve behavior, original number/date preservation, historical posting so it is never e-invoiced, paper attachment, residual lines, strict keys, and non-stock creation. The reply content for created items is even disclosed ('the reply names them'). No output schema exists, but the return is adequately described inline.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 92%, so the schema already documents most parameters well. The description still adds genuine value: it explains the attachmentToken flow (call request_attachment_upload first, PUT bytes, then pass token), reinforces sourceDocId as the dedupe key, and clarifies productRef is required and auto-creates non-stock items. This exceeds the baseline 3 but the schema is doing most of the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource ('Bring ONE already-issued historical document... across') and immediately differentiates itself from create_invoice_draft by stating what it is NOT ('nothing is posted, nothing is drafted for individual approval, and no approval card is raised per document'). An agent can confidently distinguish it from the create_*_draft siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Extremely explicit: names the exact alternatives (create_invoice_draft / create_bill_draft) for NEW documents, states the lane must be OPEN and only before books are locked, gives the cutover date rule (documents dated AFTER the accounting start date), and orders supplier bills before sales invoices. No inference is left to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stage_master_dataStage master dataAInspect
Bring this company's MASTER DATA across from the system it is switching FROM: its CHART OF ACCOUNTS (every ledger code it keeps its books in), its PRODUCT CATALOGUE (every item it buys, stocks or sells, with the stock each was carrying at the cutover) or its CONTACT BOOK (every customer and supplier). Call intake_contract(doc_type: 'accounts' | 'products' | 'contacts') FIRST — it gives the exact fields, the rules and this company's own current counts. This WRITES NOTHING: the rows land on the owner's existing Import screens, where they confirm the columns and import them. THE UNIT IS THE WHOLE LIST: send every row of one kind in ONE call, and calling again for the same kind REPLACES the entire staged set rather than adding to it. Found a mistake? Send the corrected list in full. UNLIKE stage_opening_balances there is NO 'already imported' refusal, because master data is a CATALOGUE and not a ledger: the product upsert key is the SKU (an existing SKU is updated in place and ONLY the fields you actually sent are changed — a corrected three-field price list leaves its unit, barcode and buying-unit conversion untouched — its on-hand quantity is never touched by an import, and an opening quantity on a SKU that already exists is ignored), and contacts are de-duplicated case-insensitively by name — so staging a corrected list after an import is normal and double-counts nothing. OPENING STOCK IS THE ONE PART THAT TOUCHES THE BOOKS: a NEW SKU with both an opening quantity and a cost posts a real opening-stock journal entry (Dr Inventory / Cr Opening balance equity) when the owner approves, dated at this company's accounting start date while it is still migrating. Send the cutover quantity and the cutover cost, never a guess — a zero or missing cost still seeds the quantity but posts nothing. PRESENTATION FIELDS DO NOT MIGRATE: web addresses, storefront visibility and product images stay with the website, there is no field for them here, and you should TELL THE OWNER to keep those before the old system is switched off. Bill-of-materials links (combination SKUs built from components) are a separate sheet with their own review step and are NOT staged here — the owner uploads that sheet on the Import products page. Every contact row must say which side it is on (isCustomer / isVendor, or defaultRole for the whole list); Taokeh names any row that says neither straight back to you, so fix those before the owner sees the screen. THE CHART OF ACCOUNTS ('accounts') is the step to do FIRST, because the trial balance and every document afterwards land on its codes — and it is the quietest of the three: it writes NO journal line and carries NO figures, so opening balances still come across separately with stage_opening_balances(kind:'trial_balance'). Its key is the account CODE: an existing code is updated in place and keeps every posting attached to it, a new code is created, a code is never rewritten, and NOTHING IS EVER DELETED OR DEACTIVATED — an account missing from your list is simply left alone, so do not imply to the owner that it will disappear. A RENAME is always allowed on every account, including the ones Taokeh posts to automatically. A CHANGE OF TYPE OR NORMAL BALANCE IS REFUSED, row by row, on any account that already has postings or that a system role or bank account depends on, because it would silently restate reports the owner has already read and filed; intake_contract(doc_type:'accounts') marks exactly which accounts those are, so shape the list to avoid the refusal, and when a reclassification really is wanted, tell the owner the honest route — a new account plus a dated reclassification journal (create_journal_draft). THE COMPANY SETUP ('settings') is the LAST MILE: once the books are across, this is how the company itself stops being a form the owner types — its registered name, SSM registration number, TIN, address, state, statutory officer, the language it reads in, the measurement profile it prices by, and the date its books start. IT IS GOVERNED BY THE STRICT ALLOWLIST returned by intake_contract(doc_type:'settings'), and nothing outside it can be created: call intake_contract(doc_type:'settings') for the list, each key's meaning, its current value, and the settings an AI may NEVER write — payment credentials (a secret is never AI-written), publishing a storefront and advertising consent (outward and consent acts the owner takes), the AI and legal terms acceptances (an AI must never accept terms on the owner's behalf, least of all its own), the books lock (a money-integrity control) and every e-Invoice setting (they change statutory filing behaviour and follow the owner's own LHDN status). Those are refused BY NAME with the reason so you can tell the owner what to set rather than retrying. SEND ONLY WHAT THE OWNER TOLD YOU OR WHAT THE OLD SYSTEM PRINTS — an inferred TIN, registration number or employer number passes this import and fails at LHDN months later on a document that has already gone out; a blank is honest. The accounting start date is the highest-consequence key (it defines what counts as before the books started) and is refused if the company already has posted entries dated earlier; the measurement divisor is a NUMBER, not a label. Nothing is ever cleared: a setting missing from your list is left alone, and a row with a blank value is refused rather than read as 'erase this'. THE STAFF LIST ('employees') is the fifth kind and the one that behaves differently in two ways you must say out loud to the owner. It is ADMIN-ONLY — salaries and IC numbers are the most sensitive data in a company, and a non-admin connection is refused with the reason, deliberately. And ITS IMPORT CREATES RATHER THAN UPSERTS: employees have no unique key in Taokeh, so every ready row becomes a NEW person, anyone already in Taokeh must be left OFF the list, and a change to somebody who exists is made on their own employee page instead, or proposed by your AI with update_employee_draft. Use it for a whole payroll register being migrated AND for a single new hire the employer just described. SEND ONLY WHAT THE EMPLOYER ACTUALLY TOLD YOU: never construct or infer an IC number, a salary, a bank account, a date of birth or a statutory reference — a blank field is honest and gets filled in, a guessed one passes silently and then follows that person into EPF, SOCSO and LHDN submissions that have already gone out. A foreign employee has a passport and countryCode instead of an IC, with isMalaysian:false. The statutory flags DEFAULT ON when omitted (EPF, SOCSO, EIS, PCB, Malaysian, resident), exactly as the new-employee form does, so send false only on the employer's own word. Salary is MONTHLY and in ringgit. A date spoken loosely ('starting Monday') is yours to resolve to YYYY-MM-DD and to confirm back to them. Only ACTIVE staff consume a payroll seat: rows past the plan's allowance are reported as skipped at import rather than created, and historical leavers come across with status RESIGNED. This creates records only — it pays nobody, posts no journal line and files nothing; a payroll run is a separate act the owner takes afterwards. ⛔ ROW KEYS ARE STRICT (MD, 2026-09-09): a key the schema does not list is REFUSED BY NAME — with the key it probably meant, e.g. sell_price → unitPrice — and NOTHING is staged. Send the contract's apiField, not the spreadsheet header. Unknown keys used to be dropped in silence, so a set could stage "successfully" with its prices missing; that is the bug this refusal closes.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | 'products' = the item list / catalogue. 'contacts' = the customer and supplier book. 'accounts' = the chart of accounts. 'settings' = the company setup; its complete stageable and owner-only classification is returned by intake_contract(doc_type:'settings'). 'employees' = the staff list; admin-only and create-only. | |
| rows | Yes | EVERY row of the list, in the order the export prints them. This replaces any previously staged set for this kind. | |
| sourceName | No | What you read — the export's file name or the report's own title, e.g. 'Item list export (Financio) — 412 items'. Shown to the owner as the evidence for what they are approving. | |
| defaultRole | No | contacts only: the side a row that carries neither isCustomer nor isVendor lands on — the same choice the upload page offers as a Customers / Suppliers radio. Send it when the export is a single-sided list. Per-row flags always win over it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly=false, destructive=false, idempotent=false, openWorld=false. The description adds far more: re-calling REPLACES the entire staged set, employees is ADMIN-ONLY and CREATE-only, settings is allowlist-governed with named refusals, unknown row keys are refused by name, no deletions ever occur, and opening stock posts a real journal entry only on approval. This is exactly the beyond-annotation context the dimension rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, but the description runs to roughly 1,500 words with heavy caps emphasis and repeated admonitions ('never construct or infer') across kinds. For a five-kind mega-tool some length is warranted, yet there is noticeable redundancy that could be compressed without losing signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a very complex tool, the description covers ordering, per-kind behavior, refusal modes (unknown keys, type/normal-balance changes, admin-only, allowlist-by-name), and what does not migrate. An agent has everything needed to stage correctly and to explain outcomes to the owner.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each row field already carries a rich description, so the schema does the heavy lifting and baseline 3 applies. The description reinforces key semantics (SKU and account code as keys, case-insensitive contact dedupe, defaultRole) but adds little about parameter syntax not already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (stage/bring across) and resource (master data) and enumerates the five kinds (accounts, products, contacts, settings, employees) with what each represents. It explicitly distinguishes itself from siblings stage_opening_balances, intake_contract, create_journal_draft and update_employee_draft, so an agent can route correctly without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing (call intake_contract FIRST; accounts is the step to do FIRST), names the replacement siblings (stage_opening_balances for trial balance, create_journal_draft for reclassification, update_employee_draft for existing staff), and states when-not to use it (no BOM links, presentation fields do not migrate). Nothing about selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stage_opening_balancesStage opening balancesAInspect
Bring this company's OPENING BALANCES across from the accounting system it is switching FROM: its closing TRIAL BALANCE, its AGED RECEIVABLES (invoices customers still owed) or its AGED PAYABLES (bills still owed to suppliers). Call intake_contract(doc_type: 'trial_balance' | 'aged_receivables' | 'aged_payables') FIRST — it gives the exact fields, the rules, and for the trial balance this company's own chart of accounts. This posts NOTHING: the rows land on the owner's existing "Switch to Taokeh" review screens, where they match each account to their chart (trial balance) or check each party (aged lists) and post it themselves. THE UNIT IS THE WHOLE SET: send every row of one kind in ONE call, and calling again for the same kind REPLACES the entire staged set rather than adding to it — a trial balance only balances as a whole, so a partial patch would produce a set that ties to no report. Found a mistake? Send the corrected set in full. Taokeh re-derives every total from the rows you sent and shows it beside the total you say was printed on the report, so the owner can see the set tie to the cent; send that printed total in sourceTotal, and never adjust a row to make a total work — report the gap instead. THESE ARE THE CUTOVER, so unlike stage_document there is no "too old" fence: an aged invoice dated before the accounting start date is exactly right, and that is what makes it an opening document. REFUSED in two cases, both permanent until the owner acts: the books are LOCKED, or that kind's opening balances are ALREADY POSTED (staging over them would double-count — only the owner can undo them). A CREDIT balance on an aged list (an overpayment, an unapplied credit note) is not an opening invoice or bill: Taokeh will not enter it, so leave those rows out and TELL THE OWNER they exist, naming the party and the amount, so they can enter them by hand. ⛔ ROW KEYS ARE STRICT (MD, 2026-09-09): a key the schema does not list is REFUSED BY NAME — with the key it probably meant, e.g. sell_price → unitPrice — and NOTHING is staged. Send the contract's apiField, not the spreadsheet header. Unknown keys used to be dropped in silence, so a set could stage "successfully" with its prices missing; that is the bug this refusal closes.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | 'trial_balance' = the closing balances of every account in the old system. 'aged_receivables' = the invoices customers still owed at the cutover, one row each. 'aged_payables' = the bills still owed to suppliers, one row each. | |
| rows | Yes | EVERY row of the set, in the order the report prints them. This replaces any previously staged set for this kind. | |
| sourceName | No | What you read — the export's file name or the report's own title, e.g. 'Trial Balance as at 31-12-2025 (Financio)'. Shown to the owner as the evidence for what they are approving. | |
| sourceTotal | No | The grand total PRINTED on the report (for a trial balance, the DEBIT total). Advisory: Taokeh re-derives its own figure and shows the two side by side. This is the single most useful check in the whole migration — send it whenever the report states one. | |
| sourceTotalCredit | No | trial_balance only: the CREDIT total printed on the report, if it is stated separately. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations all false (no readOnly, no idempotent, no destructive hint), the description carries full disclosure responsibility. It does so thoroughly: states it posts NOTHING, explains the replace semantics (calling again replaces the entire set), details denial conditions (locked books or already posted), and warns that unknown row keys are refused by name and nothing is staged. It even references the historical silent-drop bug this refusal fixes. This goes well beyond what the annotations reveal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place; it is dense with critical operational rules rather than filler. It is front-loaded with the core purpose and the intake dependency, then covers whole-set semantics, exceptions, and strictness. While it could be shortened slightly, the length is justified by the tool's complexity. It is well-organized and not tautological.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex staging tool with no output schema and minimal annotations, the description leaves nothing critical out. It covers preconditions (intake_contract), the whole-set requirement, the cutover nature, refusal cases, handling of credit balances, strict row-key validation, and the sourceTotal reconciliation check. An agent can invoke the tool correctly without guessing about any essential behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already has detailed parameter descriptions. The description adds extra meaning: for sourceTotal it calls it 'the single most useful check' and instructs sending it whenever the report states one; it clarifies that outstanding is what gets booked when both amount and outstanding are present; and it instructs never adjusting a row to make a total work. This is additive context beyond the schema, so a score above the baseline of 3 is warranted, but not a 5 because the schema already covers the core semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise statement of what the tool does: bring a company's opening balances (trial balance, aged receivables, aged payables) across from the old accounting system. It distinguishes itself from stage_document ('THESE ARE THE CUTOVER, so unlike stage_document...') and from intake_contract (which it instructs to call first). The verb ('Bring... across') and resource (the three kinds of opening balances) are explicit and unique.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage instructions: call intake_contract first with the right doc_type, send the whole set in one call, and replace on re-call. It contrasts with stage_document on the 'too old' fence, and lists two permanent refusal conditions. It also tells the agent exactly what to do with credit balances (leave them out and tell the owner) and how to handle mistakes (send corrected set in full). This is comprehensive when-to and when-not-to guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_levelStock levelsARead-onlyInspect
How many units of a product you have on hand right now — search by product name or SKU. An item marked "Service — no stock" comes back with isService:true and onHand:null — it holds no stock, so never report a quantity for it. If this company uses Counter (the till), the figure is as at the last day-close: counter sales move stock once, when the day is closed, not at each scan.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover read-only/no-open-world. The description adds real behavioral payload: service items return isService:true and onHand:null and must never be reported as a quantity, and Counter companies only reflect sales at day-close rather than per scan. These are non-obvious semantics an agent could easily get wrong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core answer, then two dense caveat sentences. Every sentence carries distinct information; the Counter explanation is niche but genuinely affects interpretation of the number.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining the return shape, and it does partially (isService, onHand). It does not describe the full field set or what happens on no-match, but the critical edge cases are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'query' param is undocumented in the schema, but the description compensates by stating it accepts a product name or SKU, which is exactly what an agent needs to form the input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (units on hand), a precise time scope ('right now'), and the retrieval key ('product name or SKU'). This cleanly separates it from sibling tools like stock_movements (history) and low_stock (threshold list) without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains what the returned figure means and when a quantity is not applicable (service items, Counter day-close timing), which gives contextual usage. However, it never says when to prefer this over stock_movements or low_stock, nor any prerequisites, so routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_movementsStock movementsARead-onlyInspect
Read the STOCK MOVEMENT LEDGER in bulk — every recorded change to on-hand quantity, across products, in one call. stock_level says how many you have RIGHT NOW for one product; this says WHAT MOVED, WHEN, WHY and WHAT IT LEFT BEHIND. Use it for stock-turn analysis, shrinkage hunting, reorder timing, or reconstructing how a balance got where it is. Every filter is optional and a bare call is a legitimate "what moved lately" pull: narrow with product (case-insensitive contains against the product name OR SKU), from/to (a window on each row's docDate) and sourceType (exactly one cause — see the list). Each row: ts (when it was posted), docDate (the document's own date, YYYY-MM-DD — a voided or edited bill's reversal keeps the bill's date, while a voided credit or debit note's reversal is dated the day of the void, like its journal entry; where no document survives, the Malaysian day it was posted), product {sku, name}, qtyDelta (signed — negative took stock out), balanceAfter (on-hand immediately after that move, as the posting path recorded it), unitCost, reason (free text the person typed, where there was one), sourceType and sourceId (the id of the document that moved it — pair it with search_documents to see which one). ⚠ WHICH MOVEMENTS EXIST AT ALL DEPENDS ON THIS COMPANY'S STOCK MODE, and the answer says which mode it is in (stockMode) with a note. In modified_periodic — the mode of companies created before 28 Sep 2026 unless they switched; newer companies start perpetual — selling does NOT move stock: invoices, delivery orders and credit notes write no movement row, and stock is trued up at stock take. Seeing no 'sale' rows there means the company is periodic; it does NOT mean nothing was sold, and it is NOT shrinkage. Only a perpetual company ORIGINATES sale / credit_note rows — but a periodic company can still HOLD them, and can still gain new ones: rows written while it ran perpetual stay, and a VOID takes its truth from the document's own movement rows rather than from today's setting, so voiding a perpetually-booked invoice or supplier return writes a fresh sale_void / debit_note_void row in a company that is periodic now. So a periodic company with movement rows is not a contradiction and not a bug. Read the mode before you interpret the rows, and read a row's own sourceType before you attribute it to the mode. balanceAfter is the total across the whole company, not per location; on a multi-location company each row also carries location. Rows come back newest first, capped at 200 with total, shown and more — when more is true, narrow by date and pull the periods in turn rather than treating a partial page as the whole. Values are verbatim as recorded at the time, never re-derived. ⛔ WHAT IT WILL NOT DO: it does not value your inventory (balance_sheet does), does not compute COGS or margin (income_statement, profit_drivers), and does not tell you what to reorder (low_stock). Nothing is written, and no draft is created.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Latest docDate (YYYY-MM-DD; the document date, else the Malaysian posting day). | |
| sku | No | NOT a filter on this tool. Pass it as `product` — one case-insensitive contains matched against BOTH the product name and the SKU. | |
| date | No | NOT a filter on this tool. Pass a date WINDOW as `from` and/or `to` (YYYY-MM-DD). A single date is `from` and `to` set to the same day. | |
| from | No | Earliest docDate (YYYY-MM-DD; the document date, else the Malaysian posting day). The main way to break a >200-row pull into honest slices. | |
| item | No | NOT a filter on this tool. Pass it as `product` — one case-insensitive contains matched against BOTH the product name and the SKU. | |
| name | No | NOT a filter on this tool. Pass it as `product` — one case-insensitive contains matched against BOTH the product name and the SKU. | |
| type | No | NOT a filter on this tool. Pass the movement cause as `sourceType`. | |
| limit | No | NOT a filter on this tool. The page size is fixed at 200 movements. Narrow with from/to, product or sourceType and pull the periods in turn. | |
| query | No | NOT a filter on this tool. Pass it as `product` — one case-insensitive contains matched against BOTH the product name and the SKU. | |
| dateTo | No | NOT a filter on this tool. Pass the end of the window as `to`. | |
| reason | No | NOT a filter on this tool. `reason` is free text the person typed and is RETURNED on each row, not a filter. To narrow by cause use `sourceType` (e.g. 'adjustment'). | |
| source | No | NOT a filter on this tool. Pass the movement cause as `sourceType`. | |
| product | No | Case-insensitive contains matched against the product NAME or SKU. Omit for every product. | |
| dateFrom | No | NOT a filter on this tool. Pass the start of the window as `from`. | |
| productId | No | NOT a filter on this tool. This search matches product TEXT, not ids — pass the name or SKU as `product`. | |
| sourceType | No | Exactly one movement cause. One of: purchase, purchase_void, debit_note, debit_note_void, adjustment, stock_take, import_opening, assembly_out, assembly_in, disassembly_out, disassembly_in, sale, sale_void, credit_note, transfer_out, transfer_in. Note that sale, sale_void, credit_note are only ever CREATED by a perpetual company — a company that has since switched to periodic still holds the ones it wrote (and a void of one of those documents still writes its reversal leg). | |
| productName | No | NOT a filter on this tool. Pass it as `product` — one case-insensitive contains matched against BOTH the product name and the SKU. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover readOnlyHint and openWorldHint; the description carries a heavy load beyond that: the company stockMode caveat (periodic companies do not originate sale rows, absent sale rows are not shrinkage), which sourceTypes can exist per mode, that balanceAfter is company-wide and not per location, that rows are verbatim never re-derived, and the 200-row hard cap with total/shown/more pagination. It also confirms nothing is written and no draft is created.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the single most decision-relevant distinction (ledger vs current level) and organised with signal markers (⚠, ⛔). It is long, and the docDate-reversal sub-clause is dense enough to slow a reader, but in a domain with this many mode-dependent edge cases most sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the full row shape (ts, docDate, product, qtyDelta signed, balanceAfter, unitCost, reason, sourceType, sourceId, location) plus the pagination envelope. Given 17 parameters and the mode-dependent semantics, nothing an agent needs to call and correctly interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds real meaning: `product` is a case-insensitive contains against name OR SKU, `from`/`to` are a window on each row's docDate, and the docDate semantics (a voided bill's reversal keeps the bill date, a voided credit/debit note's reversal is dated the day of the void). It also reinforces that reason is returned not filtered and that limit is fixed. The sourceType enum list is repeated from the schema rather than extended.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource (read the stock movement ledger in bulk) and immediately contrasts with the sibling stock_level ('how many you have RIGHT NOW for one product' vs 'WHAT MOVED, WHEN, WHY'). The scope is unambiguous and no sibling could be confused with it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names concrete use cases (stock-turn analysis, shrinkage hunting, reorder timing, reconstructing a balance) and explicitly routes elsewhere for adjacent needs via the ⛔ block: valuation to balance_sheet, COGS/margin to income_statement and profit_drivers, reorder to low_stock. It also legitimises the bare no-filter call, so the agent knows when an unfiltered pull is acceptable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tax_positionTax positionARead-onlyInspect
Your current tax posture in one read: (1) SST — the current bi-monthly period's SST payable if you're SST-registered (registration 'none' ⇒ not registered, nothing to remit); and (2) e-invoice consolidation — whether monthly consolidation is on, the open month, last month's filing due date, and whether last month's consolidated document is generated / LHDN-validated. Basis: SST-02 return figures + the consolidated-e-invoice register.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, and the description reinforces this with 'one read' and 'Basis: SST-02 return figures + the consolidated-e-invoice register.' It adds meaningful conditional behavior about unregistered users and the e-invoice consolidation statuses, going beyond what the annotation alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a front-loaded summary sentence and numbered components. Every clause adds useful information, and the parenthetical clarifications are compact rather than redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only status tool with no output schema, the description is complete: it names both reporting areas, the source basis, the conditional SST behavior, the open period, due date, and validation status. An agent has enough to know what this tool returns and when the data is meaningful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema coverage is effectively 100%, so there is no parameter-meaning burden for the description to carry. The baseline for a no-parameter tool is 4, and the description does not need to add anything further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and target ('Your current tax posture in one read') and then enumerates the two specific components covered: SST payable and e-invoice consolidation. This makes the tool's scope unmistakable and distinct from the many summary siblings like cash_position or business_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies this is the tool for a consolidated tax-status check, and it clarifies the conditional handling for SST registration ('registration 'none' ⇒ not registered, nothing to remit'). However, it does not explicitly name alternative tools or state when not to use it, so it misses the explicit exclusion guidance that would earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_bill_draftFile a pending correction draft for a posted bill (admin approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a supplier bill ALREADY POSTED in Taokeh — the door for a bill that went onto the books with a wrong line, a wrong quantity or cost, the wrong date, the wrong payment term or the wrong payment method. It files a pending DRAFT an ADMIN reviews as a DIFF (current → proposed) and approves; only that tap rewrites the bill, and it keeps the SAME bill number, the original attached document, and the link to the purchase order the bill was converted from. Pass purchaseId, the bill's REAL id from search_documents — a bill number is refused, because two suppliers can print the same number and correcting the wrong one is worse than asking. ROUTING — three different mistakes, three different doors. (1) THE PURCHASE NEVER HAPPENED (keyed twice, entered on the wrong company, the wrong supplier's bill): the owner VOIDS it — Purchasing → the bill → Void, in Taokeh. There is deliberately NO AI lane for voiding. Tell them where to click. (2) THE PURCHASE HAPPENED BUT THE BILL IS WRONG (a wrong line, quantity or unit cost, the wrong date, term, due date or payment method): THIS tool. (3) SOMETHING CHANGED AFTER THE PURCHASE (goods went back to the supplier, the supplier credited a price difference later): a DEBIT NOTE against the bill, which the owner raises in Taokeh (Purchasing → Debit notes) — a real event that belongs on the books as its own document. WHAT YOU CAN CHANGE: purchaseDate, paymentMethod ('CASH' paid on the spot, or 'CREDIT' on account), term (the supplier's payment term, stored as written — changing an on-account bill to one of the five presets, Due on Receipt / Net 7 / Net 14 / Net 30 / Net 60, re-derives the due date from the bill date unless you also send dueDate), dueDate (only when the due date ITSELF is wrong — correcting just the bill date already moves the due date by the same number of days), and lines — the FULL REPLACEMENT SET, so send every line the bill should have, including the ones that were already right. Everything you do NOT send is carried across exactly as the bill has it: the supplier, the currency and its frozen exchange rate, the location the goods were received into, the batch (lot) each line was received into, and each line's SST code. A replacement line of a product the bill already carries is received into the same batch as the original line of that product. WHAT YOU CANNOT, each refused by name: the SUPPLIER (a different supplier is a different document — the owner voids and re-enters); the BILL NUMBER (preserved across an edit; Taokeh's own Edit screen cannot change it either); the CURRENCY or the RATE (frozen at posting); and a bill posted with a whole-document tax rate instead of per-line SST codes (Taokeh does not store that rate, so a correction from here would drop the input tax — the owner corrects it on the app's Edit screen). WHAT BLOCKS AN EDIT OUTRIGHT, checked when you file and again when the admin taps, each refusal naming the real next step: the document is not a bill (an opening balance or a debit note has its own page); it is validated on MyInvois as a self-billed e-invoice; a live debit note has been raised against it; a payment has been allocated to it; or it is dated inside a period the company has locked. A proposal that changes nothing is refused rather than filed as an empty diff. If the bill is edited by someone else after you read it, the admin is shown BOTH versions and the one-tap approval is refused — that is by design; call revise_draft (kind 'bill_update') to re-read and re-propose. An edit in Taokeh re-books the bill under a NEW id with the same number: a pending draft FOLLOWS it, and calling this tool with the old id is refused with the new id named — re-read the bill before proposing against it.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note in the reviewer's language: what was wrong, what you changed, and anything they should double-check before they approve a rewrite of a posted bill. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name; a proposal the bill already agrees with is refused rather than filed. | |
| purchaseId | Yes | REQUIRED — the posted bill's real id, from search_documents. A bill number is refused: correcting the wrong document silently is worse than asking which one they mean. | |
| needsReview | No | Set true when something gave you pause — it flags the draft for the reviewer. | |
| printedTotal | No | The grand total the corrected bill should show, if the user told you one — cross-checked against the server's own derivation and flagged on the review screen. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say write, non-destructive, non-idempotent; the description goes far beyond that, disclosing that nothing changes until a human approves a diff, what fields are carried across unchanged, what is refused by name, the concurrent-edit dual-version behavior, and that a re-booking gives the bill a new id that the draft follows. This is exactly the behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical constraint is front-loaded, which is good, but the body is a very dense wall of text with heavy all-caps emphasis and some repetition. The complexity justifies length, yet it could be tightened substantially without losing edge-case coverage, so it is only adequate on conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only lightweight annotations, the description carries the full load and does so completely: purpose, routing, permitted and forbidden changes, blocking conditions, conflict handling, and recovery steps are all covered. An agent has everything needed to invoke this correctly in a complex mutation workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents each parameter, but the description adds material semantics the schema does not: unsent fields are carried across exactly, lines must be the full replacement set, replacement lines inherit the original batch, and purchaseId must be the real id rather than a bill number. It stops short of adding format or value guidance beyond what the schema already supplies, so it is a 4 rather than a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a specific verb and resource — it files a pending correction draft for a supplier bill already posted in Taokeh — and immediately scopes it against alternatives (voiding, debit notes, create_bill_draft, revise_draft). An agent can tell exactly what this tool does and where it sits among siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The ROUTING section explicitly maps three distinct mistakes to three distinct paths (void in Taokeh, this tool, debit note), names the case with no AI lane, and lists the blockers that reject an edit outright. It also names revise_draft as the recovery path after a concurrent edit, so when-to-use and when-not-to-use are both stated rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_contact_draftFile a pending contact-change draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose changes to a customer or vendor ALREADY in Taokeh — the correction door for an address book with missing tax ids, stale phone numbers or wrong e-Invoice flags (e.g. after a migration). This does NOT change anything: it files a pending DRAFT the owner reviews as a DIFF (current → proposed, changed fields only) and approves with one tap; only that tap writes. Pass partyKind ('customer' or 'vendor'), the contact's REAL contactId from resolve_customer / resolve_vendor — a name is refused, because a near-match would edit the wrong company — and proposed, an object holding ONLY the fields you want changed. Settable fields, mirroring the contact screen exactly: name, email, phone, address, tin (Tax ID), brn (registration no.), idType ('BRN' or 'NRIC'); on a CUSTOMER also requiresEinvoice, einvCity, einvPostcode, einvState (LHDN 2-digit code), remindersMuted; on a VENDOR also selfBill, alwaysExpense, state. Pass null on a text field to CLEAR it. Anything else is refused by name — do not invent fields. You cannot move a contact between customer and vendor here; they are separate records. A field the contact already agrees with is dropped, and a proposal that changes nothing is refused rather than filed as an empty diff for a human to tap. One contact per call — loop for a sweep, because 100 reviewable diffs is the point: the owner sees each change before it lands. BE HONEST: never guess a tax id or a phone digit; if you are unsure, leave the field out and say so in notes with needsReview. If the record has been edited since you read it, the owner is shown BOTH values and the one-tap approval is refused — that is by design; call revise_draft (kind 'contact_update') to re-read and re-propose.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note, in the reviewer's language: what you changed and why, and anything they should double-check. Leave empty when there is nothing to flag. | |
| persons | No | Changes to the PEOPLE filed under this contact (the individuals you deal with there). Each entry is one change. A removal names only personId. A person needs a name to be added. These ride the SAME draft and the SAME single approval as the field changes — the owner sees a removal shown as a removal, in words. | |
| proposed | No | ONLY the fields you want changed. A field the contact already agrees with is dropped; an unknown or wrong-side field is refused by name. Omit it entirely when you are only changing the PEOPLE below. | |
| contactId | Yes | REQUIRED — the contact's real id from resolve_customer / resolve_vendor. A name is refused: editing the wrong company silently is worse than asking. | |
| partyKind | Yes | Which side of the book the contact is on: 'customer' (someone you sell to) or 'vendor' (someone you buy from). They are separate records — this tool cannot move one to the other. | |
| needsReview | No | Set true when something gave you pause — a tax id you read off a blurry document, a phone number the user was unsure about. It flags the draft for the reviewer. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only convey readOnly=false, openWorld=false, idempotent=false, destructive=false, so the description carries the behavioral burden — and it exceeds it. It discloses the draft-only semantics ('nothing changes until a human approves'), the diff review flow, rejection behaviors (empty proposals refused, agreeing fields dropped, unknown fields refused by name), null-clears-text semantics, one-contact-per-call, and the concurrency design where simultaneous edits force re-proposal. Nothing in the description contradicts the annotations; in fact, the draft-only framing explains why destructiveHint=false is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (~350 words) and repeats the core draft-only message three times (title, opening sentence, third sentence). However, it is front-loaded with the most decision-critical fact and the remaining length is dense with safety-critical semantics — particularly the 'never guess a tax id' honesty rule and the concurrency behavior. A small trim of the redundant draft-only restatements would make it exemplary, but every substantive paragraph earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a high-complexity tool — 6 parameters, two nested objects (proposed, persons), conditional customer/vendor fields, concurrency edge cases, no output schema — and the description still leaves nothing essential unexplained. It covers field eligibility rules, the people sub-object, failure modes (name refused, empty diff refused, stale-record refusal), looping guidance for sweeps, and honesty requirements. No output schema exists, so some return-value details are absent, but for a tool whose semantics are this intricate the description is remarkably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value on top: 'Pass null on a text field to CLEAR it', 'Anything else is refused by name — do not invent fields', the consolidated customer-vs-vendor field grouping ('on a CUSTOMER also requires...'), and 'Omit it entirely when you are only changing the PEOPLE below'. These operational meanings are not fully inferable from the schema's property descriptions alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair: 'Propose changes to a customer or vendor ALREADY in Taokeh' and 'FILES A PENDING DRAFT ONLY'. It clearly distinguishes itself from the create sibling by scoping to existing contacts ('the correction door'), and the title adds the human-approval workflow. An agent can immediately tell this is the update/amend tool, not the creation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use context (missing tax ids, stale phone numbers, wrong e-Invoice flags after a migration) and explicit exclusions: 'You cannot move a contact between customer and vendor here' and 'A name is refused'. It also routes to alternatives: resolve_customer/resolve_vendor for the real ID and 'call revise_draft (kind 'contact_update')' for the concurrency case. This is actionable routing guidance, not just a purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_credit_note_draftFile a pending correction draft for a posted credit note (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a SALES CREDIT NOTE already posted in Taokeh — the door for a credit note that went onto the books with a wrong line, quantity, price, date or reason. An ADMIN reviews it as a DIFF (current → proposed) and approves; only that tap re-books the note, and it keeps the SAME credit-note number, the invoice it credits and any attached document. Pass creditNoteId, the note's REAL id from search_documents (docType credit_note) — a credit-note number is refused. ROUTING — three different mistakes, three different doors. (1) THE CREDIT NEVER SHOULD HAVE EXISTED (keyed twice, wrong customer, raised against the wrong invoice): the owner VOIDS it — Credit notes → the note → Void, in Taokeh; voiding keeps the note on the register marked VOID, and there is deliberately no AI lane for it. (2) THE CREDIT IS RIGHT BUT TYPED WRONG: THIS tool. (3) SOMETHING NEW HAPPENED LATER (more goods came back, a further allowance was agreed): a NEW credit note with create_credit_note_draft — a later event is its own document, not a rewrite of this one. WHAT YOU CAN CHANGE: cnDate, reason, and lines — the FULL REPLACEMENT SET (send every line the note should have, including the ones that were already right). ⛔ EVERY LINE MUST ANSWER goodsReturned (true = these goods physically came back into stock, false = money only). There is no default. Your answer only PRE-FILLS the admin's per-line question on the full review page — the admin's answer is what posts — and a ONE-TAP approval never moves stock: it is refused whenever the note already put goods back or any line says goods came back, so those corrections are approved on the review page. Leave taxCode out and the line inherits the note's own SST code for that product; a NEW product on a taxed note has nothing to inherit and is refused by name. WHAT YOU CANNOT, each refused by name: the CUSTOMER, the credit-note NUMBER, and the INVOICE it credits. THINGS THAT BLOCK AN EDIT OUTRIGHT, checked when you file (the server runs the real correction and rolls it back, so every refusal approval would give you is given NOW) and again when the admin taps: the note is voided; it is validated on MyInvois, exported for MyInvois, or submitted to LHDN (the owner voids it instead — the void clears the export and says what to do at MyInvois); a customer refund is allocated to it; the invoice it credits is gone or is not an invoice; its date is inside a LOCKED period; it put goods back under the company's previous stock setting; or the corrected note would credit more than the invoice has left, or put back more than the invoice shipped. A proposal that changes nothing is refused rather than filed. If the note is edited by someone else after you read it, the one-tap approval is refused — call revise_draft (kind 'credit_note_update') to re-read and re-propose.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note in the reviewer's language: what was wrong, what you changed, and — above all — whether the goods really came back. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name. | |
| needsReview | No | Set true when something gave you pause. | |
| creditNoteId | Yes | REQUIRED — the posted credit note's real id, from search_documents (docType credit_note). A credit-note number is refused. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false/absent, so the description carries the full burden — and it delivers: draft semantics ('NOTHING CHANGES UNTIL A HUMAN REVIEWS'), pre-fill-not-posting behavior, one-tap approval refused when goods-returned is involved, same number/invoice/doc preservation, validation-by-rollback at file time, full blocker list, and concurrency refusal. This is precisely the behavioral context an agent needs and would otherwise discover only through failed calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but earns its length: routing, blockers, and refusal conditions are all actionable and non-redundant with the schema, and the scannable ALL-CAPS headers plus numbered list structure tames the density. The most critical fact — 'NOTHING CHANGES UNTIL A HUMAN REVIEWS' — is front-loaded. Minor redundancy exists ('admin's answer is what posts' restates the opening), so not a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with nested objects, a human approval gate, and no output schema, this is near-complete: it covers routing, blockers checked twice, concurrency, inheritance rules, and edge cases (no-change proposals refused, new product on taxed note). The only genuine gap is that it never states what the tool returns on successful filing, which matters more because there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description adds real value beyond the schema: lines as a FULL REPLACEMENT SET, forbidden keys (customer, number, invoice), taxCode inheritance behavior, and the no-default on goodsReturned. Deducted one point because much of the per-parameter detail (patterns, positivity, 'omit lines' note) already lives in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'FILES A PENDING DRAFT ONLY' for a 'SALES CREDIT NOTE already posted in Taokeh', scoped to corrections for 'wrong line, quantity, price, date or reason'. Clearly distinguishes itself from the sibling create_credit_note_draft by the posted-vs-new-note boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The ROUTING section is exemplary: it enumerates three mistake types and maps each to a different door — void in Taokeh, this tool, or create_credit_note_draft — with explicit conditions ('THE CREDIT IS RIGHT BUT TYPED WRONG: THIS tool'). It also points to revise_draft for the concurrent-edit recovery path. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_delivery_order_draftFile a pending delivery-order correction draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a DELIVERY ORDER that already exists and has not been invoiced — a wrong quantity or price, a missing or extra line, the wrong date or notes. An admin or bookkeeper reviews a plain current → proposed DIFF and approves. A delivery order posts nothing and moves NO stock (stock moves when it is invoiced), so the correction touches no ledger and no stock. ⚠ Approving re-captures each line's cost from the products' CURRENT average cost — exactly what Taokeh's own Edit screen does; it is a dispatch-time reference that posts nothing. ROUTING — (1) the delivery never happened: the owner voids the DO in Taokeh (Delivery orders → the DO → Void); there is no AI lane for voiding. (2) the DO is wrong and still OPEN: THIS tool. (3) the DO has already been INVOICED: it is closed to edits — correct the invoice instead (update_invoice_draft). WHAT YOU CAN CHANGE: doDate, notes (null clears them) and lines. WHAT YOU CANNOT, each refused by name: the CUSTOMER, the DO NUMBER, the CURRENCY or exchange rate, and invoicing or voiding. ONLY AN OPEN DELIVERY ORDER can be corrected — checked when you file and again when the reviewer taps. Leave a line's taxCode out and it inherits this DO's own SST code for its product — unless the DO carries that product under more than one code, when you must state it. WHAT HAPPENS ON APPROVAL: the document is updated IN PLACE — same id, same number — through the SAME update Taokeh's own Edit screen runs. Every field you leave out is carried across exactly as it stands (the party, the attention contact, the currency and its frozen exchange rate, the tax) and is never reset. LINES: lines is the FULL replacement set. READ THE LINES FIRST with search_document_lines — each row's lineNo is what a { keep: } entry refers to, and keep is how a line that is already right survives untouched (a measured line keeps its frozen unit and its rounding row). A changed or new line is written in full, in the create tool's line shape, and the server re-derives its quantity and amount. A proposal that changes nothing is refused rather than filed as an empty diff. STALE: if a person edits the document after you read it, the one-tap approval is refused and the reviewer is shown both versions — call revise_draft with this draft's kind to re-read and re-propose. Pass the document's REAL id from search_documents; a document number is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note in the reviewer's language: what was wrong, what you changed, and anything they should check before approving. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name; a proposal the delivery order already agrees with is refused rather than filed. | |
| needsReview | No | Set true when something gave you pause — a quantity you inferred, a price the user was unsure about. It flags the draft for the reviewer. | |
| printedTotal | No | The grand total the corrected delivery order should show, if the user told you one — cross-checked against the server's own figure and flagged on the review screen. | |
| deliveryOrderId | Yes | REQUIRED — the delivery order's real id, from search_documents. A DO number is refused. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false and carry little safety information, so the description carries the full burden. It discloses that nothing changes until human approval, no stock/ledger movement, in-place update, omitted fields carried over, average-cost recapture on approval, and refusal of empty diffs and stale proposals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but well-structured and front-loaded with the most critical fact that nothing changes until human approval. Some redundancy exists, such as repeated 'refused by name' and restating the in-place/same-update point, but the density is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description covers when to use the tool, what can and cannot change, line-level mechanics, stale-draft handling, and post-approval behavior. It is complete enough for an agent to invoke this complex tool correctly without external context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, but the description adds substantial meaning beyond it: deliveryOrderId must be a real id not a number, notes null clears, lines are a full replacement set, keep semantics refer to lineNo, taxCode inheritance behavior, and server re-derivation of quantities and amounts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: file a pending correction draft for an existing, not-yet-invoiced delivery order. It clearly differentiates this from voiding in Taokeh and from update_invoice_draft for already-invoiced orders.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit routing rules: if the DO never happened, void in Taokeh; if wrong and open, use this tool; if invoiced, correct the invoice instead. It also instructs reading lines first with search_document_lines and using revise_draft when stale.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_employee_draftFile a pending employee-change draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose corrections to an employee ALREADY in Taokeh's payroll — a wrong IC or passport number, a missing EPF/SOCSO/income-tax number, a salary keyed a digit out, a stale designation, the statutory applicability flags. This does NOT change anything: it files a pending DRAFT an admin (or a Bookkeeper + payroll login) reviews as a DIFF (current → proposed, changed fields only) and approves with one tap; only that tap writes. Pass the employee's REAL employeeId — from payroll_summary with staff: true (the whole roster, works even before the first payroll run) or perEmployee: true — and proposed, an object holding ONLY the fields you want changed. A name is refused: two people can share one, and editing the wrong person's salary silently is much worse than asking. ⚠ FORWARD ONLY, and you must tell the user this: an approved change applies to FUTURE payroll runs. It never re-opens, recomputes or restates a payslip or payroll run that has already been filed — including the year-to-date PCB, which is summed from payslips already submitted. If a filed month is wrong, that is a human decision taken in Payroll → Payroll Runs, not something this tool can do. Settable fields, mirroring the employee form exactly: staffNo, name, icNo, passportNo, countryCode, dateOfBirth, designation, joinDate, resignDate, status ('ACTIVE' or 'RESIGNED'), email, isMalaysian, basicSalary, epfEmployeeRate, bankName, bankAccountNo, epfNo, socsoNo, incomeTaxNo, epfApplicable, socsoApplicable, eisApplicable, pcbApplicable, skbbkEnrolled, maritalStatus ('SINGLE', 'MARRIED' or 'SINGLE_PARENT'), spouseWorking, numChildren, taxResident, residencyChangeMonth ('YYYY-MM'), zakatMonthly, employmentStatus (CP8D code '1'–'6'), contractEndDate, holidayState, notes. NOT settable here, each refused by name with the screen that owns it: the TP3 prior-employer figures (a signed declaration, and requesting one EMAILS the employee), the TP2 benefits-in-kind values (owned by the approved election), and the flat-15% tax regime (approval-conditional — an AI cannot verify it, and halving somebody's tax is not a diff-door change). bankName and bankAccountNo ARE settable because they print on the payslip and on the giro list the owner uploads at their own bank — TAOKEH NEVER PAYS ANYONE; the owner authorises every payment at their own bank, and nothing on this connector can move money. A field the record already agrees with is dropped, and a proposal that changes nothing is refused rather than filed as an empty diff. One employee per call — loop for a sweep, because each reviewable diff is the point. BE HONEST: never guess an IC, a bank account or a salary; leave the field out and say so in notes with needsReview. ADMIN AND BOOKKEEPER + PAYROLL ONLY: payroll sits behind its own role and this tool refuses on any other connection (a plain Bookkeeper included).
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note, in the reviewer's language: what you changed, why, and what they should double-check. | |
| proposed | Yes | ONLY the fields you want changed. A field the record already agrees with is dropped; an unknown or deliberately-excluded field is refused by name. | |
| employeeId | Yes | REQUIRED — the employee's real id, from payroll_summary (`staff: true` for the roster, or `perEmployee: true` for a run). A name is refused. | |
| needsReview | No | Set true when something gave you pause — an IC read off a blurry photo, a salary the user was unsure about. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (only readOnlyHint=false, etc.), so the description carries the full burden. It discloses the non-destructive nature ('does NOT change anything'), the forward-only behavior ('applies to FUTURE payroll runs. It never re-opens, recomputes or restates a payslip or payroll run'), the auto-drop of agreeing fields, refusal of empty diffs, refusal of a name as identifier, and the 'Taokeh never pays anyone' constraint. It also explains the approval flow as a diff review. This is rich behavioral context well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but it is front-loaded with the most critical fact ('FILES A PENDING DRAFT ONLY') and structured with clear sections: the draft behavior, the identifier requirement, the forward-only warning, the settable fields list, the non-settable exclusions, and the role gate. Every sentence earns its place given the high stakes (money and statutory data). It is verbose but well-organised; a slightly tighter version could exist, but the detail is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 4 parameters (including a nested object with ~30 fields), no output schema, and high stakes, the description covers everything an agent needs: the identifier source, the allowed fields, the disallowed fields with reasons, the approval workflow, the forward-only consequence, the money-safety guarantee, the role restrictions, and the empty-diff handling. It also explains how to loop for a sweep. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% description coverage, including detailed per-field explanations (e.g., dateOfBirth drives EPF/SOCSO age bands). The description adds usage-level semantics not in the schema: how to obtain employeeId ('from payroll_summary with staff: true or perEmployee: true'), the rule that only changed fields go in 'proposed' ('A field the record already agrees with is dropped'), and the convention for 'needsReview' and 'notes'. These are valuable beyond the schema, but the schema does most of the heavy lifting, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, specific statement: 'FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH.' It names the exact resource (employee in payroll) and the action (propose corrections as a draft). It explicitly distinguishes from siblings by listing what it is not (TP3, TP2, flat-15% tax regime) and by naming the alternative screens. The verb 'file' and the 'pending draft' concept are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use: 'Propose corrections to an employee ALREADY in Taokeh's payroll' with examples of valid corrections. It also states when not to use it: 'NOT settable here, each refused by name with the screen that owns it' for TP3, TP2, and flat-15% tax regime. It names the alternative for a wrong filed month: 'a human decision taken in Payroll → Payroll Runs.' It also gives role prerequisites: 'ADMIN AND BOOKKEEPER + PAYROLL ONLY.' No inference needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_expense_draftFile a pending correction draft for a posted paid expense (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a PAID EXPENSE already posted in Taokeh (an entry on the Expenses page — find its id with search_expenses). The owner or bookkeeper reviews it as a DIFF (current → proposed) and approves; only that tap re-books the entry, exactly as Expenses → Edit does, and the receipt already on it carries across. ROUTING — (1) THE EXPENSE NEVER HAPPENED (keyed twice, the wrong company's receipt): the owner DELETES it in Taokeh (Expenses → the expense → Delete); there is deliberately no AI lane for that. (2) IT HAPPENED BUT WAS KEYED WRONG (the date, the amount, the input SST, the category, the account it was paid from, the payee, the reference): THIS tool. (3) SOMETHING NEW HAPPENED LATER (a refund came back, a second payment was made): a new entry — create_expense_draft for a further payment, or a reversing journal via create_journal_draft for money that came back. WHAT YOU CAN CHANGE, sending ONLY what is wrong (everything you leave out is carried from the entry as it stands): date, amount (the gross ringgit paid), inputTax (the input SST inside that amount), debitAccountCode (the category) and creditAccountCode (paid from) — both from expense_accounts —, memo and reference. Ringgit only; the currency and the receipt are not changeable here. BLOCKED OUTRIGHT, checked when you file (the server runs the real correction and rolls it back) and again at approval: the entry is not a paid expense; its date is inside a LOCKED period (or your new date is); a category that is not a cost account; a paid-from account that cannot pay; or an entry split across more lines than the expense form can show. A proposal that changes nothing is refused. If the expense is edited after you read it, the one-tap approval is refused — call revise_draft (kind 'expense_update').
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note: what was wrong and where the corrected figure came from. | |
| entryId | Yes | REQUIRED — the paid expense's journal-entry id, from search_expenses. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name. | |
| needsReview | No | Set true when something gave you pause. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are sparse (only false hints), so the description carries the full burden, and it delivers: it discloses the pending/approval workflow, the diff review, that the receipt carries across, server-side rollback checks, hard blocked conditions, and the one-tap approval refusal on stale edits. This is far beyond what annotations provide and materially changes how an agent should reason about the call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, with all-caps emphasis that makes scanning harder than necessary. However, every section earns its place: the routing rules, patchable fields, and blocked conditions are all operationally critical and are organized under clear headers, with the central pending-draft behavior front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, nested-object mutation tool with no output schema, the description is complete enough to invoke correctly: it explains how to find the entry, what can be changed, what cannot, what will be rejected, and how to recover if the draft becomes stale. No critical operational context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds meaning the schema alone cannot convey: send only what is wrong, omitted fields carry forward, ringgit-only amounts, currency/receipt are immutable, null clears values, and unknown/patchable keys are refused by name. It also maps debitAccountCode/creditAccountCode to expense_accounts, enriching the agent's understanding of valid values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Files a pending draft' for a correction to a posted paid expense) and immediately distinguishes itself from siblings by emphasizing that nothing changes until human approval. It also names the source for finding the entry id (search_expenses) and contrasts with create_expense_draft, create_journal_draft, and the delete lane.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The ROUTING section is explicit about when to use this tool versus the delete lane, create_expense_draft, or create_journal_draft. It also covers blocked cases and the fallback to revise_draft when the underlying expense changes after reading, leaving no ambiguity about alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_invoice_draftFile a pending correction draft for a posted invoice (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to an invoice ALREADY POSTED in Taokeh — the door for an invoice that went onto the books with a wrong line, a wrong date or the wrong payment method. This does NOT change anything: it files a pending DRAFT an ADMIN reviews as a DIFF (current → proposed) and approves; only that tap rewrites the invoice, and it keeps the SAME invoice number, the original attached document and any payment-reminder history. Pass saleId, the invoice's REAL id from search_documents — an invoice number is refused, because two documents can carry the same printed number and correcting the wrong one is worse than asking. ROUTING — three different mistakes, three different doors, and picking the wrong one is expensive. (1) THE SALE NEVER HAPPENED (keyed twice, entered on the wrong company, cancelled before delivery): the owner VOIDS it — Sales → the invoice → Void, in Taokeh. There is deliberately NO AI lane for voiding; erasing a document from the books is a person's decision at the screen. Tell them where to click. (2) THE SALE HAPPENED BUT THE INVOICE IS WRONG (a wrong line, a wrong quantity or price, the wrong date, the wrong payment method): THIS tool. (3) SOMETHING CHANGED AFTER THE SALE (goods came back, a price was renegotiated, a discount was agreed later): create_credit_note_draft — a real event that belongs on the books as its own document, not an erasure of history. WHAT YOU CAN CHANGE: saleDate, paymentMethod ('CASH' or 'CREDIT'), refNo (the customer's own reference / PO number), term (the payment term, stored as written — changing an on-account invoice to one of the five presets, Due on Receipt / Net 7 / Net 14 / Net 30 / Net 60, re-derives the due date from the invoice date unless you also send dueDate, and the admin sees the resulting due date before approving), dueDate (only when the due date ITSELF is wrong — correcting just the invoice date already moves the due date by the same number of days, so the agreed credit period survives), and lines — the FULL REPLACEMENT SET, so send every line the invoice should have, including the ones that were already right. ON A MEASUREMENT COMPANY, do NOT send the small negative rounding line the engine adds under a per-foot line: it is Taokeh's own line, not one of the invoice's, it is preserved untouched by a header-only correction, and it is re-derived from whatever timber line you do send. You cannot mark a line as one either — that is a fact Taokeh reads off the document, never something a caller asserts. WHAT YOU CANNOT, each refused by name: the CUSTOMER (a different customer is a different document — credit this one and raise a new invoice to the right party); the INVOICE NUMBER (preserved across an edit; Taokeh's own Edit screen cannot change it either); and the lines of an invoice that spreads its revenue over time under MFRS 15 (which new line inherits which schedule is a guess, and a guess about deferred revenue is not something a reviewer can check on a diff — header-only corrections still work WHILE the schedule is only part-way through, and keep it intact). On a deferred invoice whose revenue has ALREADY been recognised — any period posted, up to and including a finished schedule — the invoice is closed to edits altogether, header fields included, because rewriting it would re-state revenue in periods that may already be closed and filed; the remedy is a credit note, and the refusal says so. THESE BLOCK AN EDIT OUTRIGHT, checked when you file and again when the admin taps, and each refusal names the real next step: the document is not an invoice; it is validated on MyInvois; it has been submitted to LHDN; it was consolidated from delivery orders; a payment has been allocated to it. A LIVE CREDIT OR DEBIT NOTE against the invoice blocks only a lines change: a header-only correction (date, payment method, refNo, term, due date) still files and approves, and the note keeps its number and is re-linked to the corrected invoice. It still refuses if that note has gone to LHDN, or if the new date would fall after the note's date. AND A LOCKED PERIOD BLOCKS THE APPROVAL (2026-09-19): an edit reposts the invoice and removes the original entry, and removing a line from a month the company has already locked and filed is refused — so if the invoice is dated on or before the lock date, the admin's tap is refused by name, even when your proposal re-dates the invoice into an open month. It is checked at APPROVAL, not at filing, so a draft can still be filed against a locked invoice; do not do it — tell the owner instead that the correction has to be a credit note dated today (create_credit_note_draft), which leaves the filed months as filed, or that an admin has to unlock in Opening balances → Review & lock first. A field the invoice already agrees with is dropped, and a proposal that changes nothing is refused rather than filed as an empty diff for a human to tap. If the invoice is edited by someone else after you read it, the admin is shown BOTH versions and the one-tap approval is refused — that is by design; call revise_draft (kind 'invoice_update') to re-read and re-propose. An edit in Taokeh re-books the invoice under a NEW id with the same number: a pending draft FOLLOWS it (it shows as changed since you read it, and revise_draft on the draft still works), and calling this tool with the old id is refused with the new id named — re-read the invoice before proposing against it.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note in the reviewer's language: what was wrong, what you changed, and anything they should double-check before they approve a rewrite of a posted document. | |
| saleId | Yes | REQUIRED — the posted invoice's real id, from search_documents. An invoice number is refused: correcting the wrong document silently is worse than asking which one they mean. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name; a proposal the invoice already agrees with is refused rather than filed. | |
| needsReview | No | Set true when something gave you pause — a quantity you inferred, a price the user was unsure about. It flags the draft for the reviewer. | |
| printedTotal | No | The grand total the corrected invoice should show, if the user told you one — cross-checked against the server's own derivation and flagged on the review screen. Reconcile a disagreement in chat; never quietly average the two. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is not read-only, not destructive, not idempotent; the description carries far more: nothing changes until a human taps, the approval-time lock check, the credit-note interaction, the re-booking to a new id, and the both-versions-refused behavior. No contradiction with destructiveHint=false since the tool files a draft rather than mutating directly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very long, but front-loaded with the single most important fact in caps, and the length is largely earned by the routing and refusal detail. There is some duplication of schema-level parameter rules (term, dueDate, quantity) that could be trimmed without losing routing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-risk mutation tool with no output schema, the description covers the approval model, every refusal class by name, the locked-period interaction, deferred-revenue handling, and the recovery path via revise_draft. Nothing an agent needs in order to call it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description goes beyond it by naming which fields can be changed and the rules around them (lines as a full replacement set, omitting the engine rounding line, the term→dueDate re-derivation). It does repeat a fair amount of what the schema already documents, so it is not a full point above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with unusual precision: it 'files a pending DRAFT' to correct an 'invoice ALREADY POSTED', and immediately contrasts that with voiding (Sales → Void) and create_credit_note_draft. An agent can distinguish this from every sibling draft tool without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates three routing cases (never happened → void; invoice wrong → this tool; something changed after the sale → create_credit_note_draft) and names the alternative for each. It also states when the tool must NOT be used (locked period, deferred revenue already recognised) and names the real next step each time.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_logo_draftFile a pending logo-change draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a new COMPANY LOGO for this company — the mark Taokeh prints on every invoice, quote, receipt and statement it generates, and that fronts the online storefront. This does NOT change anything: it files a pending DRAFT an ADMIN reviews by looking at the logo they have now beside the one you are proposing, and approves with one tap; only that tap sets it. Send the image one of two ways, never both. Lane 1 (default, any size): call request_attachment_upload, PUT the raw bytes to its uploadUrl, and pass the returned attachmentToken here. Lane 2 (fallback): if your shell cannot reach taokeh.my — sandboxed clients sit behind a network allowlist and the PUT fails 403/blocked; that is YOUR sandbox, not Taokeh — send attachmentBase64 + attachmentMediaType inline instead, always with attachmentBytes (the file's decoded size on disk) so a truncated paste is rejected rather than filed. PNG or JPEG only: those are the only formats Taokeh can embed in a generated PDF, and anything else is refused by name. Taokeh SHRINKS oversized artwork for you rather than sending you away to resize it — a 5000×5000 export is accepted and stored at 2000px, and the review screen says so; only an image so large it is no longer a logo is refused. A LOGO IS ONE VALUE, so the proposal REPLACES the whole thing rather than patching it, and there is only ever ONE logo proposal waiting: filing a second one replaces the first, which is rejected as superseded. If the company already has a logo the review page shows both images side by side; if it has none, this is the first one. If an admin sets or removes a logo in Taokeh after you file, the one-tap doors refuse and send them to the full review page — a human's own choice is never overwritten by a proposal that never saw it. ⛔ THIS IS THE COMPANY LOGO ONLY. It is not the Pioneer program's testimonial artwork, and nothing on this connector can touch that. BE HONEST: propose only an image the user actually gave you for this purpose. Never generate a logo and file it as though they had chosen it, and tell them in the same breath that nothing on their paperwork changes until they tap Approve.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note, in the reviewer's language: where this image came from and what they should check before it goes on their invoices. | |
| needsReview | No | Set true when something gave you pause — an image you are not certain is the right mark, or one the user sent in passing rather than chose. | |
| attachmentBytes | No | The decoded byte size of the image on disk — send it alongside attachmentBase64 and the server rejects a truncated paste instead of filing half a logo. | |
| attachmentToken | No | The token from request_attachment_upload, AFTER you have PUT the image bytes to its uploadUrl. The preferred lane for any real logo file. Mutually exclusive with attachmentBase64. | |
| attachmentBase64 | No | The logo image as base64 — the fallback lane, for when your shell cannot reach taokeh.my. PNG or JPEG only. Mutually exclusive with attachmentToken. | |
| attachmentSha256 | No | The SHA-256 of the image as 64 hex chars — optional second integrity check alongside attachmentBase64. | |
| attachmentFilename | No | The original file name, e.g. 'acme-logo.png'. Shown on the review page so the admin recognises what they were sent. | |
| attachmentMediaType | No | The image's MIME type — 'image/png' or 'image/jpeg'. Required when attachmentBase64 is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false and carry almost no signal, so the description carries the full burden — and it delivers richly: supersede semantics ('filing a second one replaces the first, which is rejected as superseded'), shrink behavior ('5000×5000 export is accepted and stored at 2000px'), format refusal ('anything else is refused by name'), sandbox-induced PUT 403s, truncated-paste rejection via attachmentBytes, and human-override protection ('a human's own choice is never overwritten by a proposal that never saw it'). No contradiction with the annotations — destructiveHint=false aligns with 'This does NOT change anything'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the density of decision-relevant information justifies it: every major block (pending semantics, two lanes, formats, shrinking, one-value replacement, human-override, scope boundary, honesty rule) maps to a distinct risk the agent must handle. Critical info is front-loaded ('FILES A PENDING DRAFT ONLY') and warnings are scannable via caps and the ⛔ marker. Minor deductions for redundancy ('This does NOT change anything' restates the opening sentence) and an editorial aside ('that is YOUR sandbox, not Taokeh').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool — 8 parameters, two mutually exclusive input lanes, no output schema, no idempotency — the description covers purpose, exact invocation mechanics, side effects, concurrency behavior, failure modes, and ethical constraints. The only notable gap is the success response: nothing states what the agent should expect back after filing (e.g., a draft identifier or confirmation), which matters more because there is no output schema. That single omission keeps this from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds genuine workflow-level meaning beyond the schema: it groups parameters into two mutually exclusive lanes (attachmentToken alone vs attachmentBase64 + attachmentMediaType + attachmentBytes), explains why attachmentBytes is mandatory alongside base64 ('a truncated paste is rejected rather than filed'), and states the format constraint ('PNG or JPEG only'). It adds little on notes, needsReview, attachmentSha256, or attachmentFilename, but those are already well documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource with a decisive qualifier: 'FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a new COMPANY LOGO.' The title independently confirms it ('File a pending logo-change draft'). It also differentiates from siblings by fencing off the Pioneer testimonial artwork ('THIS IS THE COMPANY LOGO ONLY'), so an agent cannot confuse it with the large create_*_draft / update_*_draft family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and how-to-use guidance with named alternatives: Lane 1 directs the agent to call request_attachment_upload (a sibling) and pass its attachmentToken; Lane 2 specifies the exact fallback condition ('if your shell cannot reach taokeh.my... PUT fails 403/blocked') and the parameters to use instead. Scope exclusions are explicit ('It is not the Pioneer program's testimonial artwork'), and the honesty rule states when NOT to file ('Never generate a logo and file it as though they had chosen it').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_product_draftFile a pending product-change draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose corrections to a product ALREADY in Taokeh's catalogue — a selling price keyed wrong, a name that came across from the old system mangled, a missing barcode, the wrong unit, a reorder point nobody set, an item filed under the wrong category. This does NOT change anything: it files a pending DRAFT the owner reviews as a DIFF (current → proposed, changed fields only) and approves with one tap; only that tap writes. Name the product with productId (the id resolve_product returns) or with its EXACT sku — a SKU is unique inside a company, so it is an exact key; a product NAME is not, and is refused. If you pass both and they disagree, the call is refused rather than guessing. proposed holds ONLY the fields you want changed: name, unit, unitPrice, reorderPoint, barcode, and the category (pass categoryId, or category as a name matched against the categories the company ALREADY has — this tool will never create a category, because the owner is approving a change to a product, not new master data). ⛔ WHAT IT CANNOT DO, each refused by name with the screen that owns it: it cannot change STOCK ON HAND — a quantity change moves inventory and cost of goods sold together, so Taokeh only takes it at Products → Adjust stock where a counted reason is required; it cannot change the AVERAGE COST, which the purchases that set it own and which is what the inventory is worth on the balance sheet, so an edit here would restate that value with no journal behind it; it cannot change the SKU, the key everything else resolves on; it cannot set a description, which the product page has no control for at all (the catalogue import does); it cannot switch "Service — no stock" (isService) on or off — the owner ticks it on the product page; and it cannot publish or unpublish the item to the online shop, because what strangers can see and buy is not something an AI proposal should decide. A field the record already agrees with is dropped, and a proposal that changes nothing is refused rather than filed as an empty diff. One product per call — loop for a sweep, because each reviewable diff is the point. BE HONEST: never guess a price or a barcode; leave the field out and say so in notes with needsReview.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | The EXACT SKU, as an alternative to productId. Exact only — a near miss is refused, and a product NAME is never accepted. | |
| notes | No | A SHORT reviewer note, in the reviewer's language: what you changed, why, and what they should double-check. | |
| proposed | Yes | ONLY the fields you want changed. A field the record already agrees with is dropped; an unknown or deliberately-excluded field is refused by name. | |
| productId | No | The product's real id, from resolve_product. Give this OR `sku`. | |
| needsReview | No | Set true when something gave you pause — a price the user was unsure about, a barcode read off a blurry photo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say this is a non-read, non-destructive, non-idempotent write. The description adds the entire behavioral contract on top: nothing commits until a human taps approve, output is a current→proposed diff of changed fields only, unchanged fields are dropped, an empty diff is refused, identity mismatches are refused rather than guessed, and unmatched categories are rejected. That is substantial context the annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical constraint is front-loaded in the first line and the body is organized into identity, payload, exclusions, and honesty rules. It is long for its parameter count and repeats the same points (the no-op/human-approval idea three times, the never-creates-a-category idea twice), plus caps-and-symbol styling that inflates it, which keeps it out of 5 territory.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a nested payload object, the description covers identity resolution, allowed and forbidden fields, and refusal semantics thoroughly. It never says what a successful call returns or where the filed draft surfaces (presumably my_work / revise_draft), which is the one gap an agent would want closed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so 3 is the floor, but the description earns above it: it explains WHY `sku` is an exact key and why a product name is refused, defines the productId-or-sku choice and the both-disagree refusal, and frames `proposed` as changed-fields-only while naming the closed field set. It adds meaning rather than restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource: propose corrections to a product already in the catalogue by filing a pending draft. It separates itself from adjacent tools by naming where the excluded operations actually live (Products → Adjust stock, the catalogue import, the product page) and by pointing at resolve_product as the source of the id. An agent can tell this apart from stage_master_data or resolve_product without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use triggers (wrong price key, mangled name, missing barcode, wrong unit, unset reorder point, wrong category), explicit when-NOT-to-use rules for six refused field classes each with the owning screen, and an operational rule (one product per call; loop for a sweep). The alternative paths are named rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_purchase_order_draftFile a pending purchase-order correction draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a PURCHASE ORDER that already exists — a wrong quantity or cost, a missing or extra line, the wrong date, term or notes. An admin or bookkeeper reviews a plain current → proposed DIFF and approves. A purchase order posts nothing and moves no stock, so the correction touches no ledger and no stock; a SENT order stays SENT, and Taokeh emails nobody — sending the corrected order to the supplier is the user's own step. ROUTING — (1) the order should not stand: the owner cancels it in Taokeh (Purchase orders → the PO → Cancel); there is no AI lane for cancelling. (2) the order is wrong and not yet converted: THIS tool — a DRAFT or a SENT order both qualify. (3) the order was already CONVERTED into a bill, or is CANCELLED: it is closed to edits — a converted order's bill is the document now, corrected in Taokeh (Purchases → the bill → Edit). A NEW order is create_purchase_order_draft. WHAT YOU CAN CHANGE: poDate, term (null clears it), notes (null clears them) and lines. A purchase order is taxed at a whole-document rate Taokeh stores only as an amount, so when a taxed order's LINES change you must also send taxRate (see its field). WHAT YOU CANNOT, each refused by name: the SUPPLIER (an order to someone else is a different order), the PO NUMBER, the CURRENCY or exchange rate, the TAX, and marking sent, cancelling or converting. INVENTORY-ONLY lines, as on create_purchase_order_draft. WHAT HAPPENS ON APPROVAL: the document is updated IN PLACE — same id, same number — through the SAME update Taokeh's own Edit screen runs. Every field you leave out is carried across exactly as it stands (the party, the attention contact, the currency and its frozen exchange rate, the tax) and is never reset. LINES: lines is the FULL replacement set. READ THE LINES FIRST with search_document_lines — each row's lineNo is what a { keep: } entry refers to, and keep is how a line that is already right survives untouched (a measured line keeps its frozen unit and its rounding row). A changed or new line is written in full, in the create tool's line shape, and the server re-derives its quantity and amount. A proposal that changes nothing is refused rather than filed as an empty diff. STALE: if a person edits the document after you read it, the one-tap approval is refused and the reviewer is shown both versions — call revise_draft with this draft's kind to re-read and re-propose. Pass the document's REAL id from search_documents; a document number is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note in the reviewer's language: what was wrong, what you changed, and anything they should check before approving. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name; a proposal the purchase order already agrees with is refused rather than filed. | |
| needsReview | No | Set true when something gave you pause — a quantity you inferred, a price the user was unsure about. It flags the draft for the reviewer. | |
| printedTotal | No | The grand total the corrected purchase order should show, if the user told you one — cross-checked against the server's own figure and flagged on the review screen. | |
| purchaseOrderId | Yes | REQUIRED — the purchase order's real id, from search_documents. A PO number is refused. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations carry no safety hints (all false), so the description fully discloses behavior: nothing changes until human approval, the document is updated in place on approval, no ledger/stock impact, no emails, staleness refusal, and that a no-op proposal is refused. This goes far beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured with labeled sections (ROUTING, WHAT YOU CAN CHANGE, WHAT YOU CANNOT, WHAT HAPPENS ON APPROVAL, LINES, STALE) and front-loads the key safety point. Some redundancy with the title and schema exists, but every section adds operational detail, so it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers all necessary context for safe and correct invocation: when to use, what fields can be changed, what is refused, line semantics, staleness handling, and the id requirement. References sibling tools for pre-read steps. No output schema is needed for a draft-filing action, so completeness is high.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds critical context not in the schema: taxRate is required when lines change on a taxed order, lines is the full replacement set, keep refers to lineNo from search_document_lines, and purchaseOrderId must be the real id from search_documents. This materially improves correct usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'FILES A PENDING DRAFT ONLY' and clearly states the tool proposes corrections to an existing purchase order. It explicitly distinguishes from create_purchase_order_draft (new orders) and revise_draft (stale re-proposal), so an agent can pick it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains a dedicated ROUTING section that spells out when to use this tool (wrong, not converted, DRAFT or SENT), when to cancel in Taokeh instead (order should not stand), when it's closed (converted or cancelled), and that new orders go to create_purchase_order_draft. Also instructs to use revise_draft on staleness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_quote_draftFile a pending quote-correction draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a QUOTE that already exists — a wrong price or quantity, a missing or extra line, the wrong date, validity, term or notes. An admin or bookkeeper reviews a plain current → proposed DIFF and approves. A quote posts nothing, so the correction touches no ledger, no stock and no e-invoice. ROUTING — (1) the quote should not exist at all: the owner voids it in Taokeh (Quotes → the quote → Void); there is no AI lane for voiding. (2) the quote is wrong and still OPEN: THIS tool. (3) the quote was already CONVERTED into an invoice or a delivery order: it is closed to edits — correct the document it became (update_invoice_draft or update_delivery_order_draft). A NEW quote is create_quote_draft. WHAT YOU CAN CHANGE: quoteDate, validUntil (null clears it), term, notes (null clears it) and lines. WHAT YOU CANNOT, each refused by name: the CUSTOMER (a quote to someone else is a different quote), the QUOTE NUMBER, the CURRENCY or exchange rate, and converting or voiding. ONLY AN OPEN QUOTE can be corrected — checked when you file and again when the reviewer taps. Leave a line's taxCode out and it inherits this quote's own SST code for its product — unless the quote carries that product under more than one code, when you must state it. A quote taxed at a whole-document rate (no SST codes on its lines) needs taxRate whenever its LINES change (see that field). WHAT HAPPENS ON APPROVAL: the document is updated IN PLACE — same id, same number — through the SAME update Taokeh's own Edit screen runs. Every field you leave out is carried across exactly as it stands (the party, the attention contact, the currency and its frozen exchange rate, the tax) and is never reset. LINES: lines is the FULL replacement set. READ THE LINES FIRST with search_document_lines — each row's lineNo is what a { keep: } entry refers to, and keep is how a line that is already right survives untouched (a measured line keeps its frozen unit and its rounding row). A changed or new line is written in full, in the create tool's line shape, and the server re-derives its quantity and amount. A proposal that changes nothing is refused rather than filed as an empty diff. STALE: if a person edits the document after you read it, the one-tap approval is refused and the reviewer is shown both versions — call revise_draft with this draft's kind to re-read and re-propose. Pass the document's REAL id from search_documents; a document number is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note in the reviewer's language: what was wrong, what you changed, and anything they should check before approving. | |
| quoteId | Yes | REQUIRED — the quote's real id, from search_documents. A quote number is refused: correcting the wrong document silently is worse than asking which one they mean. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name; a proposal the quote already agrees with is refused rather than filed. | |
| needsReview | No | Set true when something gave you pause — a quantity you inferred, a price the user was unsure about. It flags the draft for the reviewer. | |
| printedTotal | No | The grand total the corrected quote should show, if the user told you one — cross-checked against the server's own figure and flagged on the review screen. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the sparse annotations (readOnlyHint false, destructiveHint false). It discloses that nothing changes until human approval, that approval updates the document IN PLACE through the same Edit-screen path, that omitted fields are carried across untouched, that no-op proposals are refused, and that stale edits are refused with both versions shown. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place — ROUTING, WHAT YOU CAN CHANGE, WHAT YOU CANNOT, WHAT HAPPENS ON APPROVAL, LINES, STALE. The most critical constraint ('FILES A PENDING DRAFT ONLY') is front-loaded in all caps, and the content is organized so an agent can quickly extract routing and refusal rules.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 5-parameter tool with a nested `proposed` object and no output schema, the description covers everything needed to call it correctly: required quoteId semantics, real-id vs number refusal, editable vs immutable fields, line replacement semantics, tax inheritance, stale-draft recovery, and approval-time behavior. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds critical meaning: which fields are patchable, which are refused by name, that `lines` is a FULL replacement set, that `keep` references lineNo from search_document_lines, how omitted taxCode inherits the document's SST code, and the special taxRate requirement. This substantially enriches the bare field definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear statement: 'FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH.' It states a specific verb (propose a correction), a specific resource (an existing QUOTE), and explicitly distinguishes itself from siblings like create_quote_draft, update_invoice_draft, and update_delivery_order_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The ROUTING section gives exact when-to-use conditions: use this tool when a quote is wrong and still OPEN; void through Taokeh if it should not exist; correct the converted document if it became an invoice or delivery order; use create_quote_draft for a new quote. It also names revise_draft for stale-draft handling, leaving no ambiguity about alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_recurring_invoice_draftFile a pending recurring-invoice change draft (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose changes to a recurring invoice ALREADY set up in Taokeh — a price rise across a subscription, a changed quantity, a new end date, a cadence correction, or PAUSING the billing. This does NOT change anything: it files a pending DRAFT the owner reviews as a DIFF (current → proposed, changed fields only) and approves; only that tap writes. Name the schedule with recurringId, its real id — there is NO name fallback, because a schedule title is free text and one customer often has two schedules (a monthly retainer and an annual licence), so guessing between them would re-bill the wrong contract. proposed holds ONLY what you want changed: title, customerId/customer, lines, cadence, paymentMethod, startDate, endDate, maxOccurrences, notes, and status ("paused" to stop future invoices, "active" to resume). ⚠ lines is WHOLE-REPLACE: send every line the schedule should have, including the unchanged ones — for a price rise, send the same lines with the new prices. ⚠ APPROVING CAN MOVE THE BILLING DATE even when you change only a price: Taokeh re-derives the next issue date from the start date on every save, so the review page states the exact date the next invoice would go out, and whether that has moved. ⛔ WHAT IT CANNOT DO, each refused by name: it cannot set issueMode: "auto" — that would let one approval remove the human tap from every future invoice; it cannot set status: "ended", a ONE-WAY door no later draft could undo (propose "paused" instead, which stops the billing just as completely and can be resumed); it cannot set MFRS 15 revenue spreading, an accounting-policy decision on the schedule's own page; and it cannot touch the invoices this schedule has ALREADY issued — a change reaches the next firing and no earlier one. If a past invoice was wrong, credit it (create_credit_note_draft); the schedule only decides what happens next. A field the record already agrees with is dropped, and a proposal that changes nothing is refused rather than filed as an empty diff. One schedule per call — loop for a sweep across several, because each reviewable diff is the point.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note, in the reviewer's language: what you changed, why, and what they should double-check. | |
| proposed | Yes | ONLY the fields you want changed. A field the record already agrees with is dropped; an unknown or deliberately-excluded field is refused by name. | |
| needsReview | No | Set true when something gave you pause — a price the user was unsure about, a date you inferred. | |
| recurringId | Yes | The recurring invoice's real id. Required — there is no name fallback. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are sparse (all false), so the description carries the full burden, and it delivers richly. It discloses the draft/approval two-step, the whole-replace semantics of lines, the fact that approving can move the billing date even on a price-only change, the one-way door of 'ended', the refusal of no-op proposals, and the dropping of unchanged fields. It also warns about the absence of a name fallback and why. This goes well beyond what annotations could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place: the opening warning, the use cases, the parameter semantics, the explicit refusals, and the operational notes. It is front-loaded with the most critical fact (nothing changes until human approval) and uses typographic signals (⚠, ⛔, ALL-CAPS) to structure scanning. It loses one point for length and some redundancy (e.g., the 'paused' explanation appears twice), but it remains well-organized and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested objects, 4 parameters, no output schema, and sparse annotations, the description is remarkably complete. It covers what the tool does, what it cannot do, how parameters behave, what the human review entails, and how to handle multi-schedule sweeps. The only minor gap is the lack of an explicit return value description, but since there is no output schema and the tool's purpose is to file a draft, the absence is not material. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds substantial meaning beyond the schema. It explains the real-world consequence of endDate null (extends billing indefinitely), the reviewer-facing meaning of notes, the strict line keys, the whole-replace semantics, and the status enum's behavioral implications (paused is reversible, ended is refused). It also clarifies the customer/customerId alternative and the no-name-fallback rule for recurringId. This is high-value semantic enrichment, not repetition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a loud, specific statement: 'FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH.' It names the exact verb (file/propose), the resource (recurring invoice already set up in Taokeh), and the scope (pending draft for human approval). It distinguishes itself from create_recurring_invoice_draft and update_invoice_draft by emphasizing it only proposes changes to an existing schedule, and it lists concrete use cases (price rise, quantity, end date, cadence, pausing). This is far beyond a vague purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: use it to propose changes to an already-set-up recurring invoice, and explicitly says what it cannot do (set issueMode auto, set status ended, set MFRS 15, touch already-issued invoices). It also names the alternative for past invoice corrections: 'If a past invoice was wrong, credit it (create_credit_note_draft); the schedule only decides what happens next.' It even gives a looping instruction for sweeping multiple schedules. This is textbook usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_supplier_debit_note_draftFile a pending correction draft for a posted supplier debit note (human approves in Taokeh)AInspect
FILES A PENDING DRAFT ONLY — NOTHING CHANGES UNTIL A HUMAN REVIEWS AND APPROVES IT IN TAOKEH. Propose a correction to a BUY-SIDE DEBIT NOTE already posted in Taokeh — a purchase return raised against a SUPPLIER's bill (Taokeh's /debit-notes page). This is NOT the sell-side debit note that create_debit_note_draft files (an additional charge on your own invoice); that document has no correction lane at all — the owner voids it and issues a new one. An ADMIN reviews this as a DIFF (current → proposed) and approves; only that tap re-books the note, and it keeps the SAME debit-note number, the bill it was raised against and any attached document. Pass debitNoteId, the note's REAL id from search_documents (docType debit_note). ROUTING — (1) THE RETURN NEVER HAPPENED (keyed twice, wrong bill): the owner VOIDS the note in Taokeh (Debit notes → the note → Void); voiding keeps it on the register marked VOID. (2) THE RETURN IS RIGHT BUT TYPED WRONG (a wrong line, quantity, cost, date or reason): THIS tool. (3) MORE GOODS WENT BACK LATER: a new debit note, raised by the owner in Taokeh — there is no AI lane for creating a buy-side debit note. WHAT YOU CAN CHANGE: dnDate, reason, and lines — the FULL REPLACEMENT SET. ⛔ EVERY LINE MUST ANSWER goodsReturned (true = these goods physically went back to the supplier and leave stock, false = money only). There is no default. Your answer only PRE-FILLS the admin's per-line question on the full review page; a ONE-TAP approval never moves stock and is refused whenever the note already took goods off the shelf or any line says goods went back. Stock only ever leaves for a product the BILL actually received, never more than the bill still has standing. SST is inherited from the bill — there is no tax code to send. WHAT YOU CANNOT, each refused by name: the SUPPLIER, the debit-note NUMBER, the BILL it was raised against. BLOCKED OUTRIGHT, checked when you file (the server runs the real correction and rolls it back) and again when the admin taps: the note is voided; it is validated or exported for MyInvois as a self-billed document; a supplier refund is allocated to it; the bill it was raised against is gone; its date is inside a LOCKED period; or the corrected note would credit more than the bill still owes. A proposal that changes nothing is refused. If the note is edited after you read it, the one-tap approval is refused — call revise_draft (kind 'supplier_debit_note_update').
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | A SHORT reviewer note: what was wrong, what you changed, and whether the goods really went back. | |
| proposed | Yes | ONLY what you want changed. An unknown or not-patchable key is refused by name. | |
| debitNoteId | Yes | REQUIRED — the posted BUY-side debit note's real id, from search_documents (docType debit_note). | |
| needsReview | No | Set true when something gave you pause. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide minimal coverage (readOnlyHint=false), so the description carries the full burden. It discloses that nothing changes until human approval, explains the server rollback mechanism, lists all blocked conditions, and clarifies that stock moves only via admin approval. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured with sections and front-loaded warnings. Each sentence adds substantive context (routing, constraints, blocked conditions). It could be tightened, but the density of information justifies the length for a complex tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the complexity (nested objects, no output schema, 4 params), the description covers all necessary context: how to obtain the id, what can/cannot be changed, blocking conditions, revision guidance, and the approval flow. Nothing an agent needs to call correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds valuable meaning beyond the schema: it explains that debitNoteId is the real id from search_documents, that lines must include goodsReturned with no default, and that proposed is a full replacement set. This elevates it above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: files a pending correction draft for a posted buy-side debit note, with explicit distinction from the sell-side debit note tool (create_debit_note_draft). It names the resource (buy-side debit note) and the specific verb (file a draft), making the purpose unambiguous and differentiating it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit routing guidance with three numbered scenarios: when to void, when to use this tool, and when to create a new debit note. It also directs to revise_draft if the note is edited after reading. This is textbook when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
working_capitalWorking capital (days + who pays slowly)ARead-onlyInspect
Working capital in DAYS, month by month, plus who is slow to pay — the deterministic answer to "the profit and loss says I made money, so where is the cash?". Give from and to as YYYY-MM, inclusive; the window is capped at 36 months, the same cap financial_timeseries carries. TWO HALVES COME BACK IN ONE CALL. First, months[]: per month revenue, cogs, arClose, apClose, inventoryClose and four ratios — debtorDays (arClose ÷ revenue × days: days of sales sitting unpaid in customers' hands), creditorDays (apClose ÷ cogs × days: days of cost sitting unpaid in yours), inventoryDays (inventoryClose ÷ cogs × days: days of cost sitting on the shelf) and cashConversionCycle (debtorDays + inventoryDays − creditorDays). Every row states days, the calendar days it covers, because a days ratio without its denominator cannot be checked. Second, customers and vendors: per party avgDaysToPay (amount-weighted, from each document's own date to the date the money moved), avgDaysPastDue (the same weighting measured from the due date), invoices, paid and withTerms — slowest first, up to 50 a side with total, shown and more. NULL IS AN ANSWER. A ratio whose denominator is zero comes back null — no comparable base — never 0 and never an enormous number; cashConversionCycle is null whenever any one of its three legs is, and is never the sum of the legs that happen to exist. Relay a null as "there is no base to measure that against", never as zero. ⚠ INVENTORY DAYS IS WITHHELD unless this company runs PERPETUAL costing. On Taokeh's default mode (modified periodic) a sale does not move stock and cost of goods sold is recognised only at a stock take, so the inventory balance sits still and then jumps while monthly COGS is zero in most months — a monthly inventory-days figure on that basis would measure something else. inventoryClose is still returned, because it is a real ledger balance; inventoryDays and cashConversionCycle are null and a note says why. Do NOT compute it yourself from inventoryClose and cogs: that is precisely the number the server declined to state. WHAT PAYMENT SPEED EXCLUDES, and say so when you relay it: cash sales and cash bills (they settle on their own date and would average zeros into every party), opening invoices and opening bills (those dates belong to the old system, not to this company's paper), refunds (a negative allocation is money walking back out, not a customer paying) and contra set-offs (a set-off moves no bank money). MONTHS THAT DO NOT EXIST ARE ABSENT, NOT ZERO: months ending before this company's accounting start date, and months after today. A month still running is partial: true with through = today, and its days counts only the days that have happened, so both sides of every division are month-to-date — say so, and never let that row bend a trend. The cutover month is partial with startsAt, and its days, revenue and cogs all count from there — so its revenue can be lower than financial_timeseries shows for the same month, which runs from the 1st. EQUAL BASIS: revenue and cogs are the /reports income statement, arClose and apClose the A/R and A/P aging reports' own totals as of month end, inventoryClose the ledger balance of the mapped inventory account. Nothing is re-summed, and basis states every formula with its exact inputs. DRILL DOWN RATHER THAN GUESS: a slow customer here → ar_aging for the buckets, open_invoices for that customer's unpaid documents, search_documents with docType 'payment' for the settlements behind the average, get_document_pdf for the paper itself. FIGURES ONLY — no verdict, no benchmark, no trend label, no "healthy" and no "improving". The reading is yours to write. A very wide window can exceed this call's time budget, in which case it is REFUSED (naming how far it got) rather than answered with a short series. Nothing is written.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only readOnlyHint and openWorldHint, so the description carries the full behavioral burden — and discharges it thoroughly. It discloses null-is-an-answer semantics (zero denominators never become 0), the perpetual-costing withholding with rationale, what payment speed excludes (cash sales, opening invoices, refunds, contra set-offs), absent-vs-zero months, partial/cutover month math, the exact data basis, and refusal on oversized windows. The closing 'Nothing is written' is consistent with readOnlyHint; no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but every capitalized block — NULL, INVENTORY DAYS, WHAT PAYMENT SPEED EXCLUDES, MONTHS, EQUAL BASIs, DRIL DOWN, FIGURES ONLY — carries distinct operational knowledge an agent would otherwise mis-relay. It is front-loaded with the purpose line and scannable via headers; some tightening was possible, but nothing reads as filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description alone must specify the return shape — it does, down to field names (arClose, apClose, inventoryClose), ratio formulas with their denominators, customer/vendor aggregates, and the basis of every figure. Edge-case handling (nulls, partial months, cutover, refusal) and drill-down routing are all present, so an agent can call, parse, and relay correctly without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate; it adds that from/to are 'YYYY-MM, inclusive' and that the window caps at 36 months, matching financial_timeseries. The format itself is already in the schema pattern, so the marginal contribution is the inclusivity rule and cap. Month-existence semantics (business start date, after today) are covered elsewhere in the body, making this adequate but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific deliverable — 'Working capital in DAYS, month by month, plus who is slow to pay' — and frames it against a concrete business question ('the profit and loss says I made money, so where is the cash?'). The 'deterministic answer' framing marks the scope as calculation, not judgment, and the return structure (months[] plus customers/vendors) is clearly announced.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the drill-down chain for follow-up: a slow customer here → ar_aging, open_invoices, search_documents with docType 'payment', get_document_pdf. It also states when-not: inventory days are withheld under periodic costing and 'Do NOT compute it yourself from inventoryClose and cogs', plus the wide-window refusal. The 36-month cap is tied to a sibling (financial_timeseries), giving the agent an anchor across tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
stock_movements2 fields changed- changed
Input schema / properties / from / descriptionPrevious value: -"Earliest movement date (YYYY-MM-DD, Malaysian time). The main way to break a >200-row pull into honest slices."New value: +"Earliest docDate (YYYY-MM-DD; the document date, else the Malaysian posting day). The main way to break a >200-row pull into honest slices." - changed
Input schema / properties / to / descriptionPrevious value: -"Latest movement date (YYYY-MM-DD, Malaysian time)."New value: +"Latest docDate (YYYY-MM-DD; the document date, else the Malaysian posting day)."
1 tool update
- Changed
stage_master_data1 field changed- changed
Input schema / properties / rows / items / properties / settingKey / descriptionPrevious value: -"settings: the setting's name, from the ALLOWLIST of 33 — the exact list, each key's meaning and its current value come from intake_contract(doc_type:'settings'), which is the source of truth; it spans company identity and bank details, statutory employer identifiers, locale, currency, tax and SST preferences, document colour, measurement, the accounting start date, and safe non-publishing storefront preferences. Anything else is refused row by row and NEVER created. Payment credentials, storefront publishing, ads consent, terms acceptances, the books lock, SST registration and every einvoice.* key are refused by name with the reason — they are the owner's own act."New value: +"settings: the setting's name, from the ALLOWLIST of 34 — the exact list, each key's meaning and its current value come from intake_contract(doc_type:'settings'), which is the source of truth; it spans company identity and bank details, statutory employer identifiers, locale, currency, tax and SST preferences, document colour, measurement, the accounting start date, and safe non-publishing storefront preferences. Anything else is refused row by row and NEVER created. Payment credentials, storefront publishing, ads consent, terms acceptances, the books lock, SST registration and every einvoice.* key are refused by name with the reason — they are the owner's own act."
2 tool updates
- Changed
update_bill_draft1 field changed- changed
Input schema / properties / proposed / properties / term / descriptionPrevious value: -"The SUPPLIER's payment term as it should read, stored as written (do not round a real term to the nearest familiar one). Pass null to clear it. Correcting the term does NOT move the due date; if the due date is wrong too, propose `dueDate` alongside it."New value: +"The SUPPLIER's payment term as it should read, stored as written (do not round a real term to the nearest familiar one). Pass null to clear it. On an ON-ACCOUNT bill, changing to one of the five preset terms re-derives the due date from the bill date (+ 0/7/14/30/60 days), shown to the reviewer before approval; an unrecognised term moves nothing. To change the term but KEEP the current due date, send `dueDate` alongside it — a stated dueDate always wins."
- Changed
update_invoice_draft1 field changed- changed
Input schema / properties / proposed / properties / term / descriptionPrevious value: -"The payment term as it should read ('Due on Receipt', 'Net 7', 'Net 14', 'Net 30', 'Net 60', or whatever this business actually agreed — it is stored as written, so do not round a real term to the nearest familiar one). Pass null to clear it. ON A POSTED INVOICE THE TERM IS A PRINTED LABEL AND NOTHING MORE: correcting it does NOT move the due date, because the document has already been out of the building and its due date is what the customer was told. If the due date is wrong too, propose `dueDate` explicitly alongside it."New value: +"The payment term as it should read ('Due on Receipt', 'Net 7', 'Net 14', 'Net 30', 'Net 60', or whatever this business actually agreed — it is stored as written, so do not round a real term to the nearest familiar one). Pass null to clear it. On an ON-ACCOUNT invoice, changing to one of the five preset terms re-derives the due date from the invoice date (invoice date + 0/7/14/30/60 days), and the reviewer sees the new due date before approving; a term Taokeh does not recognise moves nothing. To change the term but KEEP the current due date, send `dueDate` alongside it — a stated dueDate always wins."
4 tool updates
- Added
bank_reconciliation_report - Added
create_bank_clearing_draft - Added
create_payment_match_draft - Changed
my_work1 field changed- changed
Input schema / properties / kind / enumPrevious value: -[ - "expense", - "invoice", - "bill", - "quote", - "receipt", - "payment", - "contact", - "contact_update", - "employee_update", - "product_update", - "recurring", - "credit_note", - "debit_note", - "purchase_order", - "journal", - "invoice_update", - "bill_update", - "logo_update", - "quote_update", - "delivery_order_update", - "purchase_order_update", - "credit_note_update", - "supplier_debit_note_update", - "expense_update", - "bank", - "staged_documents", - "staged_balances", - "staged_master" -]New value: +[ + "expense", + "invoice", + "bill", + "quote", + "receipt", + "payment", + "contact", + "contact_update", + "employee_update", + "product_update", + "recurring", + "credit_note", + "debit_note", + "purchase_order", + "journal", + "invoice_update", + "bill_update", + "logo_update", + "quote_update", + "delivery_order_update", + "purchase_order_update", + "credit_note_update", + "supplier_debit_note_update", + "expense_update", + "bank_clearing", + "payment_match", + "bank", + "staged_documents", + "staged_balances", + "staged_master" +]
1 tool update
- Changed
import_bank_statement2 fields changed- added
Input schema / properties / beginningBalanceDateAdded value: +{ + "description": "The date on the BEGINNING (opening) BALANCE line, YYYY-MM-DD — the start of the period. Every row must fall between it and statementDate. Omit it only if the statement prints none.", + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" +} - added
Input schema / properties / statementDateAdded value: +{ + "description": "The statement date printed on the statement (the ENDING BALANCE date), YYYY-MM-DD — it is the end of the period the statement covers, even when the last transaction is earlier. Omit it only if the statement prints none.", + "pattern": "^\\d{4}-\\d{2}-\\d{2}$", + "type": "string" +}
10 tool updates
- Changed
my_work1 field changed- changed
Input schema / properties / kind / enumPrevious value: -[ - "expense", - "invoice", - "bill", - "quote", - "receipt", - "payment", - "contact", - "contact_update", - "employee_update", - "product_update", - "recurring", - "credit_note", - "debit_note", - "purchase_order", - "journal", - "invoice_update", - "logo_update", - "bank", - "staged_documents", - "staged_balances", - "staged_master" -]New value: +[ + "expense", + "invoice", + "bill", + "quote", + "receipt", + "payment", + "contact", + "contact_update", + "employee_update", + "product_update", + "recurring", + "credit_note", + "debit_note", + "purchase_order", + "journal", + "invoice_update", + "bill_update", + "logo_update", + "quote_update", + "delivery_order_update", + "purchase_order_update", + "credit_note_update", + "supplier_debit_note_update", + "expense_update", + "bank", + "staged_documents", + "staged_balances", + "staged_master" +]
- Changed
revise_draft3 fields changed- changed
Input schema / properties / kind / descriptionPrevious value: -"Which kind of draft to amend — the same ten kinds the create_*_draft tools file. 'sales_debit_note' is the SELL-side debit note create_debit_note_draft files (an additional charge on an invoice you issued); 'journal' is the adjusting entry create_journal_draft files."New value: +"Which kind of draft to amend — the kind the filing tool returned (each create_*_draft and update_*_draft tool files one kind; 'invoice_update' / 'bill_update' are the posted-invoice and posted-bill corrections). 'sales_debit_note' is the SELL-side debit note create_debit_note_draft files (an additional charge on an invoice you issued); 'journal' is the adjusting entry create_journal_draft files." - changed
Input schema / properties / kind / enumPrevious value: -[ - "expense", - "invoice", - "bill", - "quote", - "receipt", - "payment", - "contact", - "contact_update", - "employee_update", - "product_update", - "recurring", - "credit_note", - "sales_debit_note", - "purchase_order", - "journal", - "invoice_update" -]New value: +[ + "expense", + "invoice", + "bill", + "quote", + "receipt", + "payment", + "contact", + "contact_update", + "employee_update", + "product_update", + "recurring", + "credit_note", + "sales_debit_note", + "purchase_order", + "journal", + "invoice_update", + "bill_update", + "quote_update", + "delivery_order_update", + "purchase_order_update", + "credit_note_update", + "supplier_debit_note_update", + "expense_update" +] - changed
Input schema / properties / patch / properties / proposed / descriptionPrevious value: -"contact_update / invoice_update: the changes, REPLACING the previous proposal entirely. On an invoice_update that includes `lines` — send every line the invoice should have, not only the corrected one."New value: +"contact_update / invoice_update / bill_update / quote_update / delivery_order_update / purchase_order_update / credit_note_update / supplier_debit_note_update / expense_update: the changes, REPLACING the previous proposal entirely. When it includes `lines`, send every line the document should have, not only the corrected one (on the quote / delivery-order / purchase-order corrections a line you are not changing can be sent as { keep: <lineNo> }; on the two note corrections every line carries goodsReturned)."
- Changed
search_document_lines1 field changed- changed
Input schema / properties / docType / descriptionPrevious value: -"REQUIRED. One of: 'invoice', 'credit_note', 'sales_debit_note', 'quote', 'delivery_order', 'bill', 'debit_note'. Note the two debit notes: 'sales_debit_note' is one this business ISSUED to a customer; 'debit_note' is the BUY side, against a supplier bill."New value: +"REQUIRED. One of: 'invoice', 'credit_note', 'sales_debit_note', 'quote', 'delivery_order', 'bill', 'debit_note', 'purchase_order'. Note the two debit notes: 'sales_debit_note' is one this business ISSUED to a customer; 'debit_note' is the BUY side, against a supplier bill."
- Added
update_bill_draft - Added
update_credit_note_draft - Added
update_delivery_order_draft - Added
update_expense_draft - Added
update_purchase_order_draft - Added
update_quote_draft - Added
update_supplier_debit_note_draft
1 tool update
- Changed
search_document_lines1 field changed- changed
Input schema / properties / docType / descriptionPrevious value: -"REQUIRED. One of: 'invoice', 'credit_note', 'sales_debit_note', 'bill', 'debit_note'. Note the two debit notes: 'sales_debit_note' is one this business ISSUED to a customer; 'debit_note' is the BUY side, against a supplier bill."New value: +"REQUIRED. One of: 'invoice', 'credit_note', 'sales_debit_note', 'quote', 'delivery_order', 'bill', 'debit_note'. Note the two debit notes: 'sales_debit_note' is one this business ISSUED to a customer; 'debit_note' is the BUY side, against a supplier bill."
Related MCP Connectors
Connect Claude or Cursor to books, invoices, bills, payroll, and sealed closes.
Open-source AI accounting skills verified by licensed accountants (tax, VAT, payroll).
AI staff accountant for QuickBooks: transactions, reports, receivables, payables, month-end close
QuickBooks Online in Claude and ChatGPT: 221 tools, full ledger, multi-company, Canada + US, FR/EN.
Related MCP Servers
- AlicenseBqualityBmaintenanceEnables AI assistants to query and manage self-hosted accounting data—invoices, balances, and books—through natural language, with read-only tools by default and optional scoped write operations.10MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to query a connected accounting platform for organisations, invoices, contacts, chart of accounts, and profit-and-loss reports, supporting bookkeeping, AR/AP, and financial reporting workflows.319 npmMIT
- FlicenseNot gradedqualityDmaintenanceGives AI models native-level control over TallyPrime ERP, covering 169+ tools across all functional modules including masters, vouchers, reports, GST, payroll, and more.-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to read and write Cynco accounting data, including querying books, creating invoices, reconciling transactions, and generating financial reports.4 npm1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.