Skip to main content
Glama
ComplyEaze

ComplyEaze Bridge: TallyPrime MCP server for Claude Desktop

Official

ComplyEaze Bridge

ComplyEaze Bridge is an MCP server that connects Claude Desktop to the TallyPrime running on your own computer. It is built for CA firms and accountants. You can ask about outstanding receivables and payables with ageing, the trial balance, ledger movement and vouchers. It checks ledger names before you post, and it prepares Journal, Payment, Receipt and Contra vouchers as a file. If you turn posting on in the extension, it posts them one at a time, after you approve each one.

Current release:

for Windows x64 and Apple Silicon Macs. We check each release before we publish it: the release check confirms that each package launches, lists its tools and parses a synthetic encrypted bank statement. It does not run against TallyPrime, and nothing we can run covers every Tally edition, set of books or setting. What has been run against a real TallyPrime, and what has not, is listed below. Install it · What changed · Security and privacy

Not yet code-signed; your computer may warn you before opening it.

Try asking (on a test company first):

  • "List the loaded companies."

  • "Show the outstanding receivables and payables, with ageing, as of 31 March."

  • "Show the trial balance for 1 April to 31 March."

  • "Check these ledger names against the book before I post: …"

How it handles your books

  • Local connection only. It talks to Tally's own XML gateway on a loopback address, so it cannot be pointed at a remote Tally host. It cannot tell whether that local port is forwarded to another machine; do not forward one across the internet. The Tally connection sends nothing to a ComplyEaze server.

  • Your AI provider sees what the assistant reads, just as it sees the rest of the conversation: company names, party names and amounts. You can mask party names or drop narration. Amounts are always sent. See Before you use it with client data below.

  • Posting is off by default in the extension. If you installed an earlier version, check the setting: an earlier default may still be saved as on. When you turn posting on, each voucher waits for your approval in a separate ComplyEaze Bridge window. No Bridge tool lets the assistant approve it, and an approval counts only when that window returns a fresh one-time token. After posting, the voucher is read back from Tally so you can see what landed. Three known limits remain: ComplyEaze Bridge cannot undo a posted voucher (you correct it in Tally); a company renamed to, or loaded under, the target company's name (or one differing only in case or spacing) just after its last check could still receive the voucher, if it has the voucher's ledgers, and ComplyEaze Bridge cannot always say where it went or prevent it; and a ledger renamed and replaced in that same moment could receive the entry, and not every such change is noticed. Read Before you turn on posting first.

  • Tool calls leave receipts in a log on your computer: the company's Tally identifier, and a fingerprint of what was asked and of what came back, written whether the call succeeds or is refused.

  • You accept the Terms of Use first. The extension asks you to accept the ComplyEaze Bridge Terms of Use (version 2026-10) in its settings, and every tool refuses with terms_not_accepted until you do.

  • Open source under Apache-2.0.

Related MCP server: Tally Prime MCP Server

What has been run against a real TallyPrime

Each line below is recorded in the repository or on the linked issue or pull request, unless marked as reported by the owner. Each ran on licensed TallyPrime Silver 7.1 and synthetic companies unless stated. The MCP guide and ADR 0004 hold the full record.

  • Reads: 27 checks on an unpublished macOS arm64 build (PR #228), recorded in the 6 September 2026 assessment. Some later reads were also run live on development builds, for example the trial balance on a debug build with synthetic companies (#246, whose record does not name the Tally release or licence tier). Not every read has its own recorded live run.

  • Posting one Journal, on a development build from 22 September 2026 (issue #579), and a Payment, a Contra and two Receipts (one of three entries) with the approval dialog on macOS (PR #600), each read back as posted. These builds predate the release published on 26 September 2026 (version 0.3.0).

  • Native posts of ten batches on licensed TallyPrime Gold 7.1, in one session on 28 September 2026, on a development build and one client book (the import request was captured for nine of them); their approval step was not recorded (protocol reference).

  • The published 0.3.0 package on Windows x64, in Claude Desktop (with no paid plan; we make no claim about other plans): reads only, against licensed TallyPrime Gold 7.1 with one client book, on 28 September 2026, in a session separate from the development-build posting above (reported by the owner; no logs were kept).

  • A candidate build of 0.4.0 on Windows 11, on 1 October 2026 (the build CI produced for the version pull request: the same code as the release apart from a comment in one test file). The maintainer installed it in Claude Desktop with a new Claude account that has no paid plan, against licensed TallyPrime Silver 7.1 holding one synthetic company. With the Terms setting off, a call was refused and nothing was read. With it on, tally_status, vouchers, validate_masters, purchase_register, stock_summary and local_data_report answered. One voucher post was declined in the Windows approval window and nothing was sent; one was approved, and one Journal was posted and then verified by verify_import. The same candidate's macOS build was started and read the company list, and in Claude Desktop on macOS its tools loaded in a chat. The record is the maintainer's dated notes and screenshots, kept privately. This was one run, not a controlled test of each key of the window.

  • A build of 0.4.1 on a Mac, on 2 October 2026. What was run: the package CI built for the version pull request, not the published file. The maintainer installed it in Claude Desktop over an installed 0.4.0, as an upgrade; the Terms setting and the other settings carried over, and the extension's server process started again on its own after the upgrade (its start time was read from the process list). tally_status and list_companies then answered against licensed TallyPrime Silver 7.1 holding the lab's own synthetic companies. The record is our dated notes, kept privately. What was not run: the published file is built again on another runner, and its program file differs from the one tested. We ran both builds without Tally: each reports version 0.4.1, lists 21 tools and gives the same answer to tally_status. We have not installed the published file in Claude Desktop.

  • A build of 0.4.2 on a Mac, on 3 October 2026. What was run: the package CI built for the release candidate, not the published file. The maintainer installed it in Claude Desktop on a Mac that had 0.4.1: it did not replace 0.4.1 but installed as a second extension (the author line changed in 0.4.2, and Claude Desktop builds an extension's identity partly from it); the settings did not carry over. tally_status and list_companies then answered against licensed TallyPrime Silver 7.1 holding the lab's own companies. The record is our dated notes, kept privately. What was not run: the published file is built again on another runner, and its program file differs from the one tested. We ran the published Mac file without Tally: it reports version 0.4.2, lists 22 tools with posting off (24 with it on), and refuses a call with the Terms setting off. We have not installed the published file in Claude Desktop, and nobody on our side has installed the Windows package in Claude Desktop on a Windows PC.

Not yet run by us in a controlled test: posting with a published package against a live TallyPrime; each way of declining in the Windows approval window (one was tried); the tools answering through Claude Desktop on macOS after the Terms are accepted (tally_status and list_companies answered on a CI build of 0.4.1 and again on a CI build of 0.4.2); posting on TallyPrime Education; posting on TallyPrime Gold with its approval step recorded. Each release package is built and launched, its tool list checked and a synthetic encrypted bank statement parsed, on hosted CI runners for Windows x64 and Apple Silicon Mac.

Not in the latest release

  • Stock quantities, and stock reads on books with many stock items; sales, purchase or tax posting; creating masters; bill-wise allocation

  • Deleting or undoing a posted voucher (correct it in Tally)

  • Reads on very large books can fail or take longer than the assistant waits (#485, #703)

  • A base currency other than INR. On a book with several currencies: the foreign-currency ledgers and vouchers themselves (they are set aside or withheld, and named), ledger movement, Profit and Loss and Balance Sheet, the purchase register, and posting

  • Tally Cloud Access or any remote Tally host

  • Intel Macs, and a code-signed installer

Is this for you

It is aimed at a practising accountant or a CA firm that already keeps client books in TallyPrime and wants to ask questions of them, or post entries into them, through an AI assistant such as Claude Desktop.

What it does today

  • Reads the loaded companies, ledger masters, trial balance, vouchers in a date window, outstanding receivables and payables, and ledger movement.

  • Checks ledger names before you post. Give it the names from a bank statement or an invoice and it reports which exist in the book and which are near-misses needing your decision. Reading the ledger list first is the single biggest cause of an import being rejected wholesale when it is skipped.

  • Records what it did. Every tool call Bridge runs — read or write, and whether it succeeds or is refused — appends a receipt to a log on your own machine, identifying the company it touched and fingerprinting what was asked and what came back. Reads keep those fingerprints as evidence alongside. A prepared batch records the local endpoint it was built for, and a native posting is refused if that endpoint has changed since; that is a safety check kept in Bridge's internal import ledger, not a line in the proof report a reviewer opens. A reviewer can read the log rather than take a summary on trust.

Before you turn on posting

Whether writing is on depends on how you installed it. Everything above is reading. When writing is off, the write tools do not merely refuse — they are absent from the tool list entirely, so an assistant cannot see that they exist.

  • The Claude Desktop extension turns voucher posting off by default. Three known limits in posting remain. Tally aims an import at a company by its name and cannot bind it to a company's GUID. Bridge's last request before the post checks that exactly one loaded company has the target's GUID and name, and that no other loaded company has the same name ignoring case and spacing; otherwise it refuses the post (#607). A company renamed to, or loaded under, the target's name (or one differing only in case or spacing) in the moment after that check could still receive the voucher, if it has the voucher's ledgers. Bridge may flag afterwards that the loaded companies changed, but cannot always say where the voucher went, and cannot prevent it (accepted residual, #574). A ledger renamed and replaced in that same moment means the post can land in the replacement ledger. Bridge marks the result as needing reconciliation when it sees that the ledger now resolves to a different master; a change that leaves the company's master mark unmoved, or is reverted before that check, is not seen, and a regroup in that moment is not detected (#623). And Bridge has no tool to delete or undo a voucher it has posted, so a wrong post must be corrected by hand in Tally. It records the REMOTEID each post sends, but no delete tool exists yet (#579, #582). Turning on Allow voucher posting (Journal, Payment, Receipt, Contra) in the extension settings adds post_import, which posts one saved voucher of those types; every posting still waits for your approval in a separate Bridge dialog. Leave it off unless you accept those risks. Voucher file preparation and bank-statement parsing, which write nothing to Tally, stay available with the setting off. If you installed an earlier version, check the setting: an earlier default may still be saved as on.

  • A source build turns writing off by default. Preparing a file needs BRIDGE_AGENT_ENABLE_IMPORT; posting additionally needs BRIDGE_AGENT_ENABLE_WRITES, which grants both.

  • A source build also needs BRIDGE_TERMS_ACCEPTED=true. The extension asks for that as its "I accept" setting; without it every tool refuses.

With writing on:

  • Prepares vouchers as a local file — Journal, Payment, Receipt and Contra. Bridge writes the file; it does not send it.

  • Posts one saved Journal, Payment, Receipt or Contra per approval, and only after you approve that exact voucher in a dialog on your own machine. The assistant cannot approve it. Bridge then reads the voucher back so you can see what actually landed. You can also import a prepared file through Tally yourself; verify_import then reads that back.

  • Posting has limits. It creates no masters, posts no sales, purchase, tax or inventory entries, and never alters or deletes a voucher. A company with more than one currency defined is refused.

What it does not do

  • The Tally path uploads nothing to ComplyEaze. Bridge reads it over a local connection and hands it to the assistant you are talking to; nothing in the Tally path sends it to a server of ours.

  • It will not post anything without a separate, explicit step after the file is prepared.

  • It is not a Tally replacement, a reporting suite, or a filing tool.

The desktop app

The extension is built from the same source library as the desktop app. Packages up to 0.3.0 contained an unfinished document-upload feature and an AXAL sign-in, which no tool of the extension reached. That code was removed (#914) and releases from 0.4.0 on do not contain it; the only network client in ComplyEaze Bridge's own code connects to Tally on your own computer. No desktop installer is published. See Security and privacy.

Before you use it with client data

One thing to understand before you use it. When you ask an AI assistant for financial data through Bridge, the assistant's provider sees what it reads — company names, party names and amounts. That is a property of using a hosted assistant, not of Bridge. Bridge can mask party names or drop narration first (BRIDGE_AGENT_REDACTION), but neither setting removes amounts — figures always go with the answer. Decide this deliberately for client data.

What it costs

ComplyEaze has not set a price for ComplyEaze Bridge and does not sell licences to it; there is no account with us, subscription or licence key. The code of a release you download stays under the licence it was published with (Apache-2.0 for current releases). We have not decided whether to charge for anything in future.

What you pay or provide today:

  • TallyPrime: your own licence.

  • Claude Desktop: Anthropic's plans. On 1 October 2026 we ran a candidate build of 0.4.0 (not the published file) on Windows with a Claude account that had no paid plan, on one synthetic company; we make no claim about other plans or larger books.

  • Your clients' data: what Claude reads goes to Anthropic, under Anthropic's terms for your plan; through ComplyEaze Bridge, ComplyEaze does not receive it (Privacy Policy, sections 4 to 6).

  • Your checking: check results in Tally before you rely on them.

Support and updates are not guaranteed (Terms of Use, section 4.4). Our liability is limited as section 14 sets out, including its fallback and exceptions; read it before client work.

Installing it

Before you install, turn on Tally's HTTP gateway. TallyPrime does not listen for ComplyEaze Bridge by default. In Tally's own connectivity / client-server configuration settings, set Tally to act as a server ("acts as Both" in Tally's own words) and note its HTTP gateway port — 9000 by default, but configurable. To check it is actually on, open http://localhost:9000/status (substitute your port) in a browser: a running gateway answers with a short Tally XML response, and a browser that cannot connect means the gateway is still off — unless Tally is running in a Windows virtual machine on a Mac, in which case run this check inside that VM, or only once your local forwarding is working. A Mac browser that cannot connect may mean the forwarding described below is missing rather than that the gateway is off. If instead it hangs without answering, Tally may simply be busy behind another request — wait and retry rather than changing the setting.

The latest published release of the Claude Desktop extension is the one to install. Follow the installation guide to install and configure it. Before you do, know what it is and is not:

  • Bridge is still being developed. A release may contain errors, so try it on test data first and keep current backups. It is not yet code-signed or notarized, so your operating system may warn before opening it. Each package has a .sha256 file and a provenance record so you can confirm exactly which bytes and which source commit you downloaded.

  • Checked only as far as launching. The release build confirms the package starts and lists its tools. It does not establish that it works against your Tally, or in conversation inside Claude Desktop. What has been run against a real TallyPrime, and on which builds, is listed above; the published 0.4.2 package itself has not been run by us against a live TallyPrime.

  • Windows x64 and Apple Silicon Macs only. Intel Macs are not supported.

  • On a Mac, Tally must run on that same Mac, in a local Windows virtual machine or through approved local forwarding. Bridge only talks to Tally on your own computer, so a separate PC or a Tally elsewhere on your network cannot be reached by typing its address.

  • It does not update itself. To upgrade, install a newer release from Claude Desktop's Extensions settings. Release 0.4.2 installs beside an older release instead of replacing it (seen on a Mac; not tried on Windows): remove the older extension first, do not delete the data folder, and enter your settings again, including Response redaction, which starts at none.

The Bridge desktop application is a separate program and has no published installer; building it from source is described under Contributor quick start below.


The rest of this file is for people working on Bridge. The repository is self-contained: build and development commands resolve files relative to the clone, not to a developer-specific directory. It holds a Tauri desktop application and the MCPB packaging path for Claude Desktop, with React/TypeScript and Rust components for Tally and local database operations.

First useful result

Install the latest published release of the Claude Desktop extension with the installation guide. For source use, the contributor quick start below builds the desktop app; to run the MCP server from source, follow the source MCP setup.

Before requesting financial data through an MCP client, the client may send the selected Tally result to its AI provider, including company identity, party or open-bill details, and amounts. Source installations default to BRIDGE_AGENT_REDACTION=none; set it to mask_parties or drop_narration before launch when that better fits the workflow. These settings mask party names or drop narration; they do not remove amounts. The package installation settings expose the same choices.

For a first result, run tally_status to check that TallyPrime and its Licensed or Education mode are observed, then list the loaded companies. Select a company with exactly one observed INR currency master and request receivables or payables. In Education mode, explicitly supply an as_of date on day 1, 2, or 31; an omitted date defaults to today and may be refused. Bridge rechecks product, mode, dates and currency for the financial read; other or unobserved products/modes, no currency master, non-INR, or multiple currency masters are refused.

The desktop also offers local XML draft preparation through Prepare file. It preserves source observations beside editable proposals and saves a local draft for later review.

For contributors, use the setup and development path below.

Supported development hosts

Bridge is intended to build and run on Windows and macOS 12.4 or later. Run platform checks on a native host for each operating system; a successful build on one operating system does not verify the other.

Shared prerequisites:

  • Node.js 24 (>=24.15.0) and Corepack (.node-version pins the CI baseline)

  • the Rust toolchain pinned by rust-toolchain.toml

  • Perl 5 with Locale::Maketext::Simple for the bundled SQLCipher/OpenSSL build

  • LLVM/libclang for SQLCipher binding generation (LIBCLANG_PATH may be required)

  • the operating-system dependencies listed in the Tauri prerequisites

On Windows, install the Microsoft C++ build tools, WebView2 components, and a complete Perl distribution such as Strawberry Perl. If another incomplete perl.exe appears first on PATH, set OPENSSL_SRC_PERL to the complete Perl executable. Install LLVM as well; if libclang.dll is not discoverable, set LIBCLANG_PATH to its directory (commonly C:\Program Files\LLVM\bin). On macOS, install Xcode Command Line Tools. Bridge's macOS bundles require macOS 12.4 or later.

Contributor quick start

Run these commands from the repository root in PowerShell, Command Prompt, or a POSIX-compatible shell:

corepack pnpm install --frozen-lockfile
corepack pnpm exec playwright install chromium webkit
corepack pnpm test
corepack pnpm run build
corepack pnpm run cargo:check
corepack pnpm run tauri:dev

pnpm test includes Chromium and WebKit evidence-drawer focus suites; installing the lock-pinned browsers after dependencies is therefore required once for each developer environment. tauri:dev starts the Vite development server and desktop application. It does not require a fixed checkout location. The first Rust build can take several minutes.

For a release build, run corepack pnpm run tauri:build on each target host. CI-produced bundles are unsigned smoke artifacts only. Do not redistribute a desktop installer until the signing, notarization, provenance, and rollback gates in the release runbook are complete.

Platform verification

Before claiming support for a platform, run the following on that platform:

corepack pnpm install --frozen-lockfile
corepack pnpm run build
corepack pnpm run cargo:check
corepack pnpm run tauri:build

Also manually exercise the affected Tally workflows. Vendor integrations may require host-specific software even though repository paths and project commands are portable.

Integration trust boundaries

Bridge restricts native network and file access even if the renderer is compromised:

  • Tally connections are loopback-only (localhost, 127.0.0.0/8, or ::1). Remote plaintext Tally hosts are intentionally rejected.

  • Showing an export in the file manager accepts only a file ComplyEaze Bridge exported since it started; any other path the renderer sends is refused before anything is launched.

Privacy and safe diagnostics

Do not commit or attach real customer, company, tax, certificate, credential, financial, or document data. Before sharing logs, screenshots, fixtures, or reproduction steps, replace personal and customer data with synthetic values and remove local usernames and absolute paths. See SECURITY.md for private reporting and handling requirements.

Repository map

  • src/ - React UI and API bindings

  • src-tauri/ - Rust core and Tauri configuration

  • docs/ - architecture, roadmap, and operational guidance

  • .github/ - issue and pull-request templates plus CI configuration

Governance

License

Bridge is licensed under the Apache License, Version 2.0. Attribution notices are provided in NOTICE. The ComplyEaze logo and icon files are not licensed under Apache-2.0; see NOTICE and TRADEMARKS.md. The historical v0.1.0 release remains under the MIT license shipped with that tag; current development source is version 0.4.2 under Apache-2.0.

Available Tools

20 tools
balance_sheetA
Read-only

Return the Balance Sheet for a date range, derived from Tally's native Trial Balance: one line per reserved Balance Sheet primary group (Capital Account, Loans (Liability), Current Liabilities, Suspense A/c, Branch / Divisions, Fixed Assets, Investments, Current Assets, Misc. Expenses (ASSET)), each the signed closing balance at to (a debit negative), and result.profit_and_loss: the Profit & Loss A/c ledger's closing and the carried result (that closing plus every P&L ledger's closing; in a part-year window the year's earlier result sits in that ledger). The top-level lines is null while carried is not established, so a derived line is never shown as the statement. Reads the Trial Balance, the group tree and Tally's own Balance Sheet inside one company, mode and book-extent bracket; requires observed INR currency (a book with more than one currency master is refused) and supported date boundaries. Each line reports amount (sum, present_count, empty_count): the sum is over the amounts Tally returned, and the empty ones it left out are counted. A result is established only if every ledger is classified (a ledger under a user-created primary group, or whose group chain is incomplete, is listed in unclassified, up to 100, with unclassified_total), no Stock-in-Hand ledger carries an amount, and Tally's own Balance Sheet for the window ties line for line to the derived one (balance_sheet_gate). Otherwise it is not_established with a reason (unclassified_ledger_carries_an_amount, closing_stock_not_derivable_from_trial_balance, profit_and_loss_ledger_not_returned, tally_balance_sheet_differs, or for gross and net tally_profit_and_loss_differs); for the two differs reasons, lines names the lines that did not tie. The top-level state is observed only while carried is established, and not_established otherwise with the same reason the nested result carries. A book with stock items is expected to refuse; no inventory book has been measured. A Tally line with an amount the derivation has no counterpart for, such as a heading or a difference in opening balances, refuses rather than being guessed at. Tally's own statements carry no company identity and are bound only by the checks around the read. The gate has been measured over one full year on one book and one month on another; a window spanning more than one financial year is unmeasured. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
company_guidYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds substantial context beyond annotations: it explains the gate mechanism (balance_sheet_gate), refusal reasons, the unclassified ledger limit, stock item behavior, logging of receipt lines, and that nothing is written to Tally. This is exemplary transparency far beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely dense and long, essentially a specification document embedded in the tool description. While information-rich, it is not front-loaded or concise; it reads as a single sprawling paragraph covering derivation, states, gates, and logging, making it hard for an agent to quickly extract the essential call information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool and no output schema, the description is largely complete: it explains return structure (lines, amount with sum/present_count/empty_count, result.profit_and_loss, state), refusal conditions, and data sources. It is thorough, though its density and lack of structure slightly reduce usability.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the semantics of `to` as the closing-balance date and `from`/`to` as a date range window, and explains that company_guid bounds the read to one company. While the schema provides patterns and types, the description adds conceptual meaning (date boundaries, currency constraints) that the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the Balance Sheet'), names the data source (Tally's native Trial Balance), and enumerates the exact primary groups returned. This clearly distinguishes it from siblings like trial_balance and profit_and_loss, which it references by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through preconditions (observed INR currency, single company/mode book, supported date boundaries) and references to sibling concepts (Trial Balance, P&L A/c), but there is no explicit when-to-use-this-vs-alternatives guidance. The description explains the derivation rather than choosing between tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

egress_logC
Read-only

Everything the assistant reads from Tally through ComplyEaze Bridge in this chat, amounts included, is sent to the AI provider; redaction can only mask party names or drop narration. Never tell the user that no data has left their computer. Shows the receipts ComplyEaze Bridge keeps of its own tool calls, read from its local log file on this computer: tool, time, company, counts and fingerprints, no figures or book content (bounded to the most recent receipts). It does not show what the AI provider received. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, indicating no side effects. The description states 'Each call appends metadata-only receipt lines ... to ComplyEaze Bridge's local log on this computer,' which is a write-side effect and directly contradicts the readOnlyHint annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and includes critical security context, but it front-loads warnings rather than the tool's purpose, and it repeats the local log concept. It is adequately sized but could be better structured for quick tool selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and one parameter, the description does describe the returned fields (tool, time, company, counts, fingerprints) and privacy limitations. However, it omits any explanation of the 'limit' parameter and contains an annotation contradiction, leaving gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, 'limit', with 0% schema description coverage. The description only says results are 'bounded to the most recent receipts' but never mentions the 'limit' parameter or how to control the number of receipts returned, leaving the parameter semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool shows receipts of ComplyEaze Bridge's own tool calls, read from a local log, with specific fields listed. It distinguishes itself from Tally data tools by being an audit/egress log rather than a Tally query. However, the actual purpose is buried after two warning sentences, slightly reducing clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool shows and includes a communication warning ('Never tell the user that no data has left their computer'), but it never states when to call this tool versus alternatives, nor does it name any sibling tool for comparison. Usage context is implied at best.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_mastersA
Read-only

Return a company's ledger masters with each one's opening balance as of the start of the books (opening_balance, opening_balance_as_of), on a freshly observed supported product and mode. fields=basic (the default) or fields=compliance, which adds paired party-master observations (GSTIN, PAN, MSME, bank, contact and address), each ledger's group ancestry and its gst_duty_head. GSTIN (compliance): party_gstin is the GSTIN in force on party_gstin_as_of, which is the optional as_of (YYYYMMDD or YYYY-MM-DD, such as 20260331 for a year end) or else this computer's date; as_of without fields=compliance is refused as ledger_masters_as_of_requires_compliance. It comes from the ledger's dated registration history, or from the flat GSTIN field only when that history is empty or was not returned; an empty flat field names no GSTIN. party_gstin_status names the source: in_force, flat_field, no_gstin_in_force (a history with no GSTIN on that date; party_gstin_registration_type says whether that entry is registered) or not_reported. history_unreadable fails closed: the history came back undated, misdated, malformed, repeated or contradictory, so party_gstin is null and the flat field is not used. party_gstin_flat is always the flat field as read, and gstin_sources_disagree is true when it names a GSTIN that the in-force history entry does not; both are reported, neither is chosen. For another date, such as a transaction's, pass it as as_of or read the dated entries in compliance.gst_registrations. Duty head (compliance): gst_duty_head is recognized (with head: cgst, igst, state_tax, sgst_utgst, ut_tax or cess, and raw, the spelling Tally returned), unrecognized (raw kept), not_tax_ledger, contradictory or absent. state_tax (raw State Tax) and sgst_utgst (raw SGST/UTGST) are two spellings Tally has returned for a state-side head and are kept as two heads: a consumer summing state tax must include both. Ancestry (compliance): chain (nearest group first, each hop's own name and reserved_name), complete (true only if the chain was resolved all the way to the reserved account root) and gap (null when complete, else why resolution stopped: no_parent, group_absent, group_name_repeated, reserved_name_missing, cycle or exhausted). An incomplete chain is never padded or guessed: chain is exactly what was resolved, so check complete before treating it as exhaustive. A reserved_name beginning with U+FFFD #4; is a Tally reserved value (Tally writes it as ); U+FFFD#4; Primary is the account root, distinct from a group a user named Primary. Group filter: group filters by group name. group_scope "immediate" (the default) matches only the ledger's own parent and does NOT include ledgers under sub-groups of group; "ancestry" matches any group in the resolved chain, the whole subtree (a ledger under Bank OD A/c matches Loans (Liability)), with either fields value, and a gap in a chain never counts as a match. Any group filter reads the group collection (with fields=basic, one added paired read), and the result carries group_filter: excluded_subgroup_ledgers (count of ledgers left out because they sit under a sub-group of group, always 0 under ancestry scope, with group_count and up to 20 of those names in groups) and unresolved_ancestry_ledgers (ledgers not returned whose chain stops before reaching group, so ComplyEaze Bridge cannot say whether they belong under it; counted over the whole book, so a gap anywhere is counted). Large books (compliance): when the master-alteration mark (an upper bound on the ledgers, since every master raises it) puts the estimated response over budget, the ledgers are counted first, then read whole or, if the count does not fit, in parts by parent group; every ledger counted must come back exactly once, or the whole call is refused. Above a mark of 22,857 the ledgers are counted by AlterID span, in slices of at most 4,000, one request each for GUIDs only. Above 400,000 the call is refused before any ledger read, with cause ledger_catalogue_too_large and size, so a company with fewer ledgers may be refused. Each refusal names its cause: parent_over_budget, parent_partition_too_many_parts, parent_complement_over_budget, ledger_without_parent, parent_name_unsupported (with unsupported_parent_ledgers), parent_partition_duplicate_ledger_identity, the parent_part_* coverage causes, parent_part_response_too_large, ledger_span_slice_over_bound, ledger_span_duplicate_identity, ledger_span_census_empty, ledger_span_slice_malformed, ledger_span_identity_mismatch, ledger_span_slice_response_too_large or ledger_count_catalogue_too_large. Retrying a size refusal refuses again, and fields=basic still reads the book. For ledger_count_differs (two counts, or a count and the ledgers read, disagree) or ledger_count_company_differs (Tally's own ledger count is higher than the census's), retry once while the book is quiet; ledger_count_company_invalid means that count's answer was damaged, and ledger_count_company_response_too_large that it was larger than the response limit: retry once, then use fields=basic. A counted read reports ledger_count_cross_check.status: matched, company_count_lower, or unavailable when Tally's answer carried no count, so the check did not run. Currencies (compliance): a book with several Currency masters is read through the base Tally identifies: its plain base-currency ledgers are returned with ledgers_scope base_currency_ledgers_only, and the ledgers kept in another currency (foreign_currency_ledgers_excluded) and the base-currency ledgers whose balances Tally shows in another currency (base_currency_ledgers_mixed_excluded) are named, never read. Paging: a first page (offset 0) always reads Tally afresh and holds the read; a later page is served from it while the book extent, including ALTMSTID and ALTVCHID, is unchanged, and each result reports snapshot. Pass the first page's snapshot_id on later pages to have the call refused as listing_snapshot_changed instead of continuing from a different read (cause book_changed_since_first_page, or snapshot_not_held for an id not held). With fields=compliance, repeat the first page's as_of on later pages: a snapshot serves only pages read as of the same date, so a later page without it (or across midnight) reads afresh, or is refused when it names the snapshot. A change that moves neither mark is not seen. Each screen action measured so far moved a mark (a regroup, an opening change, a ledger create or delete, a voucher delete, a cancel, a save with no change; protocol reference section 11c.5, one run each), but a change that moves neither can leave a later page up to 10 minutes old. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofNo
groupNo
limitNo
fieldsNobasic
offsetNo
group_scopeNoimmediate
snapshot_idNo
company_guidYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/non-destructive, but the description goes far beyond them: it enumerates every refusal cause, the count-then-read strategy for large books, the 400,000 mark refusal, currency handling, snapshot staleness ('up to 10 minutes old'), and states explicitly that it writes nothing to Tally and only appends metadata-only receipt lines locally. Nothing here contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in sentence one, which is good, but the body is a very long, densely nested wall of text mixing paging, counting, currency, GST and refusal semantics. Much of it is genuinely informative, yet the sheer size makes it hard to scan and pushes it past what a tool description should reasonably hold.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 parameters, no output schema and 0% schema coverage, the description is remarkably complete: it documents the response fields, refusal causes, snapshot semantics and scope flags an agent would otherwise have to discover by trial. An agent could call this correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does so richly: `as_of` format and date semantics, `fields` basic/compliance differences, `group` filtering, the full immediate-vs-ancestry meaning of `group_scope`, `offset` first-page behavior, and `snapshot_id` propagation. Only `limit` and `company_guid` are left unaddressed, which is a minor omission against an otherwise exhaustive treatment.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource ('Return a company's ledger masters') plus the exact scope (each ledger's opening balance as of the start of the books) and names the concrete output fields. Combined with the name `ledger_masters`, an agent can tell this apart from `masters`, `validate_masters` and `ledger_movement` without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context: when to pass `as_of` (e.g. a transaction's date) versus reading `compliance.gst_registrations`, when `as_of` is refused, when to reuse `snapshot_id`, and retry guidance for count mismatches. It never explicitly compares itself to sibling tools like `masters` or `validate_masters`, so it stops short of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_movementA
Read-only

Return literal-window ledger opening, exact debit/credit movement, closing, and touched-voucher count with a freshly observed supported product/mode and an operation-valid opening boundary. Reads the full voucher window before filtering or pagination; use narrow dates. Dense windows can fail source limits. Requires one observed INR currency master: a book with several Currency masters, none, or one that is not INR is refused before any ledger read. A ledger name resolves only when spelled as in the book or differing from it only in ASCII case and spaces; ledger_match names the ledger read, how (matched: exact, or case_or_spacing, which the answer should mention by naming the ledger read) and any similar_ledgers that differ from it only in case or whitespace. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
limitNo
ledgerNo
offsetNo
company_guidYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds substantial behavior beyond them: INR currency-master requirement (refuses multi/none/non-INR books), full-window read before pagination, source-limit failure mode, ledger-name matching semantics, and the fact that it appends metadata-only receipt lines to a local log while writing nothing to Tally. This is unusually rich disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the deliverable, and nearly every clause carries operational information (currency precondition, date-window behavior, logging side effect). It is a single dense block that would read better as short sentences, but little of it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter read tool with no output schema, the description describes the return shape (opening, movement, closing, voucher count, ledger_match with match mode, similar_ledgers), preconditions, failure modes, and side effects. An agent has what it needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It explains the ledger name resolution rules and the from/to window behavior, but says nothing about limit, offset, or company_guid semantics, leaving half the parameters undocumented. It partially compensates but leaves clear gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-and-resource: returns literal-window ledger opening, debit/credit movement, closing, and touched-voucher count. It is clearly a ledger-movement read, distinguishable from sibling list/aggregate tools like trial_balance or vouchers, though the phrasing is dense enough that the core purpose is buried in the first clause.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage guidance: reads the full voucher window before filtering or pagination, so 'use narrow dates', and warns dense windows can fail source limits. It does not name alternative sibling tools for related queries, so it stops short of explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_companiesA
Read-only

Start here. Return observed company tuples (the company_guid every tool that reads a company's books needs) and identity ambiguity flags. If exactly one company is open and the user named no client, use it and say which in your first line; if the user names a client and exactly one open company matches that name, use it and say which; otherwise ask which, offering the list, and never guess a company. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, not open-world), and the description adds non-obvious behavior: every call appends metadata-only receipt lines (tool, company, counts, fingerprints) to a local log, and it writes nothing to Tally. This is exactly the extra context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with 'Start here' and organized around the routing decision, so the important guidance comes first. It is dense and the trailing logging sentence adds length, but each sentence carries load-bearing information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description must carry both selection logic and return-content expectations, and it does so, including the company_guid contract and ambiguity flags. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to disambiguate and the baseline is 4. The description instead front-loads what the call returns (tuples with company_guid and ambiguity flags), which is useful but belongs to return semantics rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Return observed company tuples') and explains the payload's role by naming the exact field other tools need (company_guid). 'Start here' explicitly positions it as the entry point, distinguishing it from the many sibling read tools that require a company.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit decision rules: single open company with no named client -> use it; named client matching one open company -> use it; otherwise ask and offer the list. It also states an exclusion ('never guess a company'), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

local_data_reportA
Read-only

Reports what ComplyEaze Bridge stores locally, by class, with counts, sizes, the age of the oldest file and the state of its import journal: how many batches were sent or found posted, how many of those are not settled, and how many have no recorded dispatch and were never verified as posted (no_dispatch_never_verified: this can include a batch imported by hand, which may be in Tally, so never treat it as proof that a batch is absent). Reads ComplyEaze Bridge's own local data folder and the per-user folder of dispatch lease locks (names, sizes and times) and names no file path. A folder, link or journal it could not read or enter is reported as such (incomplete_reason and folders_that_could_not_be_listed, and this call's evidence is partial), never as empty or absent. It covers the folder the MCP server and the desktop Journal flow share; the desktop app's other settings, its mirror database and logs live elsewhere and are not covered. The import journal and the imports folder are ComplyEaze Bridge's memory of what it already sent to Tally: never suggest deleting them. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses that unreadable folders are reported as partial evidence rather than empty, that no_dispatch_never_verified is not proof of absence, that the journal/imports folder must never be suggested for deletion, and that each call appends metadata-only receipt lines locally (consistent with idempotentHint=false). It also states explicitly that nothing is written to Tally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the opening clause and every sentence carries substantive constraints. The single dense paragraph with stacked caveats is heavier than it needs to be, but little is genuinely wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the fields an agent will encounter (incomplete_reason, folders_that_could_not_be_listed, no_dispatch_never_verified) and by delimiting what is and is not covered. Complete enough to call and interpret correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description usefully notes that it reads fixed local folders and names no file path, which is the only parameter-adjacent information an agent needs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — reporting what the local data store holds, by class, with counts, sizes, oldest-file age and import-journal state. It is clearly distinguished from siblings like tally_status and verify_import, which concern Tally-side posting rather than local storage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It defines scope boundaries (what folders are covered vs not) and warns how to interpret no_dispatch_never_verified, which is real usage guidance. But it never states when to reach for this tool instead of verify_import, read_evidence or tally_status, so the alternative-selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mastersA
Read-only

List a company's masters of one kind: voucher types with their numbering method (Automatic, Manual or Default, as Tally reports it) and active and optional flags, godowns, units (with decimal places and whether simple), stock groups, or account groups (groups). Each row of another kind carries name, guid, master_id, alter_id and parent (null where it does not apply or is absent; a parent that is Tally's reserved root is written as U+FFFD #4; Primary, as in the group tree); a groups row carries name, parent and reserved_name only. Voucher types also carry active, optional (true, false or null when Tally did not say) and numbering_method (automatic, manual, default, {"unrecognised": <raw text>} for any other value Tally reports, or null when absent); units carry decimal_places and simple. One kind per call, read inside one company, mode and book-extent bracket: the extent before and after must be equal and each collection is read twice and compared, so a book that changed during the read is refused. Whole reads only. Godowns, units and stock groups are read only when the book's master-alteration mark times an assumed worst-case row for that kind fits 16 MB (the mark counts the masters of every kind, so a book with few of this kind can be refused), and otherwise refuse before any request with masters_too_large and size (master_alter_id, estimated_bytes, limit_bytes, limit_master_alter_id); that admits marks up to 1,152 for godowns, 1,168 for units and 1,160 for stock groups. Retrying refuses again. The refusal's size carries limit_master_alter_id, the largest mark this kind is read at. Both stock-heavy client books measured, with marks of about 100,000 and 300,000, refuse these three kinds; how common such marks are across live books is unmeasured. voucher_types and groups have no size check before the read (voucher types keep the policy of ComplyEaze Bridge's other voucher-type read). After a read of any kind other than groups, each row's length is checked against the assumed worst-case row (masters_row_exceeds_bound), and the row count, the rows' AlterIDs (at or under the mark, none repeated) and the response size (each collection is read twice, and the size check runs after both reads) are checked against the mark and the admitted size; a breach of those three refuses the whole read as masters_bound_premise_violated and returns no partial list, unless the closing extent shows the book moved, which is reported instead (masters_extent_changed). A response ComplyEaze Bridge cannot read refuses at once with a masters_* cause, and a voucher_types answer with no rows refuses as masters_voucher_types_empty, because every company has predefined voucher types. Not returned: alias names, counts and hints (no company NUM* fields), stock items, ledgers (use ledger_masters), and any write. The row shape was captured from one synthetic book on one licensed TallyPrime 7.1: Default, Automatic and Manual are the only numbering methods seen, and any other value is returned raw, not refused. default is Tally's reported value, not evidence that a type numbers automatically. The company's voucher-type count (NUMVOUCHERTYPES) did not equal the rows returned on two books (35 vs 26, 33 vs 24), and on one book equalled the number-series count (inferred to count series, unmeasured), so do not check these rows against it. The completeness of the voucher-type list is unverified: absence from it is not evidence that a voucher type is absent from the book. Under mask_parties, godown and stock-group names and their parents are masked, because a job-work godown or a supplier-named stock group can carry a party's name (Tally's reserved root as a parent is left as it is). Voucher-type, unit and account-group names are not masked: they are configuration labels, not counterparties. Education mode is refused. A first page (offset 0) always captures a fresh read and holds it in memory; a later page for the same kind (offset > 0) is served from it while the company's book extent, including ALTVCHID and ALTMSTID, is unchanged, at the cost of one small extent read. Each result reports snapshot (id, master_alter_id, voucher_alter_id, read_at, reused). Pass the first page's snapshot_id on later pages to have the call refused with listing_snapshot_changed (cause book_changed_since_first_page or snapshot_not_held) instead of continuing from a different read. A change that moves neither mark is not seen, so a later page can be up to 10 minutes old after such a change. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
limitNo
offsetNo
snapshot_idNo
company_guidYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: it discloses the double-read consistency model, the 16 MB pre-read size refusal with concrete master_alter_id limits per kind, the row-count/AlterID/size premise checks, mask_parties masking behavior, snapshot reuse/refusal semantics, and metadata-only local receipt logging. None of this is derivable from readOnlyHint/destructiveHint/idempotentHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries real information, but it is delivered as one dense, unstructured block with no headings or bullets, and several caveats (unmeasured NUMVOUCHERTYPES comparison, 10-minute staleness) are buried mid-paragraph. Front-loading the purpose is good; the rest is hard to parse for an agent that needs the key operational rules quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of describing return shapes per kind, the snapshot object, and error causes, and it does so thoroughly. For a tool this complex and this side-effect-sensitive, nothing an agent needs to call it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates fully: it defines the kind enum values, explains that offset 0 takes a fresh read while offset > 0 is served from the in-memory snapshot, and specifies how snapshot_id must be passed and what happens when it is stale or absent. company_guid's scoping ('read inside one company, mode and book-extent bracket') is also clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('List a company's masters of one kind') and enumerates exactly which kinds are available, including the row shape each kind returns. It also explicitly names the sibling it is not (ledgers -> ledger_masters), so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear routing guidance ('ledgers (use ledger_masters)') and states exclusion conditions (whole reads only, education mode refused, one kind per call, one company). It does not, however, contrast with other master-related siblings such as validate_masters or voucher_schema, so the when-not guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

outstandingsA
Read-only

Answer what is outstanding: receivable and payable totals from Tally's own paired bills reports, ageing, the top parties and the open bills, on a freshly observed supported product and mode and a date valid for the operation. as_of is optional: left out, it is this computer's date (tally_status today), and the date used is always returned as result.as_of, whatever the state. top ranks parties only; page the bills with offset and limit. open_bills_total counts every open bill in the requested direction, all of which totals and the ageing cover, and open_bills_shown counts the bills on this page (limit is not lowered when the response size shortens the page): when shown is less than total, say so plainly, for example "showing 500 of 1,240 open bills; the totals and the ageing cover all 1,240" (on a later page, the bills from offset + 1; on a partial read, both counts and the sentence cover the base-currency ledgers only, so say so in it), and ask for the next bills with offset set to next_offset. receivable and payable follow the sign of each bill's balance, as Tally's own Bills Receivable and Bills Payable reports scope them, not the type of party: a customer's advance or a credit note raised to a customer appears under payable, and a supplier's advance or a debit note raised to a supplier under receivable, because those reports carry no bill type. That holds for an advance or a note kept as its own bill: an on-account advance goes to the unallocated figure instead, and a credit note set against an open invoice reduces that invoice. Measured on one synthetic book (TallyPrime Silver 7.1). Read a bill's kind as a direction, not as owed by a customer or owed to a supplier. An unallocated amount's direction is the sign of the party's net unallocated balance, so an on-account receipt and an on-account payment on one party net into one figure. A book with several Currency masters is read through the INR base Tally identifies. If it has ledgers kept in another currency (foreign_currency_ledgers_excluded, each with its currency) or base-currency ledgers whose balance Tally shows in another currency (base_currency_ledgers_mixed_excluded, each set aside with all its bills), the state is partial with partial_reason currency_ledgers_excluded and partial_reasons naming the lists that are not empty: figures cover its base-currency ledgers only, under base_currency_ledgers, never a total for the whole book, and both lists are always present, paged like the bills. Each unallocated party carries ledger_bill_wise, opening_balance (the ledger's opening as of the start of the books, with Tally's sign, so a debit opening is negative, never interpreted; absent when Tally sent none, which is unknown, not zero) and a composition: not_bill_wise_ledger (the ledger keeps no bills) or bill_wise_ledger_components_not_separated (what is left on a bill-wise ledger after its named bills: on-account entries, an unallocated opening, notes with no reference and anything else, not told apart). amount is a magnitude and direction says which side; unallocated.totals.by_composition splits the gross, receivable and payable apart, by those two over every party in the requested direction before paging (a row saved without one counts under composition_not_observed). No unallocated figure is labelled on-account. Party detail: with party (a ledger name) and detail, the result also carries a detail object for that party at the same as-of. A ledger name resolves only when spelled as in the book or differing from it only in ASCII case and spaces; the detail's ledger_match names the ledger read, how (matched: exact, or case_or_spacing, which the answer should mention by naming the ledger read) and any similar_ledgers that differ from it only in case or whitespace. Passing detail is the request to read the company's vouchers from the start of the books (or, for a named bill that Tally's bills reports list, from the earliest date they list for it) to as_of; nothing is read for a party detail without it. bill_trail (optionally one reference) lists every allocation of each bill in vouchers that are neither cancelled nor optional, oldest first, with state tied (the signed allocations equal Tally's own balance for that bill, or zero for a bill the report no longer lists), trail_does_not_tie (both numbers shown) or bill_identity_ambiguous (more than one native row or bill date for one reference, or a native row dated differently from the allocations; nothing is merged, and the native dates are shown). Naming a reference starts the read at the earliest date the reports list for it, so allocations dated earlier are not read, and a reference that carries two bill dates over the whole history, and is ambiguous there, can tie when named. The detail's own state is bills_listed, or for an empty list not_bill_wise_ledger or no_named_bill_for_party (no named bill in the vouchers or in Tally's list: a ledger that keeps no bills, one that is not a party's and a party with no bills are not told apart). The vouchers and the bills reports are two reads whose extents are not compared: a voucher posted between them usually shows as trail_does_not_tie, but two changes that compensate, or allocations that net to zero, can still read tied. unadjusted lists the party's on-account, advance and pending note allocations and compares their on-account sum with the party's unallocated amount: tied means the two figures are equal, not that the composition is proven (components that net to zero are not seen); residual_not_explained_by_vouchers gives the difference and whether it equals the ledger's opening balance; no_residual_row_for_party (with residual null) means Tally lists no unallocated amount for the ledger (a zero residual, a ledger that is not a party's and a name that matched no row are not told apart), so nothing is tied; not_bill_wise_ledger lists no rows. Its rows are row_amounts: as_allocated: each amount is the allocation as made, never net of what later allocations adjusted against its reference, so an advance shows what was received, not what is left; what is still open on a reference is Tally's own native_balance beside the row, null when the reports do not list the reference or list it more than once (so no balance is chosen), told apart only by native_rows. Either detail is window_returned_no_vouchers, with nothing tied or listed, when the voucher read returned no voucher (that read is not corroborated). Cost and limits: the detail keeps the party's entries from the whole company's vouchers, so its cost is that of a vouchers read over the same span, which is unmeasured on a large book, and any refusal of that read fails the whole outstandings call. One foreign-currency composite voucher anywhere in the window fails it (voucher_amount_invalid, or bill_allocation_amount_invalid on an allocation). At most 128 data requests are sent (each sent twice, as every read is, with its census and the company marks besides), each holding at most 42 vouchers before anything is measured, so a window of more than 5,376 of the company's vouchers is always refused, and a smaller one may be. The refusals are trail_window_too_large (name a reference that open_bills lists for the party, and the read starts at that bill's date), named_bill_window_too_large (a reference was named already, so nothing narrows it further) and unadjusted_window_too_large (nothing narrows it, so it is not available for that party on that book), each with reads.needed_at_least against reads.allowed; a refusal comes before any data request when the count shows it, otherwise when a measured part does, with window listing any part already read. The detail needs a complete read (detail_requires_a_complete_read, with the read's own partial_reason) and refuses rather than cuts an answer of over 500 allocations: trail_too_large (name a reference) or unadjusted_detail_too_large (nothing narrows it). The limit counts allocations, not bills, so a party with very many opening bills can meet agent_response_too_large instead. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
as_ofNo
limitNo
partyNo
detailNo
offsetNo
directionNoboth
referenceNo
ageing_basisNodue_date
company_guidYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, but the description adds substantial behavioral context beyond them: it states nothing is written to Tally, metadata-only receipt lines are logged locally, partial-read semantics and refusal codes are enumerated, and cost/limit behavior is described. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is an enormous, dense wall of text with nested asides and repeated explanations of partial-read behavior. While the first sentence front-loads purpose, the bulk is far too long for a tool description and contains redundancies that could be trimmed or structured into sections.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 10 parameters, no output schema, and complex partial/refusal behavior, the description is extremely complete. It covers return fields, paging, partial states, refusal conditions, cost limits, and detail semantics, so an agent has enough to call the tool correctly without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the full burden for 10 parameters. It explains as_of's default and return behavior, top ranking parties only, limit/offset paging, detail's enum values and required party, reference narrowing, direction sign rules, and party-name resolution rules. Only company_guid is left unexplained, which is self-evident as a required identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource: answering what is outstanding via receivable and payable totals, ageing, top parties, and open bills from Tally's paired bills reports. It implicitly distinguishes itself from sibling financial reports like balance_sheet and trial_balance by anchoring the operation in Tally's bills reports.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives extensive conditional guidance within the tool: when as_of is optional, when detail is required for party detail, when to name a reference, and how to page with offset/limit. However, it never names an alternative sibling tool or states when not to use this tool versus balance_sheet, trial_balance, or ledger_movement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

profit_and_lossA
Read-only

Return the Profit and Loss for a date range, derived from Tally's native Trial Balance: one line per reserved P&L primary group (Sales Accounts, Direct Incomes, Purchase Accounts, Direct Expenses, Indirect Incomes, Indirect Expenses), each the window's signed debit plus credit movement (a debit negative), with gross_result and net_result (a profit positive). The top-level lines is null while net_result is not established, so a derived line is never shown as the statement. Reads the Trial Balance, the group tree and Tally's own Balance Sheet inside one company, mode and book-extent bracket; requires observed INR currency (a book with more than one currency master is refused) and supported date boundaries. Each line reports amount (sum, present_count, empty_count): the sum is over the amounts Tally returned, and the empty ones it left out are counted. A result is established only if every ledger is classified (a ledger under a user-created primary group, or whose group chain is incomplete, is listed in unclassified, up to 100, with unclassified_total), no Stock-in-Hand ledger carries an amount, and Tally's own Balance Sheet for the window ties line for line to the derived one (balance_sheet_gate). Otherwise it is not_established with a reason (unclassified_ledger_carries_an_amount, closing_stock_not_derivable_from_trial_balance, profit_and_loss_ledger_not_returned, tally_balance_sheet_differs, or for gross and net tally_profit_and_loss_differs); for the two differs reasons, lines names the lines that did not tie. The top-level state is observed only while both gross_result and net_result are established, and not_established otherwise with the same reason the nested result carries (the weaker result decides). A book with stock items is expected to refuse; no inventory book has been measured. A Tally line with an amount the derivation has no counterpart for, such as a heading or a difference in opening balances, refuses rather than being guessed at. Tally's own statements carry no company identity and are bound only by the checks around the read. The gate has been measured over one full year on one book and one month on another; a window spanning more than one financial year is unmeasured. Tally's own Profit and Loss is read too and compared by display name in tie_out (matched, matched_empty_as_zero, differs or not_compared). Gross and net are also refused as tally_profit_and_loss_differs unless it ties: no line differs, no derived line is missing from it, and no line of its with an amount is uncompared, except its Cost of Sales : heading while that equals the derived Purchase Accounts plus Direct Expenses exactly (observed once, on one book). A stock line refuses. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
company_guidYes

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotation baseline of readOnlyHint/destructiveHint. It discloses currency preconditions (multi-currency books refused), stock refusal, the full gate/`state`/`reason` taxonomy, the exception for the `Cost of Sales :` heading, measurement coverage limits, and that it writes metadata-only receipts locally while writing nothing to Tally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the remainder is a dense multi-hundred-word run-on packing refusal reasons, gate details, and edge cases into single sprawling sentences. Much of this belongs in structured fields (e.g., an output schema or error enum) rather than prose, hurting scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden and does so thoroughly: it documents line shape, gross_result/net_result, state, reason values, unclassified handling, tie_out outcomes, and failure modes. An agent can anticipate nearly every return condition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the three parameters (company_guid, from, to) are named or documented in the description. It adds contextual constraints (single company/mode/book-extent, INR-only, supported date boundaries, multi-year windows unmeasured), but the caller still gets no per-parameter meaning, format, or example from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb (Return), resource (Profit and Loss), and scope (date range), then immediately clarifies the derivation source (Tally's native Trial Balance) and the exact line structure. This distinguishes it cleanly from siblings like trial_balance and balance_sheet.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the derivation chain and refusal conditions, but the description never explicitly says when to choose this tool over balance_sheet or trial_balance, nor what a caller should do on `not_established`. The conditions for a successful read are detailed, yet the routing guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purchase_registerA
Read-only

Read-only: a register of what the books record, not a GST return. It does not decide input tax credit eligibility or blocked credit, matches nothing against GSTR-2B or any portal, checks no GSTIN (party_gstin is returned only when the voucher carries one), does not return REFERENCEDATE yet (reference is returned only when the voucher carries one), does not classify an item invoice's purchase as taxable (has_taxable_entry is false when no entry sits on a Purchase Accounts ledger), and never sums tax across heads or vouchers. It does not treat reverse-charge journals, imports (IGST paid at customs) or input service distribution specially: a voucher that touches a Duties & Taxes ledger is listed by the rule below and nothing more. A GST duty head does not say whether a ledger is input or output, and a Debit Note can be a purchase return or a debit note issued to a customer: each row carries party_group (the voucher party's predefined group, for example Sundry Creditors or Sundry Debtors, when it resolves) and the tool does not guess which it is. It inherits the compliance read's refusals (an INR base currency is required; a book too large to list is refused; see ledger_masters) and refuses with register_master_mark_unavailable when Tally does not report the master-alteration mark. Each page re-reads the masters and the window, so rows can shift between pages. Return the Purchase and Debit Note vouchers of a date window that touch a ledger under Duties & Taxes, with the tax each entry carries taken only from the GST duty head recorded on that ledger's master -- never from a ledger name and never from an amount. Reads the full voucher window before pagination (use narrow dates) and the ledger masters twice, before and after it. Per row: tax_in_books lists each entry on a ledger whose head ComplyEaze Bridge recognises as {ledger, head, raw_head, amount}; duties_taxes_entries_without_gst_head lists entries on Duties & Taxes ledgers that carry no GST head and never assigns them one: observation not_tax_ledger is a ledger whose own tax type is not GST (usually TDS or another payable), absent is a ledger with no head whose tax type is GST or was not reported, which may be a GST ledger whose head is missing (tax_type says which); duties_taxes_entries_with_unrecognised_head lists entries whose head is not in the recognised vocabulary or contradicts the ledger's tax type, with the raw spelling and its observation; entries_on_ledgers_with_unresolved_group lists entries on ledgers whose group chain could not be resolved; taxable_entries are entries on Purchase Accounts ledgers only (a GST purchase booked to a fixed-asset or expense ledger has has_taxable_entry false); party_entries are the voucher party's own; other_entries is everything else (round-off included) with no role inferred. status is the first that applies of head_conflict, has_unrecognised_head, has_unresolved_group, has_entries_without_gst_head, has_other_entries, complete. The response state follows the rule vouchers uses: complete only when every voucher read was checked against a separate count of the window (a census, which ComplyEaze Bridge sends unless the book's voucher high-water mark alone proves it small, a few dozen vouchers), otherwise partial with reason nonempty_window_unqualified and the rows still returned; an empty window is complete when its corroboration read confirms it. A row's status is separate: it says whether every entry the voucher touches classified, and the state does not change it. Amounts are as the books state them (negative is a debit), never re-signed and never summed across heads; there is no input-credit or direction field. reference, party_gstin, is_invoice, post_dated follow vouchers: absent means not observed, and cancelled, optional and post_dated vouchers are returned flagged, not excluded. Every other voucher type that touches Duties & Taxes (Sales, Journal, Payment and so on) is listed apart in other_voucher_types_touching_duties_taxes, not in items: whether it belongs in a return is the CA's call. A voucher with no resolved class is listed under unclassified_voucher_type; a voucher that touches only unplaceable ledgers under vouchers_with_unplaced_ledgers; a Purchase or Debit Note voucher with no entry on a Duties & Taxes ledger under purchase_vouchers_without_duties_taxes_entry (exempt or unregistered purchases, tax booked to a ledger filed elsewhere, or a cancelled voucher). A cancelled voucher is listed there with cancelled true whether or not it was taxed, because the cancelled vouchers measured came back from Tally with no ledger entries; a cancelled voucher that keeps its entries is not measured. Rows are in items (paged by offset and limit like vouchers); each has has_taxable_entry, false when no entry sits on a Purchase Accounts ledger (an item invoice may hold it in an inventory allocation). The side lists carry exact counts (total) and at most 100 items (listed); every ledger name in the response is masked like vouchers masks it. A voucher that names a ledger the masters do not list, a master or voucher that changed while the window was read, or a ledger set aside for its currency, refuses (ledger_snapshot_drifted, voucher_window_changed_during_read, register_ledger_currency_excluded) and releases no rows; a row dated outside the window refuses as window_not_honoured. A sgst_utgst head is a state-side head that a consumer summing state tax must include alongside state_tax. Not measured: REFERENCEDATE (not returned), item invoices whose purchase ledger sits in an inventory allocation, and books with several currencies. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
limitNo
offsetNo
company_guidYes

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world), yet the description adds substantial behavioral detail beyond them: refusal codes (`register_master_mark_unavailable`, `ledger_snapshot_drifted`, `voucher_window_changed_during_read`, `window_not_honoured`), the fact that each page re-reads masters so rows can shift, the `state` qualification rule, and the metadata-only local log write. This is exactly the extra context annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is an unusually long, single-block wall of text — roughly 600+ words with heavy repetition ('compliance read's refusals', 'never summed', 'not measured' stated more than once) and it front-loads an exhaustive negation list instead of the purpose. Much of the content is genuinely useful, but it is not appropriately sized and the key purpose statement is not front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey the return shape — and it does: per-row fields, the `status` precedence chain, the `state`/`reason` rule, every side list, refusal behavior, and known non-measured cases. Despite its verbosity, an agent has essentially everything needed to call the tool and interpret the response correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It clarifies that `from`/`to` are a date window and that `offset`/`limit` page the rows 'like `vouchers`', and it warns that the full window is read before pagination. However it never explains `company_guid` and adds no semantics (format, bounds) for the paging parameters beyond what the schema pattern already implies. Partial compensation for the coverage gap merits a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return the Purchase and Debit Note vouchers of a date window that touch a ledger under Duties & Taxes,' and explicitly frames itself against siblings (not a GST return, see `ledger_masters`/`vouchers`). That is enough to distinguish it from `sales_register` and `vouchers`. It falls short of 5 because the positive purpose is buried mid-paragraph behind a long list of negatives, so the agent must parse several hundred words before learning what the tool actually returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context and numerous exclusions (does not decide input tax credit eligibility, does not match GSTR-2B, does not check GSTIN, does not sum tax) and advises 'use narrow dates' because the full window is re-read. It routes the agent to `ledger_masters` and `vouchers` and defers the return-classification choice to the CA. It lacks a crisp single 'use this when / not when' statement and never names a directly competing sibling as the alternative, keeping it at 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_evidenceC
Read-only

Everything the assistant reads from Tally through ComplyEaze Bridge in this chat, amounts included, is sent to the AI provider; redaction can only mask party names or drop narration. Never tell the user that no data has left their computer. Shows ComplyEaze Bridge's own recent reads since it started, kept in memory on this computer: request and response fingerprints, byte counts and state, no figures or book content (bounded: the newest limit records). It does not show what the AI provider received. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, yet the description states 'Each call appends metadata-only receipt lines ... to ComplyEaze Bridge's local log on this computer' — a side-effecting write to the local environment, which contradicts a read-only hint (and is consistent with idempotentHint=false). The description is otherwise unusually rich (egress disclosure, masking limits, in-memory bounding), but the stated write side effect directly conflicts with the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is dense and every sentence is relevant, but it is not front-loaded: the opening sentence is a privacy caveat rather than the tool's purpose, and the actual 'Shows ... recent reads' statement is buried mid-paragraph. For a one-parameter tool the ~120 words could be ordered far better.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one optional param, no output schema, and annotations present, the description covers what is returned (fingerprints, byte counts, state, no figures/book content), the memory-bounded window, and side effects. The main omissions are the default limit value and the readOnlyHint conflict, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter. It does explain the sole parameter: the result is 'bounded: the newest `limit` records,' clarifying that limit selects the newest N records. It omits the default (20) and any upper bound, but the core semantic is supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: it 'Shows ComplyEaze Bridge's own recent reads since it started,' including 'request and response fingerprints, byte counts and state.' That is a concrete, distinguishable purpose. However, it never explicitly names the siblings it differs from (e.g. egress_log, local_data_report), so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use/when-not guidance. The only directive ('Never tell the user that no data has left their computer') is a behavioral prohibition, not a usage rule, and there is no indication of when an agent should call this instead of egress_log or local_data_report. Usage must be inferred from the content of what it returns.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sales_registerA
Read-only

Read-only: a register of what the books record, not a GST return. It does not decide the place of supply, the tax rate, whether tax is payable or which part of a return a sale belongs in, matches nothing against any portal, checks no GSTIN (party_gstin is returned only when the voucher carries one), does not return REFERENCEDATE yet (reference is returned only when the voucher carries one), and never sums tax across heads or vouchers. It does not treat exports, sales under reverse charge or advances specially: a voucher that touches a Duties & Taxes ledger is listed by the rule below and nothing more. A GST duty head does not say whether a ledger is input or output, and a Credit Note can be a sales return or a credit note issued to a supplier: each row carries party_group (the voucher party's predefined group, for example Sundry Debtors or Sundry Creditors, when it resolves) and the tool does not guess which it is. A Debit Note, including one issued to a customer, is not a sales row: it is listed apart by identity and ledger names, with no amount. It inherits the compliance read's refusals (an INR base currency is required; a book too large to list is refused; see ledger_masters) and refuses with register_master_mark_unavailable when Tally does not report the master-alteration mark. Each page re-reads the masters and the window, so rows can shift between pages. Return the Sales and Credit Note vouchers of a date window that touch a ledger under Duties & Taxes, with the tax each entry carries taken only from the GST duty head recorded on that ledger's master -- never from a ledger name and never from an amount. Reads the full voucher window before pagination (use narrow dates) and the ledger masters twice, before and after it. Per row: tax_in_books lists each entry on a ledger whose head ComplyEaze Bridge recognises as {ledger, head, raw_head, amount}; duties_taxes_entries_without_gst_head lists entries on Duties & Taxes ledgers that carry no GST head and never assigns them one: observation not_tax_ledger is a ledger whose own tax type is not GST (usually TDS or another payable), absent is a ledger with no head whose tax type is GST or was not reported, which may be a GST ledger whose head is missing (tax_type says which); duties_taxes_entries_with_unrecognised_head lists entries whose head is not in the recognised vocabulary or contradicts the ledger's tax type, with the raw spelling and its observation; entries_on_ledgers_with_unresolved_group lists entries on ledgers whose group chain could not be resolved; taxable_entries are entries on Sales Accounts ledgers only (a sale booked to another ledger, or whose sales ledger sits in an inventory allocation, has has_taxable_entry false); party_entries are the voucher party's own; other_entries is everything else (round-off included) with no role inferred. status is the first that applies of head_conflict, has_unrecognised_head, has_unresolved_group, has_entries_without_gst_head, has_other_entries, complete. The response state follows the rule vouchers uses: complete only when every voucher read was checked against a separate count of the window (a census, which ComplyEaze Bridge sends unless the book's voucher high-water mark alone proves it small, a few dozen vouchers), otherwise partial with reason nonempty_window_unqualified and the rows still returned; an empty window is complete when its corroboration read confirms it. A row's status is separate: it says whether every entry the voucher touches classified, and the state does not change it. Amounts are as the books state them (negative is a debit), never re-signed and never summed across heads; there is no input-credit or direction field. reference, party_gstin, is_invoice, post_dated follow vouchers: absent means not observed, and cancelled, optional and post_dated vouchers are returned flagged, not excluded. Every other voucher type that touches Duties & Taxes (Purchase, Journal, Payment and so on) is listed apart in other_voucher_types_touching_duties_taxes, not in items: whether it belongs in a return is the CA's call. A voucher with no resolved class is listed under unclassified_voucher_type; a voucher that touches only unplaceable ledgers under vouchers_with_unplaced_ledgers; a Sales or Credit Note voucher with no entry on a Duties & Taxes ledger under sales_vouchers_without_duties_taxes_entry (listed by identity only; the tool does not say why such a voucher carries no tax entry). A cancelled sale that Tally returns with no ledger entries is listed there too, with cancelled true; no cancelled sale has been read, so whether one keeps its entries is not measured. Rows are in items (paged by offset and limit like vouchers); each has has_taxable_entry, false when no entry sits on a Sales Accounts ledger (a sale typed on Tally's screen, or an item invoice of another shape than the imported one that was measured, may hold the sales ledger in an inventory allocation instead; not measured). The side lists carry exact counts (total) and at most 100 items (listed); every ledger name in the response is masked like vouchers masks it. A voucher that names a ledger the masters do not list, a master or voucher that changed while the window was read, or a ledger set aside for its currency, refuses (ledger_snapshot_drifted, voucher_window_changed_during_read, register_ledger_currency_excluded) and releases no rows; a row dated outside the window refuses as window_not_honoured. A sgst_utgst head is a state-side head that a consumer summing state tax must include alongside state_tax. Measured so far: sales_register was run against a live Tally on two synthetic companies. On the first, once per day, for one taxed Sales item invoice and one untaxed one: the taxed sale came back as one row with its CGST and SGST/UTGST heads taken from the ledger masters and its sales ledger as the taxable entry, and the untaxed one (read once by an earlier build; its voucher window is committed, its masters and the tool's answer are not) was counted under sales_vouchers_without_duties_taxes_entry. On the second, which has 44 ledgers, for one Credit Note in voucher view booked on account: one row, with its CGST and state-tax heads and its sales ledger as the taxable entry. One Sales accounting voucher (not an invoice) was also classified, in tests, against the ledger masters of the purchase register's lab book. A Credit Note is returned as a row with its signs reversed as Tally sends them: the tool neither nets nor flips, so a caller that sums tax over a window must add signed amounts. The measured Credit Note of 1,000.00 with 90.00 CGST and 90.00 State Tax came back with the sales entry -1000.00, each tax entry -90.00 and the party entry 1180.00, where a Sales row has the sales and tax entries positive and the party entry negative. The state-side tax head is state_tax (raw State Tax) on one measured book and sgst_utgst (raw SGST/UTGST) on another; both are recognised heads for the same side of the tax, so a caller must not look for one of them only. The cost of a call varies by book: 96 requests on a book with 8 ledgers and one currency, 118 on one with 44 ledgers and two currencies, which adds a voucher census and base-currency reads; the result does not report the cost. Not shown by any run: an invoice-view Credit Note; an inter-state (IGST) line; a cancelled or optional sales voucher; an unrecognised or missing duty head on a sale; more than one voucher in a window; paging; a company with a registration; a tax that Tally computes itself; a sale typed on Tally's screen; accounting-invoice mode; a post-dated sale; a REFERENCE or a populated PARTYGSTIN on a sale; REFERENCEDATE (not returned); and a ledger or voucher kept in a currency other than the book's base. A row of such a kind is returned, not withheld, and carries not_measured_live naming why (invoice_view_credit_note, inter_state_line, sales_ledger_not_an_entry, cancelled, optional, post_dated, party_gstin_present, reference_present) only where the row itself shows the kind. Kinds a row cannot show are never marked and are not vouched for: a sale typed on Tally's screen in voucher view, a tax Tally computed itself, a duty head no sales capture has (such as cess), an invoice of another shape than the one run (for example several goods lines), and a ledger or voucher kept in a currency other than the book's base; an unmarked row is not a measured one in those respects. A row is marked inter_state_line only when a tax entry's ledger master carries a recognised IGST head; an IGST ledger with no head, or an unrecognised head, is listed under the without-head or unrecognised list and the status is not complete. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
limitNo
offsetNo
company_guidYes

TDQS

A3.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations already covering readOnly/destructive/idempotent/openWorld, the description adds substantial context: enumerated refusal codes (register_master_mark_unavailable, ledger_snapshot_drifted, window_not_honoured), the fact that pages re-read masters so rows can shift, sign conventions (negative = debit, Credit Note signs not flipped), the measured/not_measured_live marking scheme, and that it writes nothing to Tally but appends metadata-only receipt lines locally. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is an enormous wall of text in which the core purpose appears roughly mid-way, after a long list of things the tool does not do. Content is dense and much of it is genuinely relevant for a complex tool, but the structure is poorly front-loaded and heavily redundant, so it fails the conciseness bar.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a highly complex, stateful read, the description compensates by enumerating every returned row group (tax_in_books, duties_taxes_entries_without_gst_head, taxable_entries, side lists, status/state rules), the state machine (complete/partial), and edge cases. An agent has essentially everything needed to interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 params, so the description must carry the load. It does convey that from/to define the date window, that offset/limit page like vouchers, and it references company context, but it never specifies the date string format, the meaning of company_guid, or limit/offset defaults. Partial compensation, so a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description ultimately states a specific verb+resource: 'Return the Sales and Credit Note vouchers of a date window that touch a ledger under Duties & Taxes,' and sharply distinguishes itself from purchase_register and vouchers. However the actual purpose is buried under a long preamble of exclusions, and the opening sentence ('a register of what the books record, not a GST return') is more negation than definition. Clear once found, but not front-loaded.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real usage context – 'use narrow dates', routes other voucher types to other_voucher_types_touching_duties_taxes, and notes cost varies by book. But it never explicitly frames when to pick this over purchase_register or vouchers; selection guidance is implied rather than stated, and the bulk is behavioral caveat rather than when/when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stock_summaryA
Read-only

Return the stock summary: the closing stock value per stock item as of a date, with the total of those values checked against Tally's own Stock Summary, and whether inventory is integrated with the accounts. The top-level state is one of three. value_total_matched: the items are returned, and their closing values add up to the sum of the top-level lines of Tally's own Stock Summary; only that total was compared, so value_total_matched can stand beside partial true when some items have no closing value. no_stock_items: Tally's own stock item count is 0, the item list is empty and the Stock Summary is empty; items is an empty list. not_established: no item is returned (items is null), reason says why and remediation says what to do next: tally_stock_summary_differs (the report has a total the items do not add up to), tally_stock_summary_shows_no_value (the items carry a value and the report has no amount; an empty report is not told apart from one Tally did not render) or stock_values_not_comparable (nothing could be compared). Such a result carries unchecked_comparison in place of tie_out, for investigation only: neither side is a stock value or a total, the items' side adds only the closing values present, and the closing-value total of a matched read is the only thing this tool checks. A not_established result is not held for paging, and it replaces any earlier read of the same date. Quantities are withheld: nothing checks them, so none is returned; totals.closing_quantity_unread_count counts those that could not be read (a compound unit, or a unit with a space), which do not refuse the read. checks says per field what is checked, not_checked or withheld: the closing-value total is checked; each value on its own, the names, parents and base units, and whether the date was honoured are not. as_of (YYYYMMDD or YYYY-MM-DD) must be a 31 March, not before the book's start or after today. The only period measured is the period ending 31 March 2026; other years' 31 March are admitted but unmeasured, and any other date is refused as stock_summary_as_of_not_measured before any request. The period runs from 1 April (or the book's start, if later) to as_of. Each item carries name, guid, parent, base_unit and closing; closing holds value only, exactly as Tally sends it. Empty is not zero: an empty closing value is returned as null and counted in totals.empty_closing_value_count, and a value sent as 0.00 is a value. Signs are kept, as in the trial balance: a negative value is a debit, which is stock held, and Tally's own Stock Summary screen shows it as a positive value (measured on one synthetic company, licensed TallyPrime 7.1 Silver). value_sum adds the values with their signs, so stock held gives a negative sum, and totals.value_sum_signs is always as_sent_negative_is_debit; value_sum is written at the scale of the values it adds, and is null with partial true whenever any item's closing value is empty. The opening quantity and value are read but not returned, because their as-at date is unmeasured. inventory reports integrated, inventory_on and batchwise as yes, no or unknown (unknown does not refuse), and basis states what Tally reported (ISINTEGRATED); how the books use these values (as closing stock, or against a Stock-in-Hand ledger) is not measured. An item valued at zero or with no value adds nothing to either total, so only Tally's own stock item count vouches for it; that count followed the one delete measured (one synthetic company, one sample), which is not proof of a complete list (checks.item_list_complete is not_checked). item_count_cross_check reports rows, tally_count and status: items are returned only when the two are equal. Otherwise the read is refused as stock_summary_item_count_differs (both numbers under counts); a count Tally did not give refuses as stock_summary_read_failed with cause stock_item_count_unavailable, and a Stock Summary this Tally does not recognise with cause stock_report_unknown. items (1 to 50 GUIDs, each once) filters what is returned from the held read; a GUID not found is listed under items_not_found, not refused, and totals, tie_out and item_count_cross_check still cover the whole book. The inventory flags, the stock items and the Stock Summary are each read twice inside one company, mode and book-extent bracket, or the read is refused. A book whose inventory is off is refused as stock_not_enabled, Education mode is refused, and a company split by year, whose sibling companies share the GUID, is refused (stock_summary_read_failed, cause company_flags_not_one_row). Small books only: the items are read whole only when the master-alteration mark (it counts masters of every kind) times an assumed worst-case row fits 16,000,000 bytes, a mark of at most 874; a larger book is refused before any item request as stock_summary_too_large, with size, and retrying refuses again. Typical stock-heavy client books refuse today. A row count or response size past what was admitted refuses as stock_summary_bound_premise_violated, or stock_summary_extent_changed if the book moved. Limits: company totals only, with no godown or batch split and no rates. The row shape and the tie were measured on one synthetic book on one licensed TallyPrime 7.1 (the tie also once on a client book); other releases, dates and books are not measured. An item's name and parent are masked under mask_parties, because stock names can carry a customer's or supplier's name (Tally's reserved root as a parent is left as it is); guid is not masked and is what items filters on. A first page (offset 0) always reads afresh and holds the read; a later page for the same date is served from it while the book extent, including ALTVCHID and ALTMSTID, is unchanged, and each result reports snapshot. Pass the first page's snapshot_id on later pages to have the call refused as listing_snapshot_changed instead of continuing from a different read (cause book_changed_since_first_page, or snapshot_not_held for an id not held). A change that moves neither mark is not seen, so a later page can be up to 10 minutes old after such a change. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofYes
itemsNo
limitNo
offsetNo
snapshot_idNo
company_guidYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: it discloses refusals and their causes, snapshot/paging semantics with a 10-minute staleness window, masking behavior under mask_parties, size limits (~874 master mark, 16MB premise), the read-twice consistency rule, and the metadata-only local receipt log. The annotations only declare readOnly/openWorld/idempotent/destructive hints, so the description carries real additional behavioral value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and the three `state` values are front-loaded, which is good, but the body is an exceptionally dense wall of text that mixes state semantics, refusals, paging, masking, limits and provenance. Given the tool's complexity much of it is load-bearing, but it is not economical and is hard to parse on first read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema this is remarkably complete: it enumerates every `state`, every refusal cause, per-field `checks` semantics, the `inventory` flags, tie-out behavior, and paging. An agent has essentially everything needed to call and interpret the tool, with the sole caveat of unmeasured release/date coverage being explicitly acknowledged.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage the description must compensate, and it does: it explains `as_of` accepted formats and the 31 March / book-start / today constraints, that `items` is 1–50 GUIDs filtering the held read, and that `snapshot_id` is the first page's id to force `listing_snapshot_changed`. `limit` and `offset` are only implied through the paging narrative (offset 0 = first page), leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb and resource ('Return the stock summary: the closing stock value per stock item as of a date') and further scopes it against Tally's own Stock Summary. It clearly distinguishes itself from sibling reporting tools like trial_balance and balance_sheet, which are referenced only for sign convention.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives strong conditional guidance: when results are 'for investigation only', when a read is refused (Education mode, inventory off, oversized book), and the constraints on `as_of` (must be 31 March). However, it never explicitly routes the agent between this tool and sibling alternatives such as trial_balance or balance_sheet, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tally_statusA
Read-only

Return loopback endpoint status and observed loaded-company identity tuples. Every ComplyEaze Bridge process on this computer (the desktop app and each AI client) sends to one Tally port one request at a time. Any tool that reads may therefore refuse, having sent nothing: tally_endpoint_busy when another ComplyEaze Bridge window or AI client held the port for longer than this call's bounded wait (about 10 seconds in all per call); it carries retry_after_s, and the same call is safe to repeat after that many seconds. post_import can be refused the same way, but only repeat it when the refusal says attempt_recorded is false; once an attempt is recorded, follow the refusal's next_step (verify_import) and never call post_import again. tally_endpoint_lock_unavailable when ComplyEaze Bridge could not open its local coordination file. today is this computer's calendar date (YYYYMMDD), the date outstandings uses when as_of is left out, and ledger_masters with fields=compliance for party_gstin. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: single-request-at-a-time concurrency across all Bridge processes, a ~10 second bounded per-call wait, the tally_endpoint_busy and tally_endpoint_lock_unavailable failure modes, the retry_after_s field, and the metadata-only local receipt log with an explicit 'writes nothing to Tally' guarantee. This is exactly the behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose well, but the body is a dense run-on spanning error codes, retry rules, and a `today` explanation that mostly concerns outstandings and ledger_masters rather than this tool. The date discussion is scope creep on a zero-parameter status probe and dilutes the core message.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param read-only probe with no output schema, the description covers failure modes, retry policy, and side effects (local receipt log, no Tally writes) adequately. It never describes the shape of the returned status/identity tuple, which is the one remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4; the description instead documents the derived `today` value (this computer's calendar date, YYYYMMDD) that other tools consume. That is useful context but is not parameter semantics for this tool itself, so it stays at baseline rather than exceeding it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Return loopback endpoint status and observed loaded-company identity tuples.' An agent can tell this is a health/status probe distinct from the data-reading siblings, though the phrasing is jargon-heavy and no sibling is named as an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the contention model clearly and when refusals occur (another ComplyEaze Bridge process holding the port), plus what to do: retry after retry_after_s for reads, and for post_import only when attempt_recorded is false, otherwise follow verify_import. It does not explicitly frame which situations should trigger calling this tool first, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trial_balanceA
Read-only

Return native ledger-wise Trial Balance for a date range, using Tally's TBAL fields without scanning vouchers. Requires an INR base currency and supported date boundaries. A book with several Currency masters is read through the base Tally identifies: only its plain base-currency ledgers are returned, with ledgers_scope base_currency_ledgers_only, the ledgers kept in another currency named under foreign_currency_ledgers_excluded, and the base-currency ledgers whose balances Tally shows in another currency named under base_currency_ledgers_mixed_excluded; totals then cover those plain ledgers only (totals_scope) and are not expected to balance. Preserves empty amounts; paired source stability is not voucher-level reconciliation. Pagination limits output only. A first page (offset 0) always captures a fresh report and holds it in memory; a later page for the same period (offset > 0) is served from it while the company's book extent, including ALTVCHID and ALTMSTID, is unchanged, at the cost of one small extent read. Each result reports snapshot (id, master_alter_id, voucher_alter_id, read_at, reused). Pass the first page's snapshot_id on later pages to have the call refused with listing_snapshot_changed (cause book_changed_since_first_page or snapshot_not_held) instead of continuing from a different report. A change that moves neither mark is not seen. A voucher delete, a cancel, a save with no change and a master change made in Tally's own screens each moved a mark when measured (protocol reference section 11c.5, one run each), but a change that moves neither can leave a later page up to 10 minutes old. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
limitNo
offsetNo
snapshot_idNo
company_guidYes

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotation baseline: it discloses currency-scoping behavior, that totals are not expected to balance, empty-amount preservation, snapshot caching mechanics with concrete staleness bounds, refusal codes, and metadata-only receipt logging with an explicit 'writes nothing to Tally'. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose well, but the body is a dense wall of parenthetical detail (field names, protocol section references, staleness anecdotes) that is hard to scan. Most sentences carry real behavior, but the volume works against quick selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the returned structures (ledgers_scope, foreign_currency_ledgers_excluded, base_currency_ledgers_mixed_excluded, totals_scope, snapshot). An agent has enough to interpret the response without opening anything else.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load, and it does for most: from/to as a date range, offset semantics (offset 0 vs offset > 0), snapshot_id pass-through, and 'pagination limits output only' for limit. company_guid is never explained, leaving one gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource: 'Return native ledger-wise Trial Balance for a date range, using Tally's TBAL fields without scanning vouchers.' The 'without scanning vouchers' clause implicitly separates it from voucher-reading siblings, but no sibling (balance_sheet, profit_and_loss) is named to make the choice explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Prerequisites are stated (INR base currency, supported date boundaries) and the snapshot_id usage pattern is spelled out. However, when to choose this over sibling reports like balance_sheet or profit_and_loss is never addressed, leaving the core selection question to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_mastersA
Read-only

Bind 1–100 nonblank ledger names (at most 1024 characters each) against the live catalogue. An identifier embedded in a master name is matched before the name itself. match_state is exact, identifier, near_miss or missing; folded names remain near_miss candidates because this catalogue has no qualified scope to bind them. Only exact is admitted by build_import_xml, and a bound row alone carries exact_live_spelling. A near-miss is never resolved: it returns candidates with the rule that surfaced each, bounded to 25 names and 8192 UTF-8 bytes per requested name, with candidate_count, candidate_count_is_lower_bound and truncation reported. When candidate_count_is_lower_bound is true, the count is a conservative lower bound and must be shown as at least that many candidates. There is no ranking and no score. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
ledgersYes
company_guidYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnly, non-destructive, non-idempotent), leaving the description to disclose the real behavior – it does so richly: match_state semantics, that near-misses are never resolved, candidate bounds (25 names / 8192 bytes), the lower-bound count flag, truncation reporting, and metadata-only receipt logging. The append-per-call logging is consistent with idempotentHint=false, so there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the body is an extremely dense run of jargon-heavy clauses (near_miss, candidate_count_is_lower_bound, fingerprints) crammed into few sentences with poor scannability. Every sentence carries content, but structure could be far cleaner.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, non-trivial tool with no output schema, the description covers return behavior (match states, candidate lists, truncation, lower-bound counts) thoroughly. The only meaningful omission is any explanation of the required company_guid parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate; it partially does by specifying ledgers constraints (1–100 entries, nonblank, max 1024 chars each). But company_guid – a required parameter – is never explained anywhere, leaving a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb (bind/validate) and resource (ledger names against the live catalogue) are stated specifically, and the description carves out clear scope by noting only `exact` results are admitted by build_import_xml and that nothing is written to Tally. An agent can distinguish this from sibling tools like masters or ledger_masters without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is clear: it is a pre-import validation step, and only `exact` results pass downstream to build_import_xml. However, it never explicitly names an alternative sibling tool or states when NOT to use this one, so the routing guidance is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_importA
Destructive

Read back a manually imported local batch and write Proof-of-Post files. This never dispatches import XML to Tally. The result gives verification_status, counts, every voucher that is not posted_verified (unverified_vouchers; for a native post a voucher may be bound_not_in_window, book_rolled_back or sent_not_attributed, never absence, with post_span_binding naming how the post was attributed and a plain summary to give the person first; a book restored from a backup and then keyed past the post's voucher mark before this check is not seen as rolled back, and if Tally then reused the post's MasterIDs a bound voucher could read posted_verified while being another voucher with the same content: unmeasured, bridge#1050), duplicates, unrelated_duplicates_in_window and ambiguous_within_batch in full, never cut to fit. Only the posted_verified vouchers are paged, as items from offset out of verified_total; when the response cap shortens them it sets truncated and next_offset. To read further pages, call again with proof_sha256 set to the returned proof.sha256 and offset set to next_offset: those pages come from the persisted proof and never read Tally again. The call is refused with verification_proof_changed if a newer verification replaced that proof, and with verification_too_large_to_report if the parts never cut do not fit the response cap. The full proof is always written to disk. Reads the batch's date window from Tally, then creates or replaces the batch's saved proof files and saves a status record, and may also save a verified baseline, a masters-check record and, for a native post, the binding of its vouchers to the Tally vouchers its post created, in ComplyEaze Bridge's local folder on this computer (paging an existing proof only reads it); writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNo
batch_idYes
company_guidYes
proof_sha256No

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations present, the description adds substantial depth beyond them: it explains local proof-file writes, status-record creation, possible baseline and masters-check saves, paging via persisted proofs, specific refusal conditions, and that Tally is never written to. This is rich behavioral context that goes well beyond readOnlyHint=false and destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a dense, run-on block with nested parentheticals and semicolon-chained clauses. The first sentence front-loads the purpose well, but the remainder is very difficult to parse and is far longer than necessary for an agent to select and invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, absent output schema, and the available annotations, the description is complete enough: it covers return fields, paging, truncated/next_offset, error refusals, proof persistence, and local file writes. An agent has what it needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for four parameters. It explains offset and proof_sha256 clearly, and gives context for batch through 'the batch's date window' and 'batch's saved proof files,' but company_guid is never mentioned. Partial compensation justifies a mid-range score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Read back a manually imported local batch and write Proof-of-Post files.' It also distinguishes itself from an import-dispatch tool by stating it 'never dispatches import XML to Tally.' The purpose is unambiguous and not a restatement of the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains how to page an existing proof using proof_sha256 and offset, which is useful procedural guidance. However, it does not say when to choose this tool over sibling verification or reporting tools, nor does it name alternatives or exclusions beyond the 'never dispatches import XML' limitation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voucher_presenceA
Read-only

Answer which of 1–500 proposed vouchers are already in the book. presence is present, possibly_present or absent, and only present names a book voucher. A nonempty window is read as complete only when its rows were checked voucher for voucher against a count ComplyEaze Bridge made of the window first, as vouchers labels it; otherwise, as on a new or test company with a few dozen vouchers, it is partial with reason nonempty_window_unqualified. An empty window can still be corroborated complete. present and possibly_present never need a complete window and are produced either way, but absent means absent from the whole window and is only ever produced from one proven complete — a proposal that would otherwise be absent from a merely partial window instead comes back possibly_present with reason window_not_proven_complete. The conditional decision basis can use a voucher number on a voucher type you declare manual — unique on both sides, within an observed voucher type, and never onto a cancelled or optional voucher; or, for a voucher ComplyEaze Bridge wrote into a file a person imported by hand, the narration marker derived from the supplied batch_id and bridge_txn_id together. A native post (post_import) writes no marker, so this tool cannot identify its vouchers: one edited or re-dated in Tally can read absent here. Check a natively posted batch with verify_import, which finds its vouchers by the GUIDs its post created once its binding is made, before posting any of them again. It neither accepts nor reads client remote identifiers. Supplying only one narration identity component is an error. The marker reaches only the current writer identity scheme; older-scheme ComplyEaze Bridge writes stay unidentified rather than matched. Date, party and amount only ever produce candidates, with the rule that surfaced each and no ranking or score. Every voucher type a proposal names needs a declared numbering method; under automatic Tally discards the supplied number, so nothing can be decided from it. absent means absent from this window, so cover the dates the book could hold. Reads the full window before comparing; dense windows can fail source limits. Party names bind through the same rules as validate_masters. A reported difference on a present voucher is a finding for a person, not a work item: correcting a voucher by Alter or Cancel silently creates a duplicate instead (§9.7), and no ComplyEaze Bridge path can correct a voucher it did not write. This never dispatches import XML to Tally. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
limitNo
offsetNo
vouchersYes
numberingYes
company_guidYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/destructive annotations, it discloses side effects (appends metadata-only receipt lines to a local log, writes nothing to Tally), resource risk (reads the full window, dense windows can fail source limits), and error conditions (supplying only one narration identity component is an error). It also explains limitations of the marker scheme and that correcting a voucher creates duplicates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, but the remainder is one very long, densely parenthetical paragraph with run-on clauses. Much of the detail earns its place given the tool's complexity, yet the packing reduces scannability and could be structured better.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex comparison tool with no output schema and 0% schema coverage, the description is thorough: it defines the presence values, the complete/partial window semantics, the decision basis, error cases, and side effects. It is nearly complete, though pagination behavior and the company_guid parameter remain unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden and largely succeeds: it explains numbering_method values (manual/automatic/unknown), how voucher_number and the batch_id/bridge_txn_id pair form the decision basis, and that date/party/amount only produce candidates. However, it never clarifies the limit/offset pagination parameters or company_guid, leaving part of the schema undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a precise verb and resource: answer which of 1-500 proposed vouchers already exist in the book. It differentiates itself from siblings by explicitly naming verify_import as the tool for natively posted batches, which this tool cannot identify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent explicitly: use verify_import for natively posted batches, and it explains the conditions that yield complete vs partial and present/possibly_present/absent. It also states exclusions (native posts, older-scheme writes, client remote identifiers), which is exactly the when-not guidance needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vouchersA
Read-only

Return literal-window voucher evidence with curated metadata and redaction. Reads the full source window before selectors and output pagination; limit does not reduce Tally work. A complete window is held in memory for its later pages: a page after the first (offset above 0) is served from that read, with one marks read instead of the whole window, while the company's two marks (vouchers and masters) are unchanged, and its result carries snapshot (id, master_alter_id, voucher_alter_id, read_at, reused). Pass the first page's snapshot_id on later pages to have the call refused with listing_snapshot_changed (cause book_changed_since_first_page or snapshot_not_held) instead of continuing from a different read. Without it, a later page whose window is no longer held or whose book moved reads the window again, and rows can shift between pages. A partial window is not held. A later page is served only for the same question: the same dates (a date is the same however it is written), the same voucher-type selector and the ledger argument exactly as typed on the first page (a differently spelled ledger is a different question and reads the whole window again). Without snapshot_id, a page whose held window found the book moved reads afresh and carries earlier_snapshot (offsets_do_not_continue true): start again from offset 0. snapshot_id is read on later pages only. Only a page whose snapshot.reused is true continues the earlier pages: a later page with reused false or no snapshot is a fresh read whose offsets may not continue them. A change that moves neither mark is not seen; not established: whether a company feature or configuration change that alters export content moves a mark, whether a restored copy of the company with the same marks is told apart, and whether a remote writer's save shows in the marks at once. So a held page can be up to ten minutes old after its read finished; a first page is always read fresh. Use narrow dates; dense windows are unqualified and can fail source limits. state is complete only when the window's rows were checked voucher for voucher against a count ComplyEaze Bridge made of the window first, or the window was empty and corroborated; otherwise it is partial with reason nonempty_window_unqualified. ComplyEaze Bridge makes that count on every book too large to read whole without one, so in practice only a new or test company with a few dozen vouchers ever reads partial for this reason. Selectors apply after the window is labelled, so a zero from a complete window is a checked zero. A voucher created, altered or deleted between the count and the read refuses with voucher_window_part_not_admitted, cause part_census_mismatch; call again. Each item carries cancelled and optional (always booleans; a source that omits or cannot assert either fails the whole read rather than guess). post_dated behaves differently: a real capture has shown Tally omitting that tag entirely rather than asserting No, so it is boolean only when Tally asserted Yes/No, and the key is absent from the item when Tally did not report it. Absent is not evidence of false — it means ‘Tally did not say’, not ‘Tally said no’, and a caller must branch on key presence, not on falsiness, before treating a voucher as not post-dated. Neither this tool nor voucher_presence filters out post-dated (or optional/cancelled) vouchers; the caller decides what a non-posting or unobserved status means for its own computation. reference, is_invoice and party_gstin follow the same absent-means-not-observed convention as post_dated: each key is present only when Tally reported a non-empty value for it, and its absence must not be read as false or as an empty string. is_invoice is a boolean exactly like post_dated; reference and party_gstin are non-empty strings when present. A captured book with no GSTIN recorded against a party's ledger has shown party_gstin absent on every voucher for that party, which is not evidence Tally cannot report one. Filter by voucher type with at most one of: voucher_class (a reserved class such as Purchase, matched however the book has renamed its types, and including their child types, by Tally's own class functions), voucher_type_guid (exactly one type; a GUID that is not this company's is refused as voucher_type_guid_foreign), or voucher_type (one display name, matched ignoring ASCII case as Tally does). Voucher-type names are editable in Tally, so a display name that is a class name, or the reserved name of any type in scope, is refused as voucher_type_ambiguous whenever the types of that name are not exactly the types of that class or reserving that name; the types involved are listed in candidates (bounded to a quarter of the response budget, with candidates_total and candidates_truncated). Use voucher_class or voucher_type_guid instead. That check sees only the vouchers in scope: the window read, after any ledger filter. A type with no voucher in scope is not seen, and when nothing is in scope nothing is ambiguous. When a voucher_type name selects no voucher, one more read lists the book's voucher types: a name no type carries is refused as unknown_voucher_type, with the name as requested and every type in candidates, nearest name first (bounded as above); a name some type carries keeps its zero. A filtered result carries voucher_types: the types included and every type in_scope (name, GUID, own reserved name, class and row count), and each item carries voucher_type_guid, voucher_type_reserved_name and voucher_class (null outside the measured classes). A row whose type Tally cannot resolve, or whose class answers contradict each other or its reserved name, refuses the whole read. A voucher whose amount Tally stored as a foreign-currency composite (-$ 100.00 @ I₹ 86/$ = -I₹ 8600.00) is withheld, not read: it still passes every date, ledger and type check, and is then listed in withheld_vouchers (GUID, date, type, number and cause foreign_currency_amount_unparsed, up to 100) with an exact withheld_total, the same on every page. items and total then exclude it, state is partial with reason vouchers_withheld, and coverage says so; the voucher_types row counts still include it, and the listing is also bounded to a quarter of the response budget. Any other amount Tally did not return as a plain decimal still refuses the whole read, as does a composite whose foreign and base amounts differ in sign while the foreign amount is not zero. A ledger name resolves only when spelled as in the book or differing from it only in ASCII case and spaces; ledger_match names the ledger read, how (matched: exact, or case_or_spacing, which the answer should mention by naming the ledger read) and any similar_ledgers that differ from it only in case or whitespace. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
fromYes
limitNo
ledgerNo
offsetNo
snapshot_idNo
company_guidYes
voucher_typeNo
voucher_classNo
voucher_type_guidNo

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotation bar: it discloses snapshot reuse and the up-to-ten-minute staleness window, why later pages can shift rows, the exact refusal causes (listing_snapshot_changed, voucher_window_part_not_admitted, voucher_type_ambiguous, unknown_voucher_type, voucher_type_guid_foreign), withheld foreign-currency vouchers, and the absent-means-not-observed convention. This is exceptional behavioral disclosure for a read tool whose only annotations are readOnlyHint/destructiveHint/idempotentHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single enormous run-on paragraph of well over a thousand words with no headings or bullet structure, mixing snapshot internals, selector rules, edge cases and logging into one block. Some length is warranted by the complexity, but the total is far past what an agent can parse efficiently and the lead is not cleanly front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, it explains return shape (items, total, state/reason, coverage, snapshot, earlier_snapshot, withheld_vouchers, voucher_types, ledger_match) and enumerates failure modes and edge cases thoroughly. For a 10-parameter, high-complexity read tool this leaves almost nothing an agent needs to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage across 10 parameters, the description carries the full burden and largely does: it defines limit/offset semantics (limit does not reduce Tally work), the snapshot_id contract, ledger spelling rules with ledger_match, the three mutually-exclusive type selectors, and date equivalence. company_guid is left unexplained, but the rest is richly specified well beyond the bare schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening states a specific action (return voucher evidence for a date window) and the body distinguishes this tool from the sibling voucher_presence by noting neither filters post-dated/optional/cancelled vouchers. However, the lead phrase 'literal-window voucher evidence' is jargon-heavy and the core purpose is buried under snapshot mechanics rather than stated cleanly up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives strong parameter-level routing guidance (prefer voucher_class or voucher_type_guid over voucher_type, use narrow dates, dense windows are unqualified) and mentions voucher_presence, but never frames when to choose this tool over the many sibling listing tools (sales_register, purchase_register, ledger_movement) that overlap in scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voucher_schemaA
Read-only

Return the fail-closed local voucher-file schema; no Tally request is sent. Each call appends metadata-only receipt lines (tool, company, counts, request and response fingerprints; no book content) to ComplyEaze Bridge's local log on this computer; it writes nothing to Tally.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false. The description adds valuable context beyond annotations by detailing that each call appends metadata-only receipt lines to a local log, specifies what data is logged, and confirms no writes to Tally. It does not violate any annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loads the purpose, and includes relevant behavioral details without unnecessary verbosity. It is appropriately sized for a zero-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, no parameters, and no output schema, the description provides sufficient information about what it returns and its side effects. There is no output schema, so the description covers the essential return behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, the baseline is 4. The description does not need to explain any parameters, and it appropriately focuses on the tool's behavior rather than parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a fail-closed local voucher-file schema and explicitly says no Tally request is sent. It distinguishes itself from siblings like vouchers or voucher_presence by emphasizing the local schema-return behavior, though it does not name any specific sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. It implies usage by stating it returns a schema and writes a local receipt, but provides no conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.4.2
    • Changedvouchers1 field changed
      • addedInput schema / properties / snapshot_id
        Added value: +{
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
  2. 20 tool updatesv0.4.1
    • First observedbalance_sheet
    • First observedegress_log
    • First observedledger_masters
    • First observedledger_movement
    • First observedlist_companies
    • First observedlocal_data_report
    • First observedmasters
    • First observedoutstandings
    • First observedprofit_and_loss
    • First observedpurchase_register
    • First observedread_evidence
    • First observedsales_register
    • First observedstock_summary
    • First observedtally_status
    • First observedtrial_balance
    • First observedvalidate_masters
    • First observedverify_import
    • First observedvoucher_presence
    • First observedvoucher_schema
    • First observedvouchers

TDQS

A3.6/5.0

Scored across 20 tools

Disambiguation4/5

Most tools target clearly distinct TallyPrime resources or workflows, such as balance_sheet, profit_and_loss, trial_balance, vouchers, and the GST registers. The main overlap is between egress_log and read_evidence, which both report local tool-call receipts, and tally_status also returns company identity tuples alongside list_companies. These are distinguishable from their descriptions, but not entirely free of potential misselection.

Naming Consistency4/5

All tool names use snake_case, which is consistent and readable. However, the set mixes noun-style names (balance_sheet, masters, vouchers, outstandings) with verb_noun names (list_companies, validate_masters, verify_import), so the pattern is not uniformly predictable. This is a minor deviation rather than a coherence problem.

Tool Count4/5

Twenty tools is on the heavy side for an MCP server, but the domain—TallyPrime accounting, GST registers, import verification, and local evidence—is broad enough that most tools serve distinct workflows. The count is slightly over the ideal range rather than excessive. No obvious duplicate tools inflate the count.

Completeness3/5

The surface covers read/reporting, masters, vouchers, GST registers, outstandings, validation, presence checks, and import verification well. However, descriptions repeatedly reference post_import and build_import_xml, yet neither tool appears in the exposed set, leaving a notable gap in the import dispatch lifecycle. This is a meaningful missing operation for a bridge that otherwise supports import validation and verification.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables natural language interaction with TallyPrime accounting software via Claude AI, allowing querying reports, creating vouchers, and managing ledgers without manual navigation.
    17
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables Large Language Models to access and query Tally Prime ERP data, including financial reports, masters, and inventory summaries, via the Model Context Protocol.
    85
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Connects Tally Prime ERP data to AI assistants via MCP, enabling natural language queries for financial reports, stock summaries, and ledger balances.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Bridges Tally Prime ERP with AI assistants, enabling querying financial reports, managing masters, creating vouchers, and analyzing GST data through natural language.
    AGPL 3.0