Skip to main content
Glama

schedule-iii

Server Details

Deterministic Schedule III statements for Indian companies: trial balance in, Excel workbook out.

If you are the author of this connector, you can claim ownership with GitHub, an HTTP challenge, or a DNS record. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP
URL
Tool DescriptionsA

Average 4.5/5 across 67 of 67 tools scored. Lowest: 3.5/5.

Server CoherenceA
Disambiguation5/5

Each tool targets a distinct resource or action — get_* reads, save_* writes, confirm_* approves, preview_* shows consequences before approval. Even the management-data trio (budgets, allocations, variance) is cleanly separated by surface. Two-step flows like preview_chart_rebaseline → confirm_complete_chart are clearly sequenced, so an agent won't confuse the stages.

Naming Consistency5/5

Tool names follow a highly consistent verb_noun pattern: get_* for reads, list_* for discovery, save_* for section writes, confirm_* for approvals, create_* for new entities/centres, preview_* for pre-approval checks. The few one-offs (ingest_upload, upload_trial_balance, set_header_row) still fit the verb-first convention. No camelCase or style mixing.

Tool Count2/5

At 67 tools this is well past the 'too many' threshold. While the Schedule III domain genuinely is broad — statutorily mandated sections, two-phase approval flows, readiness checks, and a separate management-data area — the surface is heavy; an agent will spend real effort just surveying the tool list. Some consolidation of the save_reserves/provisions/assets movements or merging preview+confirm pairs is possible.

Completeness4/5

The surface covers the full lifecycle: upload → mapping/costing → grouping → capture (all statutory sections) → declarations → readiness → generate → finalise → download, plus entity setup and consolidated statements. Minor gaps: no tool directly exposes historical version diffing beyond list_snapshots, and the management-data section (budgets, allocations, variance) feels bolted on rather than integral to the core flow.

Available Tools

67 tools
assert_previous_year_no_activityRecord that the previous year genuinely had no activityAInspect

Record your user’s assertion that the previous year genuinely had no activity, so this statement’s all-nil previous-year Profit and Loss is a fact rather than a closing-entry artefact. This clears Datavrn’s refusal to finalise it. THIS IS A PROFESSIONAL ASSERTION ABOUT A CLIENT’S ACCOUNTS, RECORDED IN THE NAMED MEMBER’S NAME AND KEPT WITH THE STATEMENT. Only send it when your user has told you, in their own words, that the previous year had no activity — a dormant entity is the case it is for. NEVER write the reason yourself and never paraphrase it into something firmer: send what they said. If they are unsure, or if the previous year DID trade and the trial balance was simply taken after closing, do not call this — the other two remedies in get_comparative_source_state are the correct ones. Call get_comparative_source_state first, read out the finding, get their explicit go-ahead, and send the confirm_token it returned. Datavrn refuses if the previous-year figures changed after you read the state, and nothing is recorded. It can be withdrawn in Datavrn if it was recorded in error; withdrawing restores the refusal.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesYour user’s own words: why the previous year had no activity. Recorded with the statement and shown to whoever reviews it.
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
confirm_tokenYesThe approval from get_comparative_source_state. Call it first and present its finding to your user.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint=false annotation (which already marks it as a write), the description discloses significant behavioural traits: it 'clears Datavrn’s refusal to finalise', that Datavrn refuses if previous-year figures changed after reading state, and that the record can be withdrawn (restoring the refusal). It also stresses that the assertion is recorded in the named member's name and kept with the statement. This is rich, non-obvious context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but appropriately so given the professional and irreversible nature of the action. It is front-loaded with the core purpose and conditions, then details the workflow and caveats. A few sentences could be tightened, but every sentence serves a purpose, so it is well-structured and not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no output schema, the description covers all necessary context: prerequisites (get_comparative_source_state), the exact user interaction (read out finding, get go-ahead), failure modes (changed figures), and withdrawal (restores refusal). An agent has everything needed to call it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter already has a description, but the tool description adds essential semantic guidance: reason must be 'your user’s own words' and never paraphrased, confirm_token comes from get_comparative_source_state, and on_behalf_of has API-key vs OAuth rules. This is exactly the kind of extra meaning that helps an agent use parameters correctly beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('record') and a precise resource ('your user's assertion that the previous year genuinely had no activity'), along with the broader intent (making the all-nil previous-year P&L a fact, not a closing-entry artefact). It distinguishes this from the 'other two remedies' in get_comparative_source_state, so an agent can separate it from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is explicitly conditioned: 'Only send it when your user has told you, in their own words, that the previous year had no activity' and 'If they are unsure... do not call this'. It also spells out the required workflow: call get_comparative_source_state first, read out the finding, get explicit go-ahead, and send the confirm_token. This leaves no ambiguity about when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_capture_reviewConfirm a section has been reviewedA
Destructive
Inspect

Record your user’s confirmation that they have REVIEWED a whole section and it is complete — the entire previous-year comparative column, or the entire disclosure set. A review confirmation is your user’s professional assertion, recorded as authorised by them. Before calling this, show them what you are confirming — the whole comparative, or the whole disclosure set — and get their explicit go-ahead. Never confirm a review that has not happened. Saving figures or text does NOT complete these two sections and never has; only this confirmation does. The confirmation is pinned to the exact set that was reviewed, so ANY later save to that section withdraws it — if a confirmation appears not to stick, the next step is to re-review and confirm again, never to retry. Confirming again after such a change supersedes the earlier confirmation: it is marked withdrawn (it stays on the record) and the response names what was withdrawn. Re-confirming also invalidates any finalise approval you already hold. Only two sections take a review confirmation: the previous-year comparative and the disclosure set. Every other section is answered by saving its rows, or by declare_capture_na. Datavrn notifies the member you name that this was recorded in their name. Generate a fresh version after your last capture change — finalisation checks the version’s frozen capture state, not today’s.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
client_idYesThe entity (client) id — from list_clients.
period_idYes
reason_codeYes
template_idYes
capture_kindYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations already marking destructiveHint=true and readOnlyHint=false, the description adds deep behavioral context: confirmation is pinned to the reviewed set, later saves withdraw it, re-confirmation supersedes and marks the prior as withdrawn, it invalidates finalise approvals, notifies the named member, and requires a fresh version. This goes far beyond the annotations and is critical for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense. It front-loads the core purpose, then addresses exclusions, consequences, and prerequisites in a logical order. Every sentence carries meaningful guidance (e.g., the version-generation note, the notification behavior). While it could be trimmed slightly, the complexity of the tool justifies the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of destructive side effects, the description covers the essential context: prerequisites (show and get go-ahead), exclusions (only two sections), behavioral consequences (withdrawal, supersession, invalidation), and operational requirements (fresh version). It does not describe the return value, but no output schema is present and the description implies what the response names (withdrawn confirmations). Overall, it is sufficiently complete for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description should compensate, but it focuses on behavior rather than parameter details. It does explain that reason_code enumerates the two valid sections, and that capture_kind likely indicates which capture type, but it doesn't explicitly map parameters to their roles. The schema itself provides descriptions for client_id and on_behalf_of, but note is undocumented in both. The description adds context that helps infer some parameters, but not enough to fully compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Record your user's confirmation that they have REVIEWED a whole section') and clearly identifies the two resources ('previous-year comparative column' and 'entire disclosure set'). It distinguishes itself from saving operations, saying 'Saving figures or text does NOT complete these two sections'. This is precise and removes ambiguity against sibling tools like save_disclosures or save_py_values.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance ('Before calling this, show them what you are confirming... and get their explicit go-ahead') and when-not-to-use ('Never confirm a review that has not happened'). It also explains that every other section uses saving or declare_capture_na, and explicitly warns against retries when confirmation doesn't stick. This fully routes the agent to the correct context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_centre_mappingsConfirm cost-centre mappingsA
Destructive
Inspect

Persist only the explicit account-to-centre decisions the user approved. Before calling, show the proposal grouped by confidence tier and target with exact counts, call out every medium/low-confidence row, and get a clear approval for the enumerated items. Omitted accounts stay unchanged; there is no apply-all, auto-confirm, or use-suggestions flag. After the write, report confirmed, unmapped_total, and unmapped_with_balance so the user knows exactly what remains. This tool never returns rupee amounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesEvery account decision explicitly approved by the user; never a blanket flag.
client_idYesThe entity (client) id — from list_clients.
removal_countNoThe exact removal count returned by the removal preview.
removal_tokenNoOnly include the short-lived token returned by the removal preview for this exact proposal.
effective_fromYesThe effective date shown to the user.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (destructiveHint=true, readOnlyHint=false), the description discloses key behaviors: omitted accounts remain unchanged, no auto-confirm, and the tool never returns rupee amounts. It also states the post-write report fields (confirmed, unmapped_total, unmapped_with_balance), adding clarity about side effects and output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences, each serving a purpose: purpose, pre-call requirement, behavioral constraints, output summary, and a safety note. It is front-loaded with the primary action and remains concise, though slightly longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by explaining what the agent should report after the write. Combined with the detailed input schema and annotations, it provides enough context for an agent to use the tool correctly, including preconditions and post-conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add new parameter-level meaning beyond the schema; it references 'enumerated items' and approval but doesn't elaborate on removal_token or effective_from specifics, which are already described well in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Persist only the explicit account-to-centre decisions the user approved,' which specifies a clear verb (persist) and resource (account-to-centre decisions). This distinguishes it from sibling tools like confirm_column_mapping or confirm_groupings by focusing on centre mappings specifically.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use guidance: 'Before calling, show the proposal grouped by confidence tier... and get a clear approval.' It also indicates a when-not by stating there is no apply-all or auto-confirm flag. However, it does not explicitly name alternative tools for generating proposals, only implies them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_column_mappingConfirm column mappingA
Destructive
Inspect

Confirm the column→field mapping for a staged upload and run validation. Returns the full validation result (row counts, warnings, blocking issues). Mapping suggestions are never auto-applied — pass exactly the mapping your user approved. Review any warnings with your user before ingesting.

ParametersJSON Schema
NameRequiredDescriptionDefault
mappingYes
optionsNo
save_asNo
upload_idYesThe upload session id returned by upload_trial_balance.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true, but the description adds useful non-obvious behaviors: it runs validation and returns the full result (row counts, warnings, blocking issues), never auto-applies mapping suggestions, and requires user-approved mappings. It also instructs reviewing warnings before ingesting. This goes beyond what annotations alone convey, though it doesn't explicitly describe side effects or what is destroyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each earning its place: the first defines the action, the second states the return value, and the third provides a critical caveat. It is front-loaded with the main purpose and wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a nested `options` object and no output schema, the description does well to state the return value (full validation result with row counts, warnings, blocking issues). It also places the tool in the workflow ('staged upload', 'before ingesting'). However, it omits any explanation of the optional `options` parameters and `save_as`, which would be needed for full understanding of all inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only upload_id has a description). The description partially compensates by explaining that 'mapping' must be the exact mapping the user approved, and it references the upload_id as coming from a staged upload. However, the nested 'options' object and 'save_as' parameter remain completely unexplained, leaving a significant gap for a tool with four parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action: 'Confirm the column→field mapping for a staged upload and run validation.' It names the exact resource (column→field mapping) and the verb (confirm/run validation), distinguishing it from sibling confirm_* tools that target other mappings (e.g., centre_mappings, groupings, reporting_lines).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives contextual workflow guidance: 'for a staged upload' and 'before ingesting.' It also provides a key directive: 'Mapping suggestions are never auto-applied — pass exactly the mapping your user approved' and advises reviewing warnings before ingestion. While it doesn't explicitly compare to alternative tools, it clearly implies this is the step after upload_trial_balance and before ingest_upload.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_complete_chartConfirm an upload is the entity’s whole chartAInspect

Record your user’s explicit confirmation that one upload is an entity’s COMPLETE CURRENT chart of accounts, so Datavrn starts treating it as evidence when checking whether future files belong to that entity. Recorded as authorised by the member you name. THIS IS A PROFESSIONAL JUDGMENT, NOT A CALCULATION. Datavrn will never make it from the numbers, and neither should you: call preview_chart_rebaseline for that upload first, read its counts to your user, and ask them plainly whether that upload is the entity’s whole book now. Confirming a file that is NOT the entity’s chart teaches Datavrn the wrong chart and weakens wrong-entity detection for that entity from then on. Send back the state_digest and confirm_token from the SAME preview_chart_rebaseline response, unchanged. Datavrn refuses and changes nothing if: the entity’s uploads moved after you were shown those figures, someone already confirmed this upload, the upload no longer needs confirming, or it is one Datavrn asked about directly — for those, answering that question is what settles it. Nothing about the landed figures changes: this affects only which uploads count as evidence of identity.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
state_digestYesThe state_digest from the same preview_chart_rebaseline response, unchanged. It pins the figures your user was shown; a different one is refused.
confirm_tokenYesThe approval from preview_chart_rebaseline. Call it first and present its counts to your user.
source_snapshot_idYesThe upload’s id — the `snapshot_id` field of a list_chart_rebaselines row.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/destructiveHint annotations, the description explains that this action marks the upload as evidence, affects future wrong-entity detection, refuses and changes nothing under certain conditions, and does not alter landed figures. This is substantial behavioral context not available from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured: purpose first, then required workflow, warnings, and refusal conditions. Every sentence contributes to safe use, though it is more verbose than strictly necessary for a five-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Prerequisites, side effects, refusal conditions, and authorization are all covered. The main gap is that it does not describe what a successful response contains, especially since there is no output schema. This is minor for a confirmation action but still a slight omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds critical invariants: state_digest and confirm_token must come unchanged from the same preview_chart_rebaseline response, source_snapshot_id comes from list_chart_rebaselines, and on_behalf_of is required on API-key connections but omitted on OAuth. This meaningfully exceeds schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: recording explicit confirmation that an upload is an entity's complete current chart of accounts. It clearly differentiates this from sibling confirm_* tools by tying it to preview_chart_rebaseline and emphasizing it is a professional judgment, not a calculation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs the agent to call preview_chart_rebaseline first, read counts, ask the user plainly, and send back the state_digest and confirm_token from the same response unchanged. It also lists conditions under which Datavrn refuses and notes when answering a direct Datavrn question is the alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_groupingsConfirm account groupingsA
Destructive
Inspect

Persist USER-approved account→line groupings. Omitted accounts stay unchanged. Only explicit leaf_code:null clears a saved grouping. When clearing a saved grouping, use the current grouping_version from list_grouping_suggestions. An actual clear first returns an approval request; nothing changes then. Resend the unchanged request with the approval details to proceed. Clearing an already-unclassified account is an idempotent no-op. Every row must be explicit — there is deliberately no "apply all suggestions" option.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
decisionsYes
template_idYes
removal_countNo
removal_tokenNo
grouping_versionNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations already marking destructiveHint=true, the description adds substantial behavioral detail: omitted accounts stay unchanged, only explicit leaf_code:null clears, clearing returns an approval request before any change, and the exact resend/approval flow. It also discloses idempotent behavior for already-unclassified accounts and deliberately no apply-all option, far exceeding what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Six sentences deliver multiple high-value caveats without redundancy. The opening sentence states the primary purpose, followed by concise, ordered details on omission, clearing, approval flow, idempotency, and the no-apply-all constraint. Every sentence contributes operational guidance, and the length is appropriate for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex two-step clearing behavior and no output schema, the description covers most critical workflow aspects well. It explains what changes, what does not, idempotency, and the approval request flow. It falls short only in not naming the exact removal parameters or describing the success response, leaving slight ambiguity around the approval details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, so the description carries heavy responsibility. It clarifies critical parameter semantics: decisions must be explicit, leaf_code:null means clear, and grouping_version must be current from list_grouping_suggestions. However, it does not explicitly explain removal_count and removal_token, only hints at 'approval details', and optional fields like suggested_reason, is_cash_equivalent, and suggested_confidence are not addressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Persist USER-approved account→line groupings', a specific verb-plus-resource statement that clearly identifies the tool's function. It distinguishes itself from sibling confirm_* tools by naming the exact resource type (account groupings) and emphasizes it is a user-approved confirmation action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: after user approval, with explicit rows, and notes 'no apply all suggestions' alternative. It also names a prerequisite source (list_grouping_suggestions) for grouping_version when clearing. However, it does not explicitly compare against sibling confirm_* tools, though the resource-specific wording makes the distinction implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_reporting_linesConfirm reporting-line mappingsA
Destructive
Inspect

Persist only the explicit reporting-line decisions the user approved. Before calling, show the proposal grouped by confidence tier and target with exact counts, flag every medium/low-confidence row, and get clear approval for the enumerated decisions. Omitted accounts stay unchanged. Sending leaf_code:null permanently removes that account saved reporting line; send it only when the user explicitly asked to clear that row. There is no apply-all or auto-confirm flag. The response tells you how many were confirmed, cleared, and whether the balance-bearing set is fully mapped; never claim completion without checking those fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
decisionsYes
removal_countNoThe exact removal count returned by the removal preview.
removal_tokenNoOnly include the short-lived token returned by the removal preview for this exact proposal.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, and the description adds substantial behavioral detail: 'Sending leaf_code:null permanently removes that account saved reporting line,' 'Omitted accounts stay unchanged,' and the warning to never claim completion without checking the response fields. This goes well beyond the annotations and discloses important side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences and every sentence carries essential information—purpose, preconditions, null semantics, no apply-all, and response interpretation. It is front-loaded with the core action, but the density of warnings makes it longer than absolutely minimal, which is appropriate for a destructive tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—four parameters, a nested decisions array, destructive behavior, and no output schema—the description is remarkably complete. It covers prerequisites, user approval steps, edge cases, response fields to check, and the caution not to assume completion. This fully compensates for the missing output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description enriches the most critical parameter: leaf_code:null is explicitly linked to permanent removal, and the decisions array is clarified by noting omitted accounts stay unchanged and there is no auto-confirm flag. It does not explain removal_token, but the schema already describes its purpose adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Persist only the explicit reporting-line decisions the user approved'—a specific verb (persist) and resource (reporting-line decisions). It clearly distinguishes from sibling confirm tools like confirm_centre_mappings and confirm_column_mapping by narrowing scope to reporting-line mappings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Before calling, show the proposal grouped by confidence tier... and get clear approval.' It also gives when-not-to-use cautions, such as sending leaf_code:null only when the user explicitly asked to clear the row, and notes there is no apply-all flag. However, it does not explicitly name alternative tools, so it falls short of a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

copy_capture_declarationsCopy last period’s capture answersAInspect

Copy the previous period’s "nothing this period" and "does not apply" answers into this period, for sections that have no answer yet. It NEVER copies a review confirmation — a review is about this period’s content and cannot be inherited. Last period’s answer is not evidence about this period: list what it would copy to your user, section by section, and get their go-ahead before calling it. Answers already recorded for this period are left alone. Generate a fresh version after your last capture change — finalisation checks the version’s frozen capture state, not today’s.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With all annotations false, the description carries the full burden and delivers: it never copies review confirmations, skips sections with existing answers, and warns about the version/finalisation dependency ('finalisation checks the version’s frozen capture state'). These are non-obvious behaviors not otherwise visible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each contributing distinct information: the core action, exclusions, user approval requirement, and versioning caveat. No filler or redundancy, and the key action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose, exclusions, prerequisites, existing-answer preservation, and finalisation impact. No output schema exists, so return-value documentation is not needed. It is sufficiently complete for a tool with this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (client_id and on_behalf_of have descriptions). The description indirectly clarifies period_id by distinguishing 'previous period' from 'this period', but does not explain template_id or map the parameters explicitly. It adds some semantic value but leaves gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the exact action: 'Copy the previous period’s "nothing this period" and "does not apply" answers into this period, for sections that have no answer yet.' It also distinguishes itself from sibling tools by explicitly noting it never copies a review confirmation, which separates it from confirm_capture_review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear usage context: 'get their go-ahead before calling it' sets a prerequisite, and 'Answers already recorded for this period are left alone' defines scope. It also excludes review confirmations, giving clear when-not-to-use guidance, even though it does not name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_clientCreate entityAInspect

Create a new entity (company) in the organization. Requires a Manager-role key. Use only when the user explicitly asks to set up a new entity; show them the name you will create first.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
business_model_codeNo
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false (write operation) and destructiveHint=false. The description supplements this with a 'Manager-role key' requirement and the need to show the name to the user first, adding meaningful context about auth and confirmation behavior. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of three concise sentences, front-loaded with the core purpose, then the auth requirement, and finally the usage/confirmation guideline. Every sentence adds value with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, auth requirements, and a usage constraint, which is good given the simple tool signature. However, it omits the optional business_model_code parameter entirely and provides no indication of return values or post-create behavior, leaving some important context for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only vaguely references 'the name' but does not explain the required 'name' parameter or the optional 'business_model_code' parameter, their constraints, or their purpose. This leaves an agent without crucial parameter-level understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Create') and a clear resource ('new entity (company)'), which distinguishes it from sibling tools like create_cost_centre and create_profit_centre that create different resource types. It clearly identifies what the tool does without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit condition for use: 'Use only when the user explicitly asks to set up a new entity' and adds a required confirmation step ('show them the name you will create first'). However, it does not mention alternative tools or when not to use it beyond the 'only when' clause, so it misses the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_cost_centreCreate a cost centreAInspect

Create one cost centre for an entity after showing the user the exact name, kind, parent, effective date, and reason. This is one explicit centre at a time; there is no apply-all shortcut. After creating it, call list_cost_centres again and explain which mapping suggestions can now use it.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNoOptional short code.
nameYesThe cost-centre name to create.
reasonYesWhy the user asked for this centre.
client_idYesThe entity (client) id — from list_clients.
centre_kindNoOperating or shared-support centre; defaults to operating.
descriptionNoOptional plain-language description.
effective_fromNoDate from which this centre applies; defaults to the start of the entity data.
parent_cost_centre_idNoOptional existing parent cost-centre id.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a write operation (readOnlyHint=false). The description adds valuable behavioral context: the user must be shown the exact parameters before creation, there is no bulk/apply-all behavior, and a post-create verification step is required. This goes beyond the basic write semantics and helps the agent understand the intended interaction flow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, each earning its place: the first states the action and required confirmation, the second clarifies scope (one at a time, no shortcut), and the third gives the mandatory follow-up action. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with 8 parameters and no output schema, the description covers the essential workflow: pre-conditions (user confirmation), constraints (one at a time), and post-conditions (call list_cost_centres). It does not explain return values, but the follow-up instruction effectively compensates by telling the agent how to verify the result and continue the workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers all 8 parameters with descriptions (100% coverage), so the baseline is 3. The description adds meaning by identifying which parameters are user-facing and must be confirmed (name, kind, parent, effective date, reason), and by emphasizing that only one centre is created at a time. This is a helpful layer above the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Create one cost centre for an entity') and specifies the exact fields to be confirmed with the user (name, kind, parent, effective date, reason). It also distinguishes this tool from a bulk or apply-all operation by explicitly saying there is no shortcut, setting it apart from sibling tools like create_profit_centre.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: show the user the exact details before creating, create one at a time, and follow up by calling list_cost_centres again to explain which mapping suggestions can use the new centre. It does not explicitly mention alternatives or when not to use this tool, but the workflow is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_profit_centreCreate a profit centreAInspect

Create one profit centre for an entity after showing the user the exact name, optional parent, and description. This is one explicit centre at a time; there is no apply-all shortcut. Re-list the centres after creation so the user can see the new target before any mapping confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe profit-centre name to create.
client_idYesThe entity (client) id — from list_clients.
descriptionNoOptional plain-language description.
parent_profit_centre_idNoOptional existing parent profit-centre id.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a mutating, non-destructive operation. The description adds valuable behavioral context: it confirms a user-confirmation step before creation and states that centres are re-listed afterwards, which goes beyond the basic annotation flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the primary action, and every sentence adds value: scope, a limitation, and a post-creation behavior. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 params, no output schema) and available annotations, the description covers purpose, workflow (confirmation), and post-creation listing. It does not detail return values or error scenarios, but the mention of re-listing centres supplies practical context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are fully documented in the schema. The description mentions "exact name, optional parent, and description" which maps to name, parent_profit_centre_id, and description but does not add additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states "Create one profit centre for an entity" with a specific verb and resource, distinguishing it from siblings like create_cost_centre. It also notes there is no apply-all shortcut, reinforcing its single-resource scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage context: create one profit centre at a time, after showing the user the exact name/parent/description. It also states an explicit limitation (no apply-all shortcut), though it does not name alternative tools or explicit when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

declare_capture_naRecord nothing to report, or not applicableA
Destructive
Inspect

Record that a capture section had NOTHING to report this period, DOES NOT APPLY to this entity, or that this is the entity’s FIRST YEAR (previous-year figures only). These are three different statements and are not interchangeable: "nothing this period" means the section applies but had no activity; "does not apply" means it never applies to this entity at all. This is your user’s professional assertion, recorded as authorised by them — ask which one is true, and never guess. NOT every reason is available for every section — call get_schedule3_workspace and read allowed_reason_codes on the section before you ask your user, so you never put a choice to them that Datavrn will refuse. The restrictions: Settings takes NO answer here at all (it is only answered by saving the settings); share capital and partner capital take only "nothing this period", because those sections are shown only for statement formats they apply to, so "does not apply" can never be true; and "first year" belongs only to previous-year figures. To record a REVIEW being complete (previous-year figures, disclosures) use confirm_capture_review instead; this tool cannot make that assertion. A section can only hold one active answer. If the section’s current answer is still standing, this call refuses — revoke it with revoke_capture_declaration first. If the current answer was already WITHDRAWN by later changes (Datavrn shows the section as unanswered), this call supersedes it: the old answer is marked withdrawn and stays on the record, and the response names what was withdrawn. An answer is never edited in place. Superseding or changing an answer invalidates any finalise approval you already hold — the next finalise_statement will ask you to review the current state again. Generate a fresh version after your last capture change — finalisation checks the version’s frozen capture state, not today’s.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
client_idYesThe entity (client) id — from list_clients.
period_idYes
reason_codeYes
template_idYes
capture_kindYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description goes well beyond: it explains superseding behavior (old answer marked withdrawn), refusal when a current answer stands, the need to revoke first, invalidation of finalise approvals, and the requirement to generate a fresh version. This adds rich contextual detail about side effects not captured in annotations, and there is no contradiction with them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence conveys essential information. It is front-loaded with the core purpose, then moves to restrictions, then to mutational behavior. Each paragraph handles a distinct topic without redundancy. It is dense and well-structured, though slightly lengthy; for a tool with this many nuances, the length is justified and earn-worthy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and low schema coverage, the description covers all critical operational aspects: the three declaration types, allowed reason codes per section, the supersede/withdraw/revoke lifecycle, invalidation of approvals, and version generation. It also mentions what the response names (the withdrawn answer). Nothing an agent needs to safely invoke this mutation tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 29%, so the description must compensate. It does for the most critical parameters: reason_code semantics are fully explained (three meanings and their non-interchangeability), and capture_kind restrictions are detailed (settings no answer, share capital/partner capital only 'nothing this period', 'first year' only for previous-year figures). However, it does not clarify the 'note' parameter or explain period_id/template_id beyond their obvious names. Since the core decision logic is covered, this is a strong but not perfect compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Record that a capture section had NOTHING to report this period, DOES NOT APPLY to this entity, or that this is the entity’s FIRST YEAR'. It clearly distinguishes the three distinct statements and explicitly differentiates from the sibling confirm_capture_review by stating that tool makes a different assertion. The purpose is unambiguous and separate from other tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: it tells the agent to call get_schedule3_workspace and read allowed_reason_codes before asking the user, so invalid choices are never offered. It also names an alternative (confirm_capture_review) for review completion, and provides per-section restrictions (settings takes no answer, share capital/partner capital only 'nothing this period'). This is thorough, actionable usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finalise_statementSeal the permanent client copy of a versionA
Destructive
Inspect

Seal a statement version as Datavrn’s permanent client copy, recorded as authorised by the member you name. THIS IS NOT APPROVAL OR ADOPTION OF THE FINANCIAL STATEMENTS AND IT IS NOT A SIGNATURE. It does not discharge section 134(1) for a company or section 34(3) for an LLP. THERE IS NO UNDO. A change afterwards means generating a new version and finalising that one; the version you seal here stays sealed. Call get_finalise_readiness first, show your user every gate and every warning it returns, get their explicit go-ahead, and only then send the confirm_token it gave you together with the acknowledgements. Never acknowledge a warning your user has not seen, and never write the acceptance reason yourself — it is their professional judgment in their own words. Datavrn will refuse if anything about the statement changed after you read the state, and nothing will be finalised. If the response comes back with reused set to true, a finalisation of this same version was already under way: nothing new was done and your approval was not used. Do not call finalise_statement again — the status field says where that finalisation is (queued, running, or paused, which means it hit a temporary problem and Datavrn will retry it on its own). Poll get_job with the job_id and tell your user what it says.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
snapshot_idYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
confirm_tokenYes
acknowledgmentsYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses critical behaviors: 'THERE IS NO UNDO', the operation is not approval or adoption, and 'Datavrn will refuse if anything about the statement changed after you read the state.' It also explains the 'reused' response scenario and the paused/retry behavior, adding substantial context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Though lengthy, the description is front-loaded with the core purpose and every sentence carries critical operational or safety information. It avoids redundancy by structuring warnings and prerequisites logically, and the complexity of the tool justifies the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, lack of output schema, and low schema description coverage, the description is exceptionally complete. It covers prerequisites, exact workflow, error handling (reused, refused), and post-finalisation verification via get_job, leaving no major contextual gaps for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds important meaning for confirm_token and acknowledgments (e.g., 'send the confirm_token it gave you', 'Never write the acceptance reason yourself'), addressing low schema coverage (40%). However, it does not explicitly explain snapshot_id or client_id beyond what the schema states, leaving some inference required for the version-to-snapshot relationship.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool does: 'Seal a statement version as Datavrn's permanent client copy, recorded as authorised by the member you name.' It clearly distinguishes from sibling tools like confirm_* or save_* by introducing the unique concept of a permanent, sealed copy and explicitly referencing finalise-specific prerequisites (e.g., get_finalise_readiness).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Call get_finalise_readiness first, show your user every gate and every warning...', and warns against re-calling: 'Do not call finalise_statement again.' It also directs follow-up via get_job, contrasting with alternative actions and clarifying the irreversible context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_schedule_iiiGenerate Schedule III statementsAInspect

Queue the Schedule III workbook build (returns a job_id to poll with get_job — the build runs as a background job). REFUSES when ungrouped accounts exist unless acknowledged: before acknowledging, present the ungrouped accounts to your user and obtain their explicit go-ahead; record it in acknowledge_reason and pass the exact count in acknowledge_count — an acknowledgement WITHOUT its count is always re-demanded. A multi-month statement period additionally requires acknowledge_multi_month_pnl WITH acknowledge_month_count (confirm with your user that the TBs are period movements, not cumulative). Never acknowledge anything the user has not seen. Once queued, the build usually completes in a few minutes — tell your user their statements are being prepared and poll get_job periodically; do not present the wait as a problem.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
acknowledge_countNo
acknowledge_reasonNo
acknowledge_month_countNo
acknowledge_unclassifiedNo
acknowledge_multi_month_pnlNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the sparse annotations (readOnlyHint=false, destructiveHint=false), the description discloses the background job nature, refusal conditions, acknowledgment dependencies, and expected user communication. It fully exposes the behavioral contract of the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense, with every sentence serving a purpose. It is front-loaded with the primary action and return value, then methodically covers refusal and multi-month handling. No wasted words given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 8-parameter schema and no output schema, the description covers the async flow, user-consent requirements, and polling guidance. It does not enumerate possible errors or describe the resulting workbook contents, but it is sufficient for correct invocation and follow-up.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 13%, but the description compensates by explaining the semantics of acknowledge_count/acknowledge_reason and acknowledge_multi_month_pnl/acknowledge_month_count, including their interdependencies. However, template_id and period_id remain largely unexplained beyond the schema, leaving a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action: 'Queue the Schedule III workbook build' and notes the async job_id return, distinguishing it from sibling get_*/save_* tools. It leaves no ambiguity about the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit instructions on handling refusal conditions (present ungrouped accounts, obtain user go-ahead, record reason and count) and directs the agent to poll get_job after queueing. This is precise when-to-use and how-to-proceed guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_allocation_account_figuresRead allocation account figuresA
Read-only
Inspect

MANAGEMENT data class. Read the current persisted allocation run at account grain: books figure plus spreading adjustment equals MIS figure, all as decimal-string rupees. Filter account names or minimum absolute MIS amount before paging. The summary covers the full filtered set and ties the spreading reconciliation; no target-level split or source transactions are returned. The signed page_token is source-pinned, so restart at page 1 if source_changed. known_stale and not_assessed disclose run state; neither means fresh.

ParametersJSON Schema
NameRequiredDescriptionDefault
periodYesManagement month in YYYY-MM.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per source-pinned page (default 50, max 200).
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
min_abs_misNoMinimum absolute MIS figure in rupees as a decimal string, e.g. '100000'.
account_name_patternsNoUp to 10 case-insensitive account-name substrings; any match is retained.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, openWorldHint, destructiveHint), the description discloses significant behavioral traits: the output format is decimal-string rupees, the summary covers the full filtered set and ties the spreading reconciliation, no target-level split or source transactions are returned, page tokens are source-pinned and expiration behavior is described, and run-state fields (known_stale, not_assessed) are explained. This far exceeds the annotation coverage and gives the agent a precise mental model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and information-dense. It leads with 'MANAGEMENT data class' and a clear statement of purpose, then efficiently layers in the data model, filtering, paging, summary behavior, and run-state semantics across four sentences. Every sentence adds unique information, and there is no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 6 parameters, no output schema, and no nested objects, the description is remarkably complete. It explains the output grain, the arithmetic relationship of figures, filtering options, paging behavior, what data is excluded, and how to interpret run-state flags. It leaves no major ambiguities about what the tool returns or how to handle pagination edge cases, making it fully self-sufficient for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 6 parameters with descriptions (100% coverage), so the baseline is 3. The tool description adds semantic value by explaining the data context (books + spreading = MIS, decimal-string rupees), which helps understand parameters like min_abs_mis. It also clarifies the interaction between filtering and paging, and reinforces the source-pinned nature of page_token. While it doesn't describe each parameter individually, the holistic context meaningfully aids parameter understanding, so a 4 is justified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read') and a specific resource ('the current persisted allocation run at account grain'), and explicitly clarifies the data model (books + spreading adjustment = MIS figure). It distinguishes itself from the sibling get_allocation_target_figures by noting that no target-level split is returned, making the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: it reads management allocation data at account grain, supports filtering before paging, and clarifies that target-level splits are excluded. It also gives explicit paging guidance (restart at page 1 if source_changed) and explains how to interpret run-state flags. However, it does not name a direct alternative tool or state explicit 'when not to use' conditions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_allocation_target_figuresRead allocation target figuresA
Read-only
Inspect

MANAGEMENT data class. Read the current persisted allocation run at account × target × source × mode grain, with decimal-string allocated amounts. For step-down, filter allocation_sources=['step_down']; for profit centres, filter target_types=['profit_centre','sub_profit_centre']. Use filters before paging; the signed page_token is source-pinned, so restart at page 1 if source_changed. Target labels are cost-centre names as of the run and current profit-centre names; no owners, descriptions, source transactions, or recomputation are returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
periodYesManagement month in YYYY-MM.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per source-pinned page (default 50, max 200).
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
target_typesNoRestrict to allocation target types.
min_abs_amountNoMinimum absolute allocated amount in rupees as a decimal string, e.g. '100000'.
allocation_sourcesNoRestrict to allocation sources.
account_name_patternsNoUp to 10 case-insensitive account-name substrings; any match is retained.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, but the description adds valuable behavior: 'the signed page_token is source-pinned, so restart at page 1 if source_changed' and 'Target labels are cost-centre names as of the run and current profit-centre names'. This exceeds the annotation-only safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with 'MANAGEMENT data class' and the grain. No filler—each sentence adds unique information: data grain, filtering patterns, pagination caveat, and label semantics/exclusions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters and no output schema, the description covers the essential context: grain, amount format, recommended filters, pagination behavior, and what is not returned. This gives an agent enough to invoke correctly and set expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are documented. The description adds meaning beyond the schema by explaining how to use enum values (allocation_sources, target_types) and the page_token behavior ('source-pinned, restart at page 1 if source_changed'). Minor gap: doesn't elaborate on min_abs_amount or account_name_patterns, but the schema covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Read the current persisted allocation run at account × target × source × mode grain, with decimal-string allocated amounts.' It clearly distinguishes from the sibling get_allocation_account_figures by specifying the target grain and returned data format.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: 'For step-down, filter allocation_sources=["step_down"]; for profit centres, filter target_types=["profit_centre","sub_profit_centre"]' and 'Use filters before paging'. It also states exclusions ('no owners, descriptions, source transactions, or recomputation are returned'), making it clear when NOT to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_budgetRead budgetA
Read-only
Inspect

MANAGEMENT data class. Read one versioned budget: identity-free header, its pinned P&L tree, and filtered/paginated cells with entered-versus-inferred truth. Amounts and locked FX rate are decimal strings. Filter months, lines, centres, or inference before paging; the signed page_token is source-pinned, so restart at page 1 if source_changed. Cells across different P&L lines are not one meaningful grand total, so no cross-line grand total is exposed. No people, ownership, editability, or approval identities are returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
monthNoOptional budget month in YYYY-MM.
budget_idYesBudget id from list_budgets.
client_idYesThe entity (client) id — from list_clients.
inferenceNoCell inference state (default all).
page_sizeNoCells per source-pinned page (default 100, max 200).
line_codesNoUp to 25 P&L line codes.
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
cost_centre_idsNoUp to 25 cost-centre ids.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=true, destructiveHint=false), the description discloses substantive behavioral traits: amounts and FX rate are decimal strings, page_token is source-pinned and requires restart on source_changed, no cross-line grand total is exposed, and identity-related fields are omitted. This goes well beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense but well-structured: it opens with the core purpose, then covers data types, pagination/filtering, a caveat about grand totals, and exclusions. Every sentence contributes unique value, and the text is front-loaded with the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 parameters and no output schema, this description is notably complete. It explains the return contents (header, P&L tree, cells), pagination/filtering semantics, and important limitations (no grand total, no identity fields), giving an agent enough context to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 8 parameters have schema descriptions (100% coverage), setting a baseline of 3. The description adds extra meaning by mentioning filter dimensions ('months, lines, centres, or inference') and the source-pinned page_token behavior, which clarifies how parameters interact beyond the schema's individual property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Read') and a distinct resource ('one versioned budget'), and it enumerates the returned components (identity-free header, pinned P&L tree, filtered/paginated cells). It differentiates itself from sibling get_* tools by focusing on budget data and including a data-class qualifier ('MANAGEMENT').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context about what the tool returns and notes exclusions (no people/ownership/approval identities), but it does not explicitly indicate when to use this tool versus alternatives like list_budgets or get_tb_rows. Usage guidance is implied by the tool name and the first sentence, but no explicit alternatives or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_comparative_source_stateCheck the previous-year comparative sourceAInspect

Check whether this statement’s previous-year comparative can be sealed, and get the approval assert_previous_year_no_activity needs. Datavrn refuses to finalise a statement whose previous-year figures come from a trial balance drawn AFTER the year-end closing entries: that derives a previous-year Profit and Loss of all zeroes which foots perfectly and is not last year’s results. When post_closing_detected is true, READ THE WHOLE finding TO YOUR USER — what the state is and all three ways out — and let them choose. Never choose for them. Two of the three remedies are things only they can do (upload the pre-closing trial balance, or enter last year’s signed figures as previous-year values, then generate a fresh version). The third is an assertion that the previous year genuinely had NO ACTIVITY, which is a statement about their client’s accounts, in their words, recorded in their name — a dormant company is the case it exists for. The approval is single-use, expires in 15 minutes, and is tied to this statement, this connection and the member you name; if the previous-year figures change in between, the assertion will be refused and you start again from here. If nothing is wrong there is no approval to hand back, because there is nothing to assert.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide readOnlyHint=false and destructiveHint=false, so the description carries the transparency burden. It fully discloses side effects: the approval is single-use, expires in 15 minutes, tied to statement/connection/member, and can be refused if figures change. It also explains that this tool can return a post_closing_detected flag and instructs the agent to report the full finding to the user, which is a behavioral mandate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is ~140 words, long but densely packed with necessary safety-critical information. It front-loads the core purpose, then explains the problem, remedies, and constraints in a coherent flow. No redundancy, but it is long; however, given the complexity and the need to instruct the agent on handling the finding, it is justified. Slightly exceeds ideal conciseness but earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the output indirectly (post_closing_detected flag, approval object) and the exact actions the agent must take. It covers the main scenarios (post-closing detected, nothing wrong) and all constraints (expiry, ties, refusal). No output schema exists, so the description must compensate, and it does so well, though it doesn't enumerate the full return structure (e.g., whether it returns the three remedies explicitly). Minor gap, but sufficient for correct usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed descriptions for each parameter (uuid formats, template examples, on_behalf_of email rules). The tool description does not add meaning beyond the schema; it focuses on behavior and workflow rather than parameter specifics. Baseline 3 is appropriate as the schema already covers semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Purpose is crystal clear: 'Check whether this statement’s previous-year comparative can be sealed, and get the approval assert_previous_year_no_activity needs.' It names the specific verb, resource, and ties to the sibling tool. The context about the all-zeroes P&L problem distinguishes it from other tools like save_py_values or upload_trial_balance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to use it (when previous-year figures come from a post-closing trial balance) and what to do: read the whole finding, let the user choose among three remedies, never choose for them. It also states when no approval is needed. This is exemplary usage guidance, going beyond just 'when to use' to provide a decision tree.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_consolidated_statementsRead sealed consolidated statementsA
Read-only
Inspect

CONSOLIDATED data class. Read one sealed group profit-and-loss, balance-sheet, or cash-flow face for an exact periodicity and period. The response exposes presentation currency and exactly one decimal-string amount per line: section-natural for P&L/BS, signed cash movement for CFS. Amounts are persisted on the sealed run; P&L/BS labels are current display metadata and the signed page_token is source-fingerprinted across the complete safe face, so a changed value or label requires restarting at page 1. A stale sealed run is disclosed on every page and is never called current. Member names, components, eliminations, journal references and lineage are not returned. QUALIFICATIONS. A response may carry owned_share_capital_caveat (at seal, share capital owned by the group could not be eliminated, so those amounts remained inside consolidated share capital as the engine computed it; the DISPLAYED line may differ in either direction where a manual journal also moved it, so never characterise the direction from this field alone) and always carries domestic_cash_flow_caveat. These qualify specific statement lines. When a caveat is present, any figure it qualifies MUST be presented together with its qualification — never the number alone. owned_share_capital_caveat is null when the run was read and carries no such qualification, and {status: "unavailable"} when the run’s qualification record could not be read at all, which is not the same as clean: say so rather than presenting the figures as final. WITHHELD CASH FLOW. cash_flow_status may say the cash flow was not presented — for a group with a foreign member, or for a window with no opening balance sheet to measure from. An empty cash_flow then means the statement was WITHHELD, never that the group had no cash movements: say so. For the missing-opening-basis case cash_flow_refusal carries the cause, the message and remedies, each tagged with the channel that can perform it — agent_or_app you can do here, app_only needs a person in the Datavrn app. Never present an app_only remedy as something you will do. cash_flow_opening_basis states what a PRESENTED statement measured from; "not_recorded" means the run was sealed before Datavrn recorded that, so its basis is unknown — do not assume a prior period. ROW ROLES. Every statement row carries row_role: "line" participates in its section total, "total" restates it (Profit after tax), "attribution" splits a total (the amounts attributable to the owners and to the minority interest). To total a section, sum ONLY rows whose row_role is "line" — including a total or attribution row would double-count the group’s profit. Present the attribution rows as the split of profit after tax, never as additional income. Consolidated cash-flow is available only for an all-domestic group in v1. If any member uses a foreign currency, Datavrn does not present a consolidated cash-flow statement. Where it is presented, it is prepared by the indirect method from balance-sheet movements rather than from cash records, and classified into operating, investing and financing activities using each entity’s reporting-line mapping. Interest paid and taxes paid are not disclosed separately, so they remain inside the operating movement; a movement whose reporting line carries no cash classification is shown under “Unclassified movements — review” rather than assigned to an activity. Datavrn does not present other comprehensive income or total comprehensive income in the consolidated output.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupYesAn active group UUID or its exact case-insensitive display name.
periodYesThe period matching periodicity.
page_sizeNoRows per source-pinned page (default 50, max 100).
statementNoStatement face (default pnl).
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
periodicityYesPeriod grammar: monthly YYYY-MM, quarterly FYyyyy-Qn, annual FYyyyy.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a wealth of behavioral nuances beyond what annotations indicate (readOnlyHint=true, openWorldHint=false): page_token fingerprinting, stale-run disclosure, caveats like owned_share_capital_caveat and domestic_cash_flow_caveat, conditions for withheld cash flow, and row_role semantics. It goes far beyond the safety profile and explains exactly how data is presented and qualified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but deliberately structured with clear section headers (CONSOLIDATED data class, QUALIFICATIONS, WITHHELD CASH FLOW, ROW ROLES). The core purpose is front-loaded in the first sentence, and each subsequent section adds essential operational details rather than filler. Every sentence contributes to correct invocation and interpretation, justifying the length for a tool of this complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining the return format and constraints. It covers presentation currency, decimal-string amounts, row roles and summation rules, caveats and their handling, withheld cash flow semantics, and the unavailable status. It also addresses edge cases like foreign members and missing opening basis, making it remarkably complete for an agent to call and interpret correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While the schema covers 100% of parameters, the description adds meaningful semantics: page_token is described as 'signed continuation from the prior page; restart without it if source_changed', periodicity is elaborated with exact grammar (monthly YYYY-MM, quarterly FYyyyy-Qn, annual FYyyyy), and the statement enum's meaning is implied through the consolidated context. This elevates understanding beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence clearly states the verb ('Read'), the resource ('sealed group profit-and-loss, balance-sheet, or cash-flow face'), and the scope ('for an exact periodicity and period'). It distinguishes this tool from siblings like get_statement_figures or get_tb_rows by emphasizing 'sealed' and 'consolidated', leaving no ambiguity about which surface it targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: reading consolidated statements for a sealed run. It also explicitly states exclusions, e.g., 'Consolidated cash-flow is available only for an all-domestic group in v1' and explains when cash flow is withheld. While it doesn't name alternative tools, it gives strong directional guidance on use cases and constraints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_finalise_readinessCheck finalisation state and get approvalAInspect

Read the full finalisation state of one statement version, and get the approval finalise_statement needs. Call it ONCE immediately before finalising — it re-reads the stored workbook, so do not poll it. SHOW YOUR USER EVERY ROW THIS RETURNS — the gates that must be green, each warning they would be accepting and why, how many input cells are still empty, any control that could not be evaluated, and capture_live_diverged_message when it is present — before you finalise. Do not summarise the warnings away. capture_live_diverged_message means a capture answer changed after this version was generated: the version can still be sealed as it stands, and generating a fresh one is the alternative. Read it out and let your user choose. A control that "could not be evaluated" is not a pass: it is a check Datavrn did not run, and your user is entitled to know what was not checked before they seal the version. The approval is single-use, expires in 15 minutes, and is tied to this exact version, this connection and the member you name — if anything about the statement changes in between, finalise_statement will refuse and you start again from here. If your connection is not permitted to finalise, you still get the whole state; there is simply no approval to hand back.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
snapshot_idYesThe statement version id — from list_snapshots or get_schedule3_workspace.
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals important behavioral traits beyond the annotations: it re-reads the stored workbook (explaining why polling is discouraged), the approval is 'single-use, expires in 15 minutes, and is tied to this exact version, this connection and the member you name,' and it clarifies that non-permitted connections still get the full state but no approval. This adds significant context that the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but each section serves a purpose: usage timing, user display requirements, side-effect caveats, and permission behavior. It is front-loaded with the core purpose and the key warning. The length is justified by the tool's complexity, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain what the tool returns. It details every row type (gates, warnings, empty input cells, unevaluated controls, capture_live_diverged_message) and explains their meanings. It also covers the full lifecycle from call to finalise and the failure mode if the state changes, making it complete for this complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides descriptions for all three parameters (100% coverage), including format hints and which are required. The description does not add specific parameter-level semantics beyond referencing 'the member you name' and 'your user', so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Read the full finalisation state of one statement version, and get the approval finalise_statement needs.' This clearly states what the tool does and differentiates it from the sibling finalise_statement by positioning it as the prerequisite readiness check. The title also reinforces the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit timing: 'Call it ONCE immediately before finalising — it re-reads the stored workbook, so do not poll it.' This provides a clear when-to-use instruction and an exclusion (do not poll). It also explains the consequence of stale state ('finalise_statement will refuse'), which guides the agent on when to re-run the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_helpGet the agent guideA
Read-only
Inspect

Get the Datavrn agent guide: how connecting works (OAuth and API key), what an assistant can do, how reading a statement as data works, and the guarantees and limits — plus the current list of tools. Call this to answer a user's questions about how Datavrn works from canonical documentation instead of guessing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and destructiveHint=false. The description adds valuable context by specifying what the guide contains (OAuth/API key, assistant capabilities, statement-as-data, guarantees/limits, tool list) and that it serves as canonical documentation. This goes beyond the annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact: two sentences, front-loaded with the tool's purpose, followed by a structured list of contents and a direct usage instruction. Every word earns its place, with no redundancy or padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema), the description is fully sufficient. It explains what the guide covers, why to use it, and frames it as the canonical source for answering user questions, making it complete in context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, which sets a baseline of 4. The description does not need to explain parameters and instead clarifies the tool's scope, maintaining the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and names a concrete resource ('the Datavrn agent guide'), then lists its contents. It clearly distinguishes from sibling tools that perform data operations, making this uniquely about documentation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: 'Call this to answer a user's questions about how Datavrn works from canonical documentation instead of guessing.' It provides clear context but does not mention any exclusions or alternative tools, so it earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_jobGet job statusA
Read-only
Inspect

Poll a background job by id until status is succeeded or failed. A failed job carries its user-safe error reason — show it to your user. Jobs run on a background worker that claims queued work on a schedule, so a job sitting at "queued" (0 attempts) for the first few minutes is NORMAL, not a fault — keep polling every ~30–60s and reassure the user it is being prepared; do NOT report this as an error or a Datavrn bug. Only if it is still "queued" well past a few minutes should you tell the user it is taking longer than usual.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job id returned by generate_schedule_iii.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description significantly expands on the annotations. While readOnlyHint=true and destructiveHint=false are useful, the description adds critical behavioral details: the asynchronous nature of jobs, that 'queued' for a few minutes is normal, that the tool returns a user-safe error reason on failure, and that the agent should keep polling. This is far beyond the annotations and helps the agent understand the tool's runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with a clear first sentence stating the core action, followed by failure handling, then queued-status behavior, and a final clause on when to escalate. Each sentence earns its place, and the length is justified by the need to prevent misinterpretation of the queued status. There is no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description adequately covers the essential return information: statuses (succeeded, failed, queued), the presence of an error reason on failure, and the attempts count implied by '0 attempts'. It does not enumerate all possible statuses or include other fields, but for a polling tool, it provides sufficient context for an agent to act correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage for the single parameter 'job_id', stating that it is the job id returned by generate_schedule_iii. The description does not add any additional semantic detail about the parameter itself. Therefore, the baseline of 3 is appropriate, as the schema does the heavy lifting and the description adds no extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Poll a background job by id until status is succeeded or failed.' It identifies the specific resource (background job) and action (poll by id), and it distinguishes this tool from siblings by focusing on generic background job status rather than specific entity statuses. The description also clarifies the end condition, which fully captures the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives extensive usage context, including polling frequency (every ~30–60s), how to interpret the 'queued' status (normal initially), and how to communicate results to the user (show error reasons, reassure user). However, it does not explicitly name alternative tools or state when not to use this tool, so it falls short of a 5 but provides clear enough context for when it should be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_partner_capitalRead partner or owner capital scheduleA
Read-only
Inspect

Read the partner or owner capital schedule currently on file — Note 3a and Note 3b. THIS RETURNS PEOPLE’S NAMES, along with each person’s profit-sharing ratio and amounts. Call it before save_partner_capital so you can show your user what is on file and what your change would do — that save replaces the whole section, so a schedule you cannot see is a schedule you cannot safely replace. It also returns the total of the capital-account profit-sharing ratios, because Datavrn warns about a total that is not 100% only when at least two rows carry a ratio.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size (default 50, max 200).
offsetNoRows to skip (default 0).
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
account_kindNoLimit to one section: 'capital' is Note 3a, 'current' is Note 3b. Omit for both.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true and destructiveHint=false, but the description adds meaningful behavioral detail: it returns people's names, profit-sharing ratios, amounts, and the total of ratios. It also explains why the total is returned (Datavrn warns only when at least two rows carry a ratio), which goes beyond the structured annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet complete, with the core purpose in the first sentence and all additional sentences carrying actionable context (usage before save, return contents, and ratio-total behavior). No filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no output schema, the description adequately covers the return contents and why the total is included. It also ties into the save sibling. Minor gaps like pagination behavior are already covered by the schema's limit/offset descriptions, so the overall context is strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter has a thorough schema description (e.g., account_kind maps 'capital' to Note 3a and 'current' to Note 3b). The tool description reinforces the Note 3a/3b framing but does not add significant parameter-level meaning beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads the partner or owner capital schedule (Note 3a and Note 3b), with a specific verb ('Read') and resource. It differentiates from siblings such as save_partner_capital by explicitly framing itself as the read-before-write companion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit: 'Call it before save_partner_capital so you can show your user what is on file and what your change would do.' It explains why reading first is necessary because the save replaces the whole section, and even discusses the Datavrn total-ratio warning behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pending_workList pending work across entitiesA
Read-only
Inspect

Answer "what's left to do?" across every entity you can see — one row per entity, with what is blocking its Schedule III statement: whether the trial balance is in, how many accounts are still ungrouped, the latest generated version, and whether it has been finalised. Pass period_label to pick a period, or omit to default to the period most of your entities have a trial balance for (not necessarily the newest — one entity uploading a future period early will not flip the board). Rows include deep links that open the Datavrn web app (a login is needed there). It also answers a second question nothing else here does: restorable_replacements lists automatic connector syncs that REPLACED an entity's data and can still be undone, soonest-closing first, each with the date its 30-day undo window shuts — after that the replaced data cannot be put back, and no message ever announces that clock running out, so raise these with your user rather than waiting to be asked (use list_replacements and preview_replacement_restore on the ids given; restorable_replacements_omitted says how many more were not listed).

ParametersJSON Schema
NameRequiredDescriptionDefault
period_labelNoReporting period label, e.g. '2026-03'; omit to default to the period most of your entities have a trial balance for.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds substantial behavioral context: the login requirement for deep links, the default-period logic ('not necessarily the newest'), the 30-day undo window and the fact that no message announces its expiry, and the ordering of restorable_replacements. This goes well beyond what annotations provide and helps the agent understand side effects and risks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense, covering two distinct purposes and relevant nuances without fluff. It front-loads the primary purpose and the parameter behavior early, then introduces the secondary replacement feature. The structure is logical, though a bit sprawling; it could be split into bullets for readability, but it remains coherent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description fully explains the return structure (per-entity row details, restorable_replacements with their fields, counted omissions) and provides necessary context about the login requirement and the undo-window risk. It also directs the agent to relevant sibling tools for follow-up actions, making it complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, period_label, with 100% schema description coverage. The description enriches it by explaining the default behavior in more depth than the schema (e.g., that a future-period upload won't flip the board) and how to intentionally pass a label. This adds value over the schema's description, so it earns above the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific question it answers ('what's left to do?') and details exactly what it returns for each entity (trial balance status, ungrouped accounts, latest version, finalised flag) plus the second distinct purpose of listing restorable replacements. This is highly specific and clearly distinguishes it from siblings, as it explicitly claims that the replacement functionality is unique ('nothing else here does').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (to answer pending-work or replacement questions), gives parameter guidance (how to select a period or rely on the default), and suggests follow-up tools (list_replacements, preview_replacement_restore) for acting on the returned IDs. It does not explicitly state when not to use it or contrast with a sibling, but the 'nothing else here does' phrase implies a unique role.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_schedule3_workspaceGet Schedule III workspaceA
Read-only
Inspect

THE state tool: grouping progress, every required capture answer, generated/finalised versions, finalisation blockers, and bounded per-version control summaries. Report generation never marks capture complete. Exception output is rule/severity/count only — no account names or amounts. Call this to know what is left before finalising.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds non-obvious behavioral details: it will never mark capture complete, and exception output is deliberately limited to rule/severity/count without sensitive account data. This sets accurate expectations for side effects and data scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, informative sentences with no filler; each clause earns its place. Slightly heavy on domain jargon like 'bounded per-version control summaries' may impede quick comprehension, but overall it is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given its read-only nature and rich annotations, the description covers tool outputs, non-side-effect behavior, and output limitations. It lacks only a full return structure, which is acceptable without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds no parameter-specific semantics. The baseline of 3 applies because the schema already fully documents each parameter's meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly positions this as the central state tool for Schedule III, enumerating specific components (grouping progress, capture answers, versions, blockers, control summaries). This distinguishes it from sibling tools like get_statement_figures or get_setup_status by framing it as the holistic workspace view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Call this to know what is left before finalising', providing a clear when-to-use scenario. It does not name alternative tools or exclusions, but the context is sufficient for selecting this over other read-only tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_setup_statusShow setup status and the next stepA
Read-only
Inspect

Answer "how do I get started?", "what do I do next?", or help a user who seems lost setting up. Returns where they are in the journey from an empty organization to a finished Schedule III statement, and the ONE next step to take. Call it WITHOUT client_id first (the organization view): it lists the entities this credential can see, or — if there are none — the step to create the first one. Then call it again WITH one entity’s client_id for that entity’s full step-by-step path (upload trial balance → confirm groupings → capture figures → generate → download). Each step has a status (done / next / todo / blocked / web_only) and either the exact tool to call or a web-app link. NARRATE ONE STEP AT A TIME — walk the user through the single next step; do not dump the whole list unprompted. Steps marked web_only are done in the Datavrn web app and need a login — never claim you can do them yourself. This tool reports STATUS only (counts, names, what is done) — it never returns a figure or balance; read those with get_statement_figures once a statement is generated.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idNoAn entity id (from list_clients or the organization view) for that entity’s full path; omit for the organization-level overview and, on an empty org, the create-entity step.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable context: it only reports status (counts, names, what is done), never returns a figure or balance, and explains that web_only steps require login. It also discloses that on an empty organization it returns the create-first-entity step. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries actionable information: two calling modes, step statuses, narration instruction, web_only caveat, and scope limitation. It is front-loaded with the use case. The length is justified by the tool's complexity; there is zero fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully explains return semantics: org vs entity views, step statuses, tool links, and the narration rule. It even cross-references get_statement_figures to clarify what the tool does NOT return. Given the tool's dual-mode nature, this is complete and leaves no critical ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers the single parameter well (omit for org view, include for entity path). The description reinforces this and adds the step sequence and the 'create entity' fallback, deepening the agent's understanding of how the parameter affects output. Since schema coverage is 100%, this extra context raises it above a baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens by naming the exact user questions ('how do I get started?', 'what do I do next?') and states it 'Returns where they are in the journey... and the ONE next step to take.' This clearly identifies the tool's purpose (reporting setup status) and distinguishes it from sibling tools like get_pending_work or generate_schedule_iii.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance (user lost during setup) and a precise two-call protocol: first without client_id, then with one entity's client_id. It also instructs the agent to narrate one step at a time, warns against claiming to perform web_only steps, and points to get_statement_figures for figures. This is exemplary usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_spreading_reconciliationRead spreading reconciliationA
Read-only
Inspect

MANAGEMENT data class. Read the persisted books-to-MIS spreading reconciliation, not a new allocation run. Accounts view gives account-grain books plus adjustments equals MIS; adjustments view requires one account and pages its rule adjustments. Money is decimal-string rupees. Whole-run summary figures never change with account filters. The signed page_token is source-pinned, so restart at page 1 if source_changed. Raw warning context is withheld; known_stale and not_assessed are disclosure states, not freshness claims.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoaccounts (default) or adjustments; adjustments requires account_id.
periodYesManagement month in YYYY-MM.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per source-pinned page: accounts default 25/max 100; adjustments default 50/max 200.
account_idNoRequired for adjustments view; an account in this entity and run.
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
account_name_patternsNoUp to 10 case-insensitive account-name substrings; any match is retained.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds extensive behavioral detail: view-specific output semantics, decimal-string rupees, invariant summary figures, source-pinned page_token behavior, and the meaning of warning states. This significantly aids the agent in interpreting results and pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet information-dense, with every sentence contributing unique value. It is front-loaded with the core purpose and uses a logical flow (purpose, views, data format, pagination, warnings). No redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the absence of an output schema, the description covers critical contextual aspects: view contents, pagination semantics, data type, summary behavior, and disclosure states. This is a complex tool with 7 parameters, and the description provides enough context for the agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description mostly reinforces schema information (e.g., adjust view requires account_id, page_token source pinning) and adds output-related context rather than new parameter depth. It does not materially elevate parameter understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads the persisted books-to-MIS spreading reconciliation, using a specific verb and resource. It explicitly distinguishes from a new allocation run and explains the two view modes, making its purpose unambiguous relative to sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: it is for reading the persisted reconciliation, not a new allocation run, and explains when to use accounts vs adjustments view. It lacks explicit alternatives by name, but the exclusion and view guidance are sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statement_figuresRead statement figuresA
Read-only
Inspect

Read a generated Schedule III statement's figures: the balance-sheet and profit-and-loss faces, current-year and previous-year balance-sheet tie verdicts separately (a null verdict means UNKNOWN, never a pass: either no comparative was captured, or the version predates per-column balance recording), the unclassified count, and the notes listed by number. Also returns bounded exception counts by rule/severity and the frozen control changes versus the immediately previous recorded version; it never recomputes either from live books. Figures come from a generated version (the latest unless you pass a specific version) and match the workbook exactly. If the version was generated before figure reads existed it returns available:false with reason "figures_not_available" and only the legacy flat tie verdict; tell the user to generate the statement again, read the latest version, then retry. For a note's line-by-line breakdown, use its note_index entry with get_statement_notes. Amounts are decimal strings in rupees. ONE EXCEPTION: the profit-and-loss face ends with the statutory earnings-per-share rows, marked kind:"eps". Their figures are ₹ PER SHARE, not rupees of profit — report them as EPS and never add them into a face total. They are null until the weighted average share counts are saved in statement settings. Figures are Datavrn's deterministic engine output; interpretation is your assistant's.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoA specific statement version; omit for the latest.
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses extensive behaviors: null verdicts mean UNKNOWN not pass, old versions return available:false with a specific reason, it never recomputes from live books, figures match the workbook, amounts are decimal strings, and the EPS rows are per share and excluded from totals. This far exceeds what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely packed with necessary information, organized into clear sentences and paragraphs. It front-loads the primary purpose and then covers edge cases and exceptions. While longer than average, every sentence earns its place—for instance, the EPS rule and null-verdict semantics are critical for correct usage. Not excessively verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's 4 parameters, no output schema, and no enums, the description fully equips an agent: it details all return categories, handles versioning edge cases, explains the EPS exception, and provides recovery instructions for unavailable figures. The inclusion of 'frozen control changes' and 'never recomputes' clarifies non-obvious behavior. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% as all 4 parameters have descriptions. The tool description adds no additional parameter-level detail beyond what the schema already states (e.g., version meaning latest unless specified is already in the schema). Baseline 3 applies because the schema does the heavy lifting; no compensation needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description precisely states the tool reads a generated Schedule III statement's figures, enumerating the exact elements (balance-sheet and P&L faces, tie verdicts, unclassified count, notes, exception counts, control changes). It distinguishes itself from sibling get_statement_notes by explicitly routing line-by-line breakdowns to that tool. The verb 'read' and resource 'statement figures' are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly states when to use this tool versus the alternative: 'For a note's line-by-line breakdown, use its note_index entry with get_statement_notes.' It also provides recovery guidance for old versions (regenerate, read latest, retry) and clarifies that interpretation is the assistant's responsibility. No exclusions are missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statement_notesRead statement notesA
Read-only
Inspect

Read the line-by-line breakdown of a generated statement's notes — every line's current and prior-year amount, and the note total. Pass note_numbers (from get_statement_figures' note_index) to fetch specific notes, or omit for all. Use this to answer "what's in Other Expenses?" or "what makes up trade receivables?". Each line has a kind: 'component' (an additive line), 'subtotal' (a presentational group subtotal — do NOT add it into the total, or you double-count), or 'header'. Fixed-asset / intangible notes carry a block per class with gross_block, accumulated depreciation, and net (the additions/deletions movement schedule itself lives in the workbook). If the full set is too large it returns too_large:true with a note_index — fetch note_numbers in small batches. A single very large note (e.g. a PPE schedule or an ageing note) is returned in explicitly-flagged line pages: each page carries the authoritative note total, lines_page, lines_total, and has_more_lines — keep fetching lines_page until has_more_lines is false; never treat one page's lines as the whole note. Amounts are decimal strings in rupees. Figures are Datavrn's deterministic engine output; interpretation is your assistant's.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoA specific statement version; omit for the latest.
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
lines_pageNoFor a single very large note returned in line pages: the 1-based line page to fetch (fetch exactly one note; keep going until has_more_lines is false).
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
note_numbersNoSpecific note numbers to fetch (from note_index); omit for all notes.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly and non-destructive annotations, the description discloses line kinds ('component', 'subtotal', 'header'), warns against double-counting subtotals, explains fixed-asset block structure, details pagination (lines_page, has_more_lines), the too_large flag, decimal strings, and deterministic engine output. This is extensive behavioral context far exceeding annotation value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense yet front-loaded with the core purpose, then systematically covers line kinds, blocks, pagination, and data format. Every sentence delivers distinct, actionable information without redundancy, making it easy for an agent to parse and apply.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by explaining response features (line kinds, totals, pagination fields, too_large flag) and edge cases like double-counting subtotals and single large notes. It fully equips an agent to safely interpret and consume the tool's output across normal and paginated scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds operational meaning: note_numbers can be omitted for all notes, lines_page should be fetched until has_more_lines is false, and how note_index from get_statement_figures feeds into note_numbers. This deeply enriches the schema's basic parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Read the line-by-line breakdown of a generated statement's notes," and clearly states what is included (current/prior-year amounts, note total). It also distinguishes from siblings by referencing get_statement_figures' note_index and the note-level scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit use cases are given ("Use this to answer 'what's in Other Expenses?'...") and it directs users to get_statement_figures for note_numbers. However, it does not explicitly state when not to use this tool or name alternative tools for higher-level summaries, missing the 'when-not/alternatives' part of the 5-level criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tb_rowsRead trial-balance source dataA
Read-only
Inspect

Read the SOURCE DATA behind a statement: the trial-balance rows (account name, debit, credit) as landed for a period, BEFORE grouping — the pre-statement numbers, not statement figures. PREFER FILTERS over fetching everything: name_patterns (e.g. ['cash','bank','od']), side ('debit'/'credit' by net balance), and min_abs_balance return a small exact subset with its own debit/credit totals — e.g. wrong-side cash accounts = name_patterns ['cash','bank'] + side 'credit'. Paginated (page 1-based; page_size default 50, max 500). These are the CURRENT live rows: statement figures are frozen at a generated version, so if the trial balance was re-uploaded after a version was generated, these rows may not tie to that version (the response note says so). Amounts are decimal strings in rupees. Figures are Datavrn's deterministic engine output; interpretation is your assistant's.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo1-based page number (default 1).
sideNoKeep only accounts whose net balance falls on this side (debit = debits exceed credits). Accounts netting to zero match neither.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per page (default 50, max 500 — prefer filters over big pages).
period_idYesThe reporting period id — from list_periods.
name_patternsNoUp to 10 case-insensitive substrings; an account matches if its name contains ANY of them (e.g. ['gst','tds']).
min_abs_balanceNoKeep only accounts whose balance (the larger of its debit/credit) is at least this many rupees — a decimal string like '100000'.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark read-only and non-destructive, but the description adds valuable behavior: rows may not tie to a generated version after re-upload, amounts are decimal strings in rupees, output is deterministic, and pagination behavior. These go beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, filtering guidance, pagination, data freshness caveat, amount format, and interpretation responsibility. It is well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description covers the key output fields (account name, debit, credit, totals, note), filters, pagination, and data semantics. It is complete for an agent to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage, so baseline is 3. The description enhances this by providing examples for name_patterns, side, and min_abs_balance, plus clarifying pagination defaults, adding meaningful context beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource+scope: 'Read the SOURCE DATA behind a statement: the trial-balance rows (account name, debit, credit) as landed for a period, BEFORE grouping'. It clearly distinguishes from statement figures and sibling tools like get_statement_figures.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to 'PREFER FILTERS over fetching everything', gives a concrete example (wrong-side cash accounts), and notes the 'CURRENT live rows' vs frozen statement figures, indicating when not to use this tool. It also contrasts with statement figures.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_uploadGet upload statusA
Read-only
Inspect

Read an upload session: status, detected header row and columns, the confirmed mapping (if any), and the stored validation outcome. Use to check what a staged upload still needs.

ParametersJSON Schema
NameRequiredDescriptionDefault
upload_idYesThe upload session id returned by upload_trial_balance.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description aligns by saying 'Read'. It adds value by listing the specific data returned, which goes beyond the annotation. It does not mention side effects, but for a read-only tool this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the action and resource, and every sentence contributes. There is no redundant information or padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one parameter and no output schema, the description fully covers what the tool returns and when to use it. No additional behavioral context seems necessary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the upload_id parameter already has a clear description ('returned by upload_trial_balance'). The tool description adds no additional parameter semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads an upload session and enumerates exactly what it returns (status, header row/columns, confirmed mapping, validation outcome). The verb 'Read' plus resource 'upload session' is specific and distinct from siblings like get_upload_link_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Use to check what a staged upload still needs' provides clear usage context. However, it does not explicitly mention alternatives or when not to use this tool, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_variance_reportRead variance reportA
Read-only
Inspect

MANAGEMENT data class. Read the existing budget-or-prior variance report for a month; it never recalculates it. Amounts are decimal strings; a null actual or variance means unavailable, never zero. Rows view has whole-report unfiltered rollups and no cross-side grand total; filters affect only rows and filtered grain counts. Explanations view withholds internal notes and may redact structured PII. Filter lines or centres before paging; the signed page_token is source-pinned, so restart at page 1 if source_changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoReport rows (default) or safe explanations.
basisNoComparator basis (default budget).
periodYesManagement month in YYYY-MM.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per source-pinned page (default 50, max 100).
line_codesNoUp to 25 P&L line codes.
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
cost_centre_idsNoUp to 25 cost-centre ids.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true and destructiveHint=false. The description significantly expands on this with deep behavioral details: null semantics ('null actual or variance means unavailable, never zero'), rollup behavior ('whole-report unfiltered rollups and no cross-side grand total'), privacy filtering ('withholds internal notes and may redact structured PII'), and pagination semantics ('source-pinned, so restart at page 1 if source_changed'). This goes far beyond the annotations and is highly informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a dense, well-structured paragraph that front-loads the core purpose ('MANAGEMENT data class. Read the existing budget-or-prior variance report') and then packs high-value edge-case details into a compact series of clauses. Every sentence contributes distinct information (nulls, rollups, explanations view, paging) without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with 8 parameters, two views, pagination, and subtle data semantics, and there is no output schema. The description covers critical operational details: null handling, rollup behavior, explanation-view redaction, and paging restart logic. It is thorough enough that an agent could invoke this tool correctly with minimal additional inference, making it effectively complete for its complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter has a baseline description. However, the tool description adds meaningful semantics about how parameters interact: filters affect rows and filtered grain counts, and page_token is source-pinned. This goes beyond individual parameter docs, enriching understanding of view, line_codes, cost_centre_ids, and page_token. Given the high schema coverage, a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Read the existing budget-or-prior variance report for a month; it never recalculates it.' This clearly identifies the action (read), the resource (existing variance report), and the scope (budget/prior, monthly). It distinguishes this from sibling read tools like get_budget or get_statement_figures by specifying the variance-report resource and its read-only, non-recalculating nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes usage context by stating it reads an existing report and never recalculates, implying it should be used when a pre-computed variance report is needed. It also provides practical guidance: 'Filter lines or centres before paging' and the page_token restart condition. However, it does not explicitly name alternative tools or state when not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workbook_downloadDownload workbookA
Read-only
Inspect

Mint a short-lived signed URL for a frozen workbook version (the Excel file). Give the URL to your user to open in a browser — it needs no login and expires in about 10 minutes. The bytes are immutable and integrity-hashed.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
snapshot_idYesThe snapshot id from list_snapshots.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond annotations: short-lived signed URL, ~10-minute expiry, no login required, immutable bytes with integrity hashing. These details help the AI understand the tool's behavior without needing to invoke it. No contradiction with the readOnlyHint=true annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose, no redundant phrasing. The second sentence adds essential usage and behavioral details without waste. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains what the tool returns (a URL) and gives key attributes (short-lived, no login, immutable). It does not specify the exact response format (e.g., JSON field name), but for a low-complexity tool with no output schema, this covers the essential context. Slight gap in not mentioning any error conditions or prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with descriptions for client_id and snapshot_id, referencing list_clients and list_snapshots. The description adds minimal meaning for snapshot_id ('frozen workbook version') but does not elaborate on parameter syntax or format. Baseline of 3 is appropriate given high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Mint a short-lived signed URL for a frozen workbook version (the Excel file).' The verb 'mint' is specific and the resource (workbook download) is unambiguous. It distinguishes itself from sibling download/upload tools by focusing on the frozen workbook version download.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: 'Give the URL to your user to open in a browser — it needs no login and expires in about 10 minutes.' This implies the tool is for user-facing download of a frozen workbook, but it does not explicitly mention alternatives or when not to use. The context is clear but lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ingest_uploadIngest upload into the booksA
Destructive
Inspect

Commit a validated upload into the entity’s books. This is a TWO-CALL approval: if the upload has any warnings, or would permanently delete existing trial-balance rows for a period it covers, the first call writes NOTHING and refuses with every warning, the exact record counts, and a short-lived removal_token. Show your user every warning and both counts, get their explicit go-ahead, then resend the SAME call adding removal_token and removal_count exactly as returned. acknowledge_warnings is IGNORED on this connection — the token is the only acknowledgment, so sending it changes nothing. A clean, additive ingest needs no token and succeeds on the first call. Returns the ingestion outcome including any notices.

ParametersJSON Schema
NameRequiredDescriptionDefault
upload_idYesThe upload session id returned by upload_trial_balance.
removal_countNoThe exact removal count returned by the removal preview.
removal_tokenNoOnly include the short-lived token returned by the removal preview for this exact upload.
confirm_mergesNo
acknowledge_warningsNo
acknowledged_state_digestNoBrowser/REST only, and IGNORED on the agent connection exactly as acknowledge_warnings is: the destroy_state_digest returned by the confirm-mapping step, resent unchanged so the server can prove the acknowledgment was given against the data that is still there. On this connection the removal_token already pins the rows at risk, so sending this changes nothing.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, and the description goes further by describing the exact destructive scenario (permanently deleting trial-balance rows), the refusal-on-first-call behavior, and the short-lived token as the only acknowledgment. It also discloses that two parameters are ignored on this connection, which is critical for correct invocation. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries essential operational detail: purpose, conditional flow, token handling, ignored params, and return value. It is front-loaded with the core purpose. Slight verbosity, but justified given the two-call complexity; no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema and 6 parameters, the description covers the approval flow, the token lifecycle, the ignored fields, and the return value (ingestion outcome with notices). It assumes a prior 'validated upload' step, which is reasonable given sibling upload_trial_balance. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so the description is expected to fill gaps. It explicitly explains removal_token and removal_count (exact counts, token short-lived, resend unchanged) and states acknowledged_state_digest is ignored. However, confirm_merges is not mentioned at all, and acknowledge_warnings only gets a dismissal. Still, the description adds substantial meaning for the most security-sensitive parameters, going beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Commit a validated upload into the entity’s books.' It clearly differentiates this from upload_trial_balance (which creates the upload) and other confirm_* siblings by framing it as the final commit step. The two-call approval mechanism is prominent, making the tool's unique role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when a token is required (warnings or destructive over-write) vs when a simple additive ingest suffices. It instructs the agent to show warnings and counts, get explicit go-ahead, and resend with the token. It also clarifies that acknowledge_warnings is ignored, preventing misuse. This is as explicit as usage guidance gets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_account_mappingsReview account mappingsA
Read-only
Inspect

Review account-to-cost-centre mapping status and deterministic suggestions for an entity. This is status-only: it returns account names, types, target names, confidence, reasons, and balance-bearing booleans, but never debit, credit, balance, or any rupee amount. Always present the rows grouped by confidence tier and target, state exact counts, flag every medium/low-confidence row, and show the two distinct completion counts: unmapped_total and unmapped_with_balance. Do not call either count pending.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo1-based page number (default 1).
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per page (default 100, max 200).
name_patternsNoReturn accounts whose code or name contains at least one of these case-insensitive patterns.
unmapped_onlyNoReturn only accounts that still need a centre mapping.
balance_bearing_onlyNoReturn only accounts that carry a balance, without returning the balance itself.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds substantial context: it is 'status-only', never returns rupee amounts, returns deterministic suggestions, and imposes presentation rules (grouping by confidence tier, exact counts, flagging medium/low rows, and avoiding the word 'pending'). This goes well beyond the annotations and clarifies the tool's behavioral safety and output expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary purpose. The second sentence packs many required presentation rules, which is dense but each clause earns its place. It could be broken into bullet points for readability, but it is efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description takes on the burden of explaining return values: account names, types, target names, confidence, reasons, balance-bearing booleans, and the two completion counts. It also covers usage context. Gaps remain (e.g., ordering of results, definition of confidence tiers, pagination behavior) but the core information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with every parameter documented (e.g., client_id, page, page_size, name_patterns, unmapped_only, balance_bearing_only). The description does not add additional parameter-level meaning beyond what the schema already provides, and the baseline is 3 when the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb+resource: 'Review account-to-cost-centre mapping status and deterministic suggestions'. It clearly distinguishes from sibling tools like confirm_centre_mappings or list_grouping_suggestions by emphasizing 'status-only' and listing exactly what is returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly sets the context: use this to review mapping status before confirming, and it explicitly states what data is NOT returned (no amounts). However, it does not name alternative tools or explicitly say when not to use it, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_allocation_runsList allocation runsA
Read-only
Inspect

MANAGEMENT data class. Discover persisted allocation runs and their conservation heads; this does not generate or recompute allocation. Money is decimal-string rupees. Results are ordered period, version, then run id, and the signed page_token is pinned to the complete filtered source: if it reports source_changed, restart at page 1. current_only means latest generated version, not source freshness; known_stale and not_assessed are both warnings, never a claim that the source is fresh. Raw stale reasons and warning context are not returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNoOptional management month in YYYY-MM.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per source-pinned page (default 20, max 50).
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
current_onlyNoReturn only the latest generated version per period (default true); this is not a freshness claim.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, but the description adds substantial behavior beyond that: ordering by period/version/run id, page_token pinning and source_changed handling, current_only semantics (not freshness), stale warnings ('never a claim that the source is fresh'), and absence of raw reasons. This is rich context and fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each adding new information: classification, purpose, money format, ordering, pagination semantics, stale warnings, and exclusion of raw reasons. Front-loaded with purpose and no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully covers the operational semantics: ordering, pagination with source pinning, current_only meaning, stale warning interpretation, and data format (decimal-string rupees). It is complete for a read-only list tool and handles edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 5 parameters are fully described in the schema (100% coverage), so the baseline is 3. The description elevates this by adding deeper meaning for page_token ('pinned to the complete filtered source... restart if source_changed') and current_only ('latest generated version, not source freshness'), which goes beyond the schema's descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Discover persisted allocation runs') and immediately clarifies scope ('conservation heads') while explicitly stating what it does not do ('does not generate or recompute allocation'). This strongly distinguishes it from generation tools like generate_schedule_iii.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly indicates when to use (discover persisted runs) and includes an explicit exclusion ('does not generate or recompute allocation'), but it does not name alternative tools or provide broader when-to-use guidance beyond this scope. The context is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_budgetsList budgetsA
Read-only
Inspect

MANAGEMENT data class. Discover budget ids and versions without identity fields. Locked FX rate and all money-valued fields are decimal strings. Filter status or fiscal-year start before paging. The signed page_token is pinned to the complete filtered source; restart at page 1 if source_changed. This lists budget headers only, not cells, approvals identities, or a recalculated budget.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoOptional budget status.
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per source-pinned page (default 20, max 50).
page_tokenNoSigned continuation from the prior page; restart without it if source_changed.
fiscal_year_startNoOptional fiscal-year first-of-month date, e.g. 2026-04-01.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=true, destructiveHint=false), the description adds meaningful behavioral context: money-valued fields are decimal strings, page_token is pinned to the complete filtered source, and restart at page 1 if source_changed. These details are not present in annotations and enrich the agent's understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each carrying distinct information: data class, purpose, data format, pagination, and scope exclusions. The 'MANAGEMENT data class' phrase is somewhat cryptic, but the description is not bloated and front-loads the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with 5 parameters and no output schema, the description covers purpose, pagination behavior, data format, and exclusions, providing a solid mental model. It does not explicitly describe the full response structure but implies the return contains budget ids and versions, which is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description covers all parameters (100%), but the description adds usage semantics: 'Filter status or fiscal-year start before paging' and explains the page_token's pinning behavior. This goes beyond the schema's individual parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description explicitly states 'Discover budget ids and versions' and 'lists budget headers only', clearly identifying the resource (budgets) and action (list/discover). It distinguishes itself from sibling tools like get_budget by scoping to headers only and excluding cells, approvals identities, and recalculated budgets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context: filter status or fiscal-year start before paging, and explains page_token pinning with restart guidance if source_changed. It notes what the tool does NOT list, implying alternatives, but does not explicitly name sibling tools like get_budget.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_chart_rebaselinesUploads not yet counted as identity evidenceAInspect

List the uploads Datavrn is NOT counting as evidence of which entity a file belongs to. This happens when an entity’s chart of accounts grew or changed faster than Datavrn can vouch for from what it already holds — typically an acquisition, a migration, or a year-end restructure. NOTHING WAS REJECTED, BLOCKED OR CHANGED: the figures in those uploads are landed and live. What has not advanced is the evidence Datavrn compares FUTURE uploads against, so wrong-entity detection for this entity is working from a smaller picture than the entity’s real chart. Each row reports how many distinctive ledger names Datavrn already held (trusted_considered), how many the upload carries (incoming_considered), how many are on both sides (matched), and the two coverage ratios. Show your user those numbers in their own terms. shortfall names WHICH direction fell short, and it is the part to say out loud, because the three cases are different situations to an accountant: "incoming" means most of the file is ledger names Datavrn does not already treat as evidence — the chart in the file is much larger than what Datavrn holds; "trusted" means much of what Datavrn holds is missing from the file; "both" means the two charts barely overlap either way. It is null when there was nothing to compare against at all, which the row reports as reason "empty_trust_recovery". This call only lists. To confirm one of these uploads, call preview_chart_rebaseline for that upload — it re-reads the figures and issues the approval confirm_complete_chart needs. An empty list means nothing is waiting: for almost every entity that is the normal state.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description claims 'This call only lists' and 'NOTHING WAS REJECTED, BLOCKED OR CHANGED', implying a read-only operation. However, the annotation readOnlyHint is false, which suggests the operation may not be read-only. This directly contradicts the description, so per the rules the score is 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then explains the semantic distinctions of shortfall and the three cases. While lengthy, most sentences serve a clear purpose for an accountant-facing tool, though it could be trimmed without losing essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description thoroughly explains the returned row fields (trusted_considered, incoming_considered, matched, coverage ratios, shortfall, reason), the null case, and guidance on how to proceed. It does not cover pagination or error conditions, but for a list-only tool this is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and both parameters are well described in the input schema. The tool description adds no additional parameter semantics beyond what the schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('List the uploads Datavrn is NOT counting as evidence'), names the exact resource ('uploads not yet counted as identity evidence'), and differentiates from the sibling preview_chart_rebaseline and confirm_complete_chart. It is unmistakable what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it ('This call only lists') and when to use the alternative ('To confirm one of these uploads, call preview_chart_rebaseline'), and also explains the normal state of an empty list. This gives clear routing guidance with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_clientsList entitiesA
Read-only
Inspect

List the entities (companies) this credential can work with. Call this first to resolve the client_id every other tool needs. Returns each entity id and name.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safe read nature is covered. The description adds behavioral context by stating the returned data (entity id and name) and the credential-scoping behavior, which is useful beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the primary action and immediately followed by the key usage instruction. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only list tool, the description fully covers purpose, usage, return format, and relationship to other tools. No output schema exists, but the description explicitly states what is returned (entity id and name), making it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is 100% by default. The description adds meaningful context about the resource being listed and the return value, fulfilling the baseline for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('List') and resource ('entities (companies) this credential can work with'), and explicitly distinguishes the tool's purpose by explaining it resolves the client_id needed by other tools. This differentiates it from sibling tools like create_client or list_budgets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs 'Call this first to resolve the client_id every other tool needs,' giving concrete when-to-use guidance and positioning it as a prerequisite step. This is strong usage direction beyond what annotations provide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_cost_centresList cost centresA
Read-only
Inspect

List the cost centres for an entity. Use this before proposing account mappings so you can group the proposal by target name and distinguish operating from support centres. This is status-only: it returns names and kinds, never rupee amounts. Tell the user what the existing structure means before suggesting a change.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo1-based page number (default 1).
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per page (default 100, max 200).
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds meaningful context beyond annotations: it is 'status-only', returns only names/kinds (never rupee amounts), and advises informing the user about the existing structure before suggesting changes. This enriches the behavioral profile without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: primary action, workflow context, and behavioral caveat. Front-loaded with the verb+resource, no filler, and well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return content (names and kinds, no amounts) and covers purpose, usage, and user interaction guidance. Combined with detailed schema and annotations, the tool is fully contextualized for an agent to invoke and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter (client_id, page, page_size) fully described in the schema. The description adds no parameter-specific detail beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('List the cost centres for an entity') and distinguishes from siblings like list_profit_centres by clarifying it returns names and kinds only, never rupee amounts. The purpose is unambiguous and clearly scoped.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('before proposing account mappings') and what to do with the results (group by target name, distinguish operating vs support, explain structure before suggesting changes). It does not name alternatives or provide explicit 'when not to use' exclusions, but the context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_grouping_suggestionsList grouping suggestionsA
Read-only
Inspect

List ungrouped accounts with DETERMINISTIC grouping suggestions (curated rules + name/group-path matching — no AI is involved; Datavrn never applies a suggestion itself). Paginated. Each row carries a reason and a confidence tier: present them to your user GROUPED BY CONFIDENCE, and call out low-confidence and balance-bearing rows for individual attention — a single blanket approval is not a review of the low-confidence tail. Confirm only what your user approves via confirm_groupings. Returns a summary (counts by confidence tier) plus one page of suggestion rows — fetch tier by tier with the confidence filter instead of everything at once; pass include='confirmed' to see already-confirmed groupings.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size (default 50, max 200).
offsetNoRows to skip (default 0).
includeNoWhich rows to page: 'suggestions' (default), 'confirmed' (already-grouped accounts), or 'both'.
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
confidenceNoKeep only suggestion rows in this confidence tier ('none' = accounts with no deterministic suggestion). Filters rows only — the summary counts stay over the whole population.
has_balanceNoKeep only suggestion rows whose account carries a live balance (true) or not (false).
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and non-destructive behavior. The description adds substantial behavioral context: deterministic rules (no AI), Datavrn never auto-applies, pagination behavior, summary counts remain whole-population when filtering, and return format. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loaded, but somewhat long (~170 words). Every sentence adds value, though some phrases like 'a single blanket approval is not a review...' could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (8 params, pagination, confidence tiers, confirmation workflow) and no output schema, the description is remarkably complete. It explains return format, pagination, filtering strategy, and the confirmation flow, offering sufficient context for correct tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already covers 100% of parameters, so baseline is 3. The description adds usage semantics for confidence (fetch tier by tier), include ('confirmed' shows already-grouped), and has_balance (call out balance-bearing rows), which goes beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists ungrouped accounts with deterministic grouping suggestions, using a specific verb and resource. It distinguishes itself from the sibling confirm_groupings tool by explicitly noting that suggestions are never auto-applied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage guidance: present suggestions grouped by confidence, call out low-confidence rows, confirm only via confirm_groupings, and fetch tier by tier instead of all at once. It also tells when to use include='confirmed'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_periodsList reporting periodsA
Read-only
Inspect

List the reporting periods a Schedule III statement can be prepared for (periods with a live Trial Balance). Returns period ids for get_schedule3_workspace, save_py_values, and generate_schedule_iii.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and non-destructive behavior. The description adds meaningful context by specifying that only periods with a live Trial Balance are listed and that the output consists of period IDs used by specific downstream tools. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences: the first leads with the primary action and scope, the second clarifies the output and consumers. Every word adds value, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list operation with one parameter and no output schema, the description fully covers purpose, the condition (live Trial Balance), and return value (period IDs). Naming the downstream tools provides integration context, making the description complete on its own.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents the single parameter client_id, including its type and derivation ('from list_clients'), so the description need not add further detail. The description does not mention parameters at all, which is acceptable given the schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a clear resource ('reporting periods for a Schedule III statement'), and a distinguishing condition ('periods with a live Trial Balance'). It also names downstream consumers (get_schedule3_workspace, save_py_values, generate_schedule_iii), which differentiates it from other list_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: before preparing a Schedule III statement, to obtain valid period IDs. It explicitly names the tools that consume these IDs, giving clear workflow context. However, it does not explicitly state when not to use it or mention alternatives, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_profit_centresList profit centresA
Read-only
Inspect

List the profit centres for an entity. Use this to explain available targets before a user confirms any explicit mapping. This is status-only: it returns names and hierarchy, never rupee amounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo1-based page number (default 1).
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per page (default 100, max 200).
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds meaningful context: 'It returns names and hierarchy, never rupee amounts,' which clarifies the return scope without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, each earning its place: purpose, usage context, and behavioral caveat. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple read-only list tool, the description plus schema fully cover what an agent needs: return type (names/hierarchy), non-financial nature, and pagination parameters. No output schema exists, but the description adequately conveys return content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description doesn't add parameter-specific details beyond the schema, but it doesn't need to; the schema already documents client_id, page, and page_size thoroughly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'List the profit centres for an entity', a specific verb+resource with scope. It further distinguishes itself from siblings by noting it's used 'before a user confirms any explicit mapping', clarifying its role versus confirmation/mutation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear when-to-use context: 'Use this to explain available targets before a user confirms any explicit mapping.' However, it doesn't explicitly name alternative tools or state when not to use it, though the phrase implies it precedes confirm_centre_mappings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_replacementsList what automatic syncs replacedA
Read-only
Inspect

List what a connected data source’s AUTOMATIC syncs have REPLACED for one entity — each one showing what was overwritten, how many records, and whether it can still be undone. Call this when a user says figures for a past period look wrong or changed on their own, and after any surprise in a period a connector covers. Datavrn keeps the replaced data for 30 days from the moment it was replaced: within that window state is "restorable" and preview_replacement_restore/restore_replacement can put it back; after it, state is "lapsed" — the record of what happened is kept and is still listed here, but it can no longer be undone from this connection, so tell the user to contact Datavrn support if they need that data recovered. "restored" means it has already been put back. Only syncs that ran UNATTENDED are listed: a replacement someone on the team previewed and confirmed themselves is not offered for undo, by design. window_closes_at is the date the undo window shuts. If connection_attributed is false, Datavrn can no longer prove which connection made that replacement — say so rather than naming this one. To find these without a connection id, call get_pending_work: it lists every still-undoable replacement across the entities you can see, with the entity and connection to use here.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
connection_idYesThe connected data source’s id — from get_pending_work’s restorable_replacements rows, or from the entity’s data-sources page in the web app.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds substantial behavioral context: 30-day retention window, state transitions (restorable/lapsed/restored), the unattended-only filter, the connection attribution caveat, and the semantics of window_closes_at. This goes far beyond the annotations and is essential for correct use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries information. It front-loads the core purpose, then states lifecycle, exceptions, and alternatives. For the complexity (state machines, undo windows, attribution caveats), the length is justified. Slight structure improvements could split into paragraphs, but it's coherent and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description explains the return contents (what was overwritten, record counts, undoability) and all state values. It covers the full workflow: when to call, what to expect, and what to do afterwards (preview/restore or support). Nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for both parameters. The description adds provenance guidance: client_id comes from list_clients, connection_id from get_pending_work or the web app. This is valuable, though not exhaustive—it doesn't explain format or edge cases. Still, it elevates beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('list') and resource ('what a connected data source’s AUTOMATIC syncs have REPLACED for one entity') and enumerates what each entry shows. It clearly differentiates from siblings like get_pending_work by naming it as the alternative for cross-entity discovery, and preview_replacement_restore/restore_replacement as follow-up actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call: 'when a user says figures for a past period look wrong or changed on their own, and after any surprise in a period a connector covers.' It also gives a direct when-not: manual confirmations are excluded by design, and provides an alternative (get_pending_work) for the case without a connection id. No ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_reporting_linesReview reporting-line mappingsA
Read-only
Inspect

Review reporting-line classification status and deterministic suggestions for an entity and reporting period. This is status-only: it returns names, line labels, confidence, reasons, and balance-bearing booleans, but never debit, credit, balance, or any rupee amount. Present suggestions grouped by confidence tier and target, with exact counts and both unmapped_total and unmapped_with_balance; do not call either count pending.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo1-based page number (default 1).
client_idYesThe entity (client) id — from list_clients.
page_sizeNoRows per page (default 100, max 200).
period_idYesThe reporting period id from list_periods.
name_patternsNoReturn accounts whose code or name contains at least one of these case-insensitive patterns.
unmapped_onlyNoReturn only accounts without a confirmed reporting line.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and destructiveHint annotations, the description discloses that the tool returns names, line labels, confidence, reasons, and balance-bearing booleans but never debit, credit, balance, or any rupee amount. It also clarifies the output is deterministic and instructs not to label the two counts as pending, giving meaningful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the purpose and safety scope stated in the first two sentences. The final sentence adds useful presentation guidance but is somewhat opaque, especially the do not call either count pending clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates well by enumerating the key returned fields and count semantics. It does not fully specify the response structure or pagination behavior, but the schema already covers paging parameters and the description provides enough for a reasonable agent to invoke and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100 percent, so the parameters are already well documented. The description adds only the general notion of entity and reporting period, which does not materially improve on the schema, justifying the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, Review, and identifies the exact resource: reporting-line classification status and deterministic suggestions for an entity and reporting period. The explicit status-only framing clearly distinguishes it from sibling confirmation and mutation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The read-only, status-only description makes the intended review use case clear and implicitly separates it from confirm_reporting_lines and other mutation tools. It does not explicitly name exclusions or alternatives, so it falls slightly short of a perfect usage-guidance score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_snapshotsList statement versionsA
Read-only
Inspect

List the frozen Schedule III workbook versions for an entity (newest first), including each version’s period, template, and unclassified count at build time.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
template_idNoFilter to one template.
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds useful behavioral context: versions are 'frozen' (immutable), ordered 'newest first', and include 'unclassified count at build time'. This exceeds the annotation baseline without contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence packs all essential information: what is listed, the scope, ordering, and key fields included. No wasted words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully enumerates returned fields (period, template, unclassified count) and ordering. It may omit a version identifier/timestamp, but for a straightforward list tool with two parameters, this is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters have clear descriptions (client_id source from list_clients, template_id as a filter). The tool description aligns with these but does not add additional semantic nuance beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and specifies the exact resource ('frozen Schedule III workbook versions') with scope ('for an entity') and ordering ('newest first'). This distinguishes it clearly from sibling list tools like list_clients or list_budgets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for retrieving a client's frozen version history, but it does not explicitly mention when to use this tool versus alternatives such as get_workbook_download or get_schedule3_workspace. No exclusions or alternative tool references are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_statement_policy_choicesReview policy and affirmation choicesA
Read-only
Inspect

List every Significant Accounting Policy and Other Regulatory Information affirmation for this statement, with the text that will print, whether a template choice is still unresolved, and what was answered LAST YEAR. This is what makes two rules actionable rather than decorative: never resolve a bracketed choice for your user, and always tell them when an answer differs from last year. CHECK THE ROW’S captured FLAG BEFORE YOU CALL ANYTHING A POLICY CHANGE. differs_from_prior is true in two different situations and only one of them is a change: with captured true the wording was set this year and genuinely differs, which IS a change in accounting policy requiring disclosure under AS-5 / Ind AS 8; with captured false nothing has changed — last year was answered, this year has not been, and the text shown is Datavrn’s generic template wording, which is what will PRINT unless last year’s wording is entered again. Warn your user about that second case explicitly: it silently replaces a policy they wrote. The summary gives you both numbers separately — changed_total (real AS-5 changes) and not_carried_forward_total (answered last year, not yet this year); differs_from_prior_total is simply the two added together. An unresolved choice blocks finalisation, so work through them with your user before generating the version you intend to finalise. "Last year" means the SAME MONTH ONE YEAR EARLIER — the same comparative period the statement itself reports — not the period immediately before this one. The summary names it: prior_period_label is the year that was compared against, and prior_period_found tells you whether Datavrn holds that year at all. A blank last-year answer means nothing was recorded for that year — when prior_period_found is false it means Datavrn has no such reporting period, so there is nothing to compare. Either way it does NOT mean last year matched this year; say which of the two it is.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size (default 50, max 200).
offsetNoRows to skip (default 0).
client_idYesThe entity (client) id — from list_clients.
period_idYesThe reporting period id — from list_periods.
template_idYesThe statement template id (e.g. 'schedule3_v1' Division I; see list_snapshots/workspace).
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already declare readOnlyHint=true and destructiveHint=false, the description adds substantial behavioral nuance: the captured flag distinction, the two meanings of differs_from_prior, the definition of 'last year' as the same month one year earlier, and the meaning of prior_period_found. It also warns about a critical trap ('Warn your user about that second case explicitly') that annotations cannot convey. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core purpose in the first sentence. It then delivers critical rules in a structured, information-dense manner. Every sentence serves a purpose, but the overall length is substantial and could be tightened. It earns a 4 rather than a 5 due to verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of explaining return values and important nuances. It covers the listed fields, summary counts (changed_total, not_carried_forward_total, differs_from_prior_total), the definition of 'last year', prior_period_found semantics, and finalisation blocking. This is more than adequate for an agent to understand and act on the tool's results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage with descriptions for all five parameters (client_id, template_id, period_id, limit, offset). The tool description does not add parameter-specific meaning beyond that, but it does reference statement context and prior-period semantics which indirectly relate to period_id. Since schema already does the heavy lifting, a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope: 'List every Significant Accounting Policy and Other Regulatory Information affirmation for this statement, with the text that will print, whether a template choice is still unresolved, and what was answered LAST YEAR.' This clearly distinguishes the tool from siblings like save_accounting_policies or finalise_statement, making the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong contextual guidance on when to use it: 'An unresolved choice blocks finalisation, so work through them with your user before generating the version you intend to finalise.' It also warns about the silent-replacement scenario. It does not explicitly name alternative tools or state when not to use it, but the context is clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_chart_rebaselineShow one upload’s figures and get the approvalAInspect

Show ONE upload’s figures as they stand right now, and get the approval confirm_complete_chart needs. Call this after list_chart_rebaselines, for the one upload your user is considering. It reports the same counts the list does — how many distinctive ledger names Datavrn already treats as evidence (trusted_considered), how many this upload adds (incoming_considered), how many are on both sides (matched), the two coverage ratios, and which direction fell short (shortfall) — re-read at this moment rather than when you listed. Read them to your user in their own terms and ask them plainly whether that upload is the entity’s whole book now. NEVER decide this from the numbers yourself — a large jump, a round number or a long gap is a reason to ASK, never a reason to conclude. Only the person who knows the client’s books can answer it. The approval is single-use, expires in 15 minutes, and is tied to this entity, this upload, the member you name, and the exact figures returned here. If the entity’s uploads change in between — including a later clean file that resolves this on its own — the approval is spent and you start again from this call. Datavrn refuses and changes nothing if this upload is not awaiting a complete-chart confirmation: it was never held out of the evidence, someone already confirmed it, or it is one Datavrn asked about directly — for that last one, answering that question is what settles it.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
source_snapshot_idYesThe upload’s id — the `snapshot_id` field of a list_chart_rebaselines row.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description discloses critical side effects: the approval is single-use, expires in 15 minutes, is tied to entity/upload/member/figures, and is spent if uploads change between the call and the confirmation. It also details refusal conditions that leave state unchanged, which is essential for an agent to predict the tool's impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but every sentence contributes: purpose, usage, return values, side-effect disclosure, and refusal conditions are all necessary for correct invocation. It is front-loaded with the core purpose and then expands logically. Could be tightened slightly, but the length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies the returned data (the five counts and two ratios) and how it relates to the list_chart_rebaselines output ('re-read at this moment'). It also covers the approval's constraints and refusal conditions, so an agent has everything needed to call it correctly and understand its effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents client_id, source_snapshot_id, and on_behalf_of. The description adds only light context (e.g., 'the one upload your user is considering' and the approval's tie to the member named), but does not clarify parameter syntax or edge cases beyond the schema. Baseline of 3 is appropriate since the schema carries the detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (show) and resource (ONE upload’s figures), and explicitly ties it to the approval that confirm_complete_chart needs, distinguishing it from siblings like list_chart_rebaselines and confirm_complete_chart. It names the specific precondition (after list_chart_rebaselines) and the entity of interest, leaving no ambiguity about the tool's role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use ('Call this after list_chart_rebaselines, for the one upload your user is considering') and when-not-to-use (refusal conditions: not awaiting confirmation, already confirmed, or a directly asked upload). It also instructs the agent to never decide based on numbers alone and to ask the user, giving clear behavioural guidance for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_replacement_restorePreview undoing an automatic syncAInspect

Show EXACTLY what undoing one automatic sync would do, counted at this moment, and get the approval restore_replacement needs. destroy_count is how many records undoing it would DESTROY — everything currently held for that period, whoever or whatever put it there, including data a colleague uploaded since. restore_count is how many records would be put back. READ BOTH NUMBERS TO YOUR USER IN THEIR OWN TERMS AND GET THEIR EXPLICIT GO-AHEAD BEFORE CALLING restore_replacement. Undoing a sync is itself a destructive act. If blocked_reason comes back non-null, the period is locked by a finalised statement or a sealed run: read that reason out, do not call restore_replacement, and no approval is issued. The approval is single-use, expires in 15 minutes, and is tied to this entity, this sync, this connection, the member you name AND the exact destroy_count returned here. Send that same number back as expected_destroy_count — never a number you adjusted. If the period’s data changes in between, the approval is spent and you start again from here.

ParametersJSON Schema
NameRequiredDescriptionDefault
event_idYesThe replacement’s id — the `id` field of a list_replacements row.
client_idYesThe entity (client) id — from list_clients.
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
connection_idYesThe connected data source’s id — from list_replacements or get_pending_work.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false) so the description carries the behavioral burden. It fully discloses that the tool is a preview that computes counts 'at this moment,' returns blocked_reason for locked periods, issues a single-use approval expiring in 15 minutes tied to specific parameters, and warns that undoing a sync is itself destructive. This exceeds what annotations convey and adds critical context about side effects and failure modes. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is approximately 250 words—long but not bloated for a safety-critical approval tool. However, it's a single paragraph without visual structure, and some phrases are repetitive (e.g., 'undoing a sync' appears twice). It is front-loaded with the main purpose, but could be tightened with bullet points or shorter sentences for easier scanning. Compared to the calibration examples, this is acceptable for complexity but not as crisp as ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description must explain return values; it does: destroy_count, restore_count, and blocked_reason. It also covers timing (counted at this moment), approval mechanics (single-use, expiry, binding), and failure handling (blocked_reason non-null). All necessary steps for correct invocation and follow-up are present. Given its role in an approval chain, the description is complete and self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover 100% of parameters, so baseline is 3. The tool description adds contextual meaning by linking parameters to the approval context: 'this entity, this sync, this connection, the member you name' maps to client_id, event_id, connection_id, and on_behalf_of. It also explains how to feed the output (destroy_count) into the expected_destroy_count parameter of the next tool. This goes beyond schema descriptions and clarifies the operational flow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Show EXACTLY what undoing one automatic sync would do' and 'get the approval restore_replacement needs.' It uses a specific verb (preview) plus resource (undoing a sync) and explicitly differentiates from the sibling restore_replacement by framing this as the prerequisite step. An agent can immediately understand what this tool does and how it fits the workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: instructs the agent to read both destroy_count and restore_count to the user, get explicit go-ahead before calling restore_replacement, and handle blocked_reason by aborting. It also names the sibling tool restore_replacement as the follow-up and warns when not to proceed. This is direct, actionable guidance with clear conditions and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_replacementUndo an automatic syncA
Destructive
Inspect

Undo one automatic sync: destroy what is currently held for that period and put back the records the sync replaced, recorded as authorised by the member you name. THIS DESTROYS DATA. Everything currently held in that period goes — including anything uploaded or synced since — and is replaced by what was there before. Datavrn keeps a record of what this destroys, but it is NOT offered back on this surface; recovering from an undo is a support operation. Call preview_replacement_restore first, read both of its numbers to your user, get their explicit go-ahead, and send back the confirm_token it gave you with the EXACT destroy_count it returned as expected_destroy_count. Datavrn refuses and changes nothing if: the period’s data moved after you were shown those numbers, the period is locked by a finalised statement or a sealed run, someone already undid this sync, the 30-day window has closed, or the replacement was one a person previewed and confirmed themselves. On success it reports how many records were destroyed, how many were put back, and — under soft_reference_relink — how many links to the restored documents could not be re-established; tell your user that residual number rather than presenting the undo as complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
event_idYesThe replacement’s id — the same one you previewed.
client_idYesThe entity (client) id — from list_clients.
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
confirm_tokenYesThe approval from preview_replacement_restore. Call it first and present both of its numbers to your user.
connection_idYesThe connected data source’s id — the same one you previewed with.
expected_destroy_countYesThe destroy_count preview_replacement_restore returned and you showed your user. Send it back unchanged — a different number is refused, because the approval covers the act your user actually saw.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint=true annotation, the description discloses concrete behavioral implications: it destroys all data held in the period (including uploads/syncs after the replacement), is not recoverable on this surface, and notes that recovering is a support operation. It also clarifies that the tool reports a residual number for failed link re-establishment, adding depth beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is verbose (approximately 150 words) but every sentence adds critical information for a high-stakes destructive operation. It is front-loaded with the destructive warning and includes structured sections for workflow, failure conditions, and post-action reporting. While not terse, the length is justified by the operational risk and complexity, so it earns a 4 rather than a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema and no annotations beyond the destructive flag, the description fully compensates: it explains what the tool does, the prerequisite workflow, all rejection conditions, what success reports (destroy count, restored count, residual link count), and that the residual number must be conveyed to the user. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with each parameter already thoroughly documented (e.g., on_behalf_of's API-key vs OAuth behavior, expected_destroy_count's exactness requirement). The description does not add new parametric information beyond what the schema already provides; it only reinforces the workflow. Baseline 3 is appropriate since the schema carries the semantic burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Undo'), a clear resource ('one automatic sync'), and precisely describes the action: destroy current data for the period and restore the replaced records. It distinguishes itself from sibling preview_replacement_restore by naming it as the prerequisite preview step, leaving no ambiguity about what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs the agent on the exact workflow: call preview_replacement_restore first, read both numbers to the user, obtain explicit go-ahead, and pass back the confirm_token and destroy_count. It also lists multiple refusal conditions (data moved, locked period, already undone, 30-day window closed, human-confirmed replacement) providing clear when-not-to-use guidance and reinforcing the safe invocation pattern.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_capture_declarationWithdraw a recorded capture answerA
Destructive
Inspect

Withdraw a recorded capture answer or review confirmation. What happens next depends on what answers the section: withdrawing a “nothing to record”/“does not apply” answer or a review confirmation makes Statement readiness show the section as UNANSWERED again; withdrawing a leftover earlier note from a section that is answered by its saved rows removes the record and the section STAYS answered. A version you have already generated is NOT affected — if you do not want that version finalised, answer the section again and generate a fresh version. Nothing is deleted: the withdrawn answer stays on the record with who recorded it and who withdrew it, and recording a new answer afterwards creates a new entry rather than overwriting the old one. One thing on this connection is affected immediately: if you already called get_finalise_readiness and hold an approval for that version, withdrawing an answer invalidates it, and the next finalise_statement will refuse and ask you to review the current state again.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
capture_kindYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations (readOnlyHint: false, destructiveHint: true). It reveals that "Nothing is deleted," clarifies the effect on existing versions, explains the creation of a new entry rather than overwriting, and discloses the side effect of invalidating approvals from get_finalise_readiness. This is rich behavioral context that an agent needs to predict side effects accurately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, then branches into consequences and edge cases. While it is verbose, each sentence adds distinct information (e.g., version impact, audit trail, approval invalidation). It is structured as a coherent narrative, though some sentences could be tightened. Overall, it is appropriately detailed for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description thoroughly covers the side effects and consequences of the withdrawal, which is essential for a destructive operation. However, it does not explain how the agent should specify which answer to withdraw (beyond implicitly via capture_kind and IDs), nor does it describe any return value or response format (no output schema exists). The request semantics are under-specified for a mutation tool, leaving gaps in what the agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has only 40% coverage, with descriptions for client_id and on_behalf_of, but none for period_id, template_id, or capture_kind. The description does not explain how these parameters map to the answer being withdrawn, nor does it clarify which specific answer is targeted. The description focuses entirely on consequences and ignores the request parameters entirely, failing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb–resource pair: "Withdraw a recorded capture answer or review confirmation." It distinguishes this from recording tools like confirm_capture_review and declare_capture_na by focusing on the act of withdrawal, and the title reinforces the purpose. The purpose is unambiguous and distinct from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on what happens after using the tool, such as readiness status changes and approval invalidation. It also suggests an alternative action: "answer the section again and generate a fresh version." However, it does not explicitly state when to use this tool versus alternatives like confirm_capture_review, nor does it give exclusions. The guidance is implied but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_accounting_policiesSave significant accounting policiesA
Destructive
Inspect

Save the Significant Accounting Policies text (Note 2) your user has chosen, one policy per title. SEND THE COMPLETE SET EVERY TIME: this replaces all of Note 2, so any title you leave out of this call is removed — including one someone answered in the Datavrn app. Call list_statement_policy_choices first and send back every title. If your call would drop a saved policy, Datavrn saves nothing and returns an approval request naming how many would be dropped — show your user, and send the approval back only if they mean to drop them. A complete resend drops nothing and saves straight away. Use the exact policy headings this statement format carries; a heading Datavrn does not recognise is refused and nothing is saved. Resolving a bracketed template choice such as "[FIFO / weighted average]" is an ACCOUNTING POLICY DECISION SPECIFIC TO THIS ENTITY: get your user’s explicit choice, and never pick one because it is the common answer. If a policy was set this year and differs from last year’s answer, that is a CHANGE IN ACCOUNTING POLICY requiring disclosure under AS-5 / Ind AS 8 — tell your user before you save it. Check list_statement_policy_choices first: a row that differs from last year while its captured flag is false is NOT a change — last year’s wording has simply not been carried forward, and unless you send it again this note prints Datavrn’s generic template wording in its place. Saving here re-opens the disclosure review — after your last change, confirm the disclosure review again with confirm_capture_review before generating. Recorded as authorised by the member you name. Generate a fresh version after your last capture change — finalisation checks the version’s frozen capture state, not today’s.

ParametersJSON Schema
NameRequiredDescriptionDefault
policiesNo
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
removal_countNo
removal_tokenNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses that the call replaces all of Note 2, that omitted titles are removed, that dropping policies triggers an approval request, that unrecognized headings are refused, and that saving re-opens the disclosure review. This fully informs the agent of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence adds critical operational detail. However, it is a single dense paragraph; breaking it into bullets or numbered steps would improve scannability without loss of content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, destructive write tool with no output schema, the description covers prerequisites, destructive consequences, approval mechanics, policy change disclosure, authorisation record, and follow-up confirmation steps. It is fully complete for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only describes client_id and on_behalf_of (29% coverage). The description compensates by explaining that policies should be a complete map of titles, warns about exact headings, and refers to the removal_count/removal_token flow via the approval request. It doesn't explicitly name each parameter, but it gives the semantic context needed to populate them correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool saves the Significant Accounting Policies text (Note 2), one policy per title, making the verb and resource explicit. This distinguishes it from sibling save_* tools like save_disclosures or save_adjustments by naming the exact statement note.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit step-by-step guidance: call list_statement_policy_choices first, send the complete set every time, follow up with confirm_capture_review, and generate a fresh version after last capture change. It also explains when a difference from last year is not a change, preventing misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_adjustmentsSave adjustment entryA
Destructive
Inspect

Save one balanced adjustment journal entry (debits = credits) as an atomic whole entry. Creating a new entry proceeds immediately. Replacing an existing entry first returns an approval request; nothing changes then. Review the existing entry in the Schedule III workspace, then resend the unchanged request with the approval details to proceed. Amounts are strings in rupees.

ParametersJSON Schema
NameRequiredDescriptionDefault
linesYes
entry_noYes
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
removal_countNo
removal_tokenNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, readOnlyHint=false), the description discloses the atomic nature of the entry, the immediate effect of new entries, and the two-step approval process for replacements, including the fact that nothing changes until the request is resent. It also clarifies that amounts are strings in rupees, adding significant operational detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the primary function, and every sentence adds essential detail about the workflow and constraints. There is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description explains the high-level workflow and approval process, it omits details about how the replacement mechanism works (e.g., what 'approval details' are, how removal_count/removal_token relate) and provides no output schema. Given the tool's complexity and multiple parameters, this is a notable gap, though the core behavior is adequately covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is only 14% (solely client_id), yet the description does not compensate for the undocumented parameters such as period_id, template_id, entry_no, lines, removal_count, and removal_token. It adds some context (rupee strings, balanced entry) but leaves the meanings and usage of most parameters unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool saves a balanced adjustment journal entry (debits = credits) as an atomic whole, distinguishing it from sibling save tools. It also distinguishes between creating a new entry (immediate) and replacing an existing one (requires approval), making the purpose highly specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides concrete procedural guidance on when to expect immediate execution vs. an approval workflow, and instructs the user to resend the request with approval details. However, it does not explicitly name alternatives or state when not to use this tool compared to sibling save_* tools, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_asset_movementsSave asset movementsA
Destructive
Inspect

Save fixed-asset movements (additions, deletions, depreciation charge, depreciation on deletions) per gross-block line for the PPE schedule.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
movementsYes
period_idYes
template_idYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds context about the content being saved (movements per gross-block line) but does not disclose any additional behavior such as whether existing data is overwritten, validation rules, or prerequisites like an open period. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It communicates the verb, object, scope, and specific movement types efficiently, making every word meaningful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 required parameters, a nested object, and no output schema, the description covers the core purpose but omits information about side effects (e.g., overwriting), return values, or prerequisites. It is adequate but not comprehensive for a mutation tool with destructive potential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, with only client_id having a description. The description compensates by explaining the movement types (additions, deletions, depreciation charge, depreciation on deletions) which map directly to the nested object properties, and 'per gross-block line' clarifies gross_leaf_code. However, it does not explain period_id or template_id in detail, though their purpose is inferable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'save' and the specific resource 'fixed-asset movements' with enumerated movement types (additions, deletions, depreciation charge, depreciation on deletions). It also scopes to 'per gross-block line for the PPE schedule,' which distinguishes it from sibling save tools like save_provision_movements or save_reserves_movements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for the PPE schedule' implies a usage context, and the title/name clearly indicates when to use it. However, it does not explicitly mention when not to use it or name alternative tools, so the guidance remains implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_disclosuresSave disclosuresA
Destructive
Inspect

Save the notes/disclosures sections the user provides for the statement. Some of these sections IDENTIFY PEOPLE BY NAME — shareholders, promoters and related parties — so send only what your user has given you, exactly as they gave it. SECTIONS YOU DO NOT SEND ARE LEFT ALONE. Send only the sections you are changing; every other section keeps exactly what is saved. Within a section you DO send, the rows you send REPLACE every row saved for that section — there is no row-level merge, so always send that section complete. To CLEAR a section, send it explicitly with its own empty value: [] for corporateInfo, contingent, shareholders5pct, promoters, relatedParties and msmeAmounts; {} for ratioReasons; null for auditorPayments, csr, proposedDividend and the DSCR amounts. The three ageing sections take [] or null. Clearing removes saved content, so Datavrn saves nothing and returns an approval request naming the sections and how many rows each holds. Show your user exactly what would be cleared; only if they mean to, resend the identical call adding removal_token and removal_count from that response. The approval is single-use and lapses after five minutes — if it lapses, call again for a fresh one. Saving here re-opens the disclosure review — after your last change, confirm the disclosure review again with confirm_capture_review before generating. ratioReasons is keyed by the ratio IDENTIFIER, never its display label: current_ratio | debt_equity | dscr | roe | inv_turnover | tr_turnover | tp_turnover | ncap_turnover | net_profit | roce | roi. msmeAmounts is a POSITIONAL list of exactly five items, in this statutory order: 1. Principal amount and interest due thereon remaining unpaid at the year end (shown separately) | 2. Interest paid under section 16 of the MSMED Act, with the payment made beyond the appointed day | 3. Interest due and payable for delay in payment (paid beyond the appointed day), excluding MSMED-specified interest | 4. Interest accrued and remaining unpaid at the year end | 5. Further interest remaining due and payable in succeeding years (section 23 disallowance) — send all five, using {"currentPaise": null, "previousPaise": null} for an item with nothing to report. Dashes and capitalisation are forgiven when matching a row label (a hyphen for an en dash is fine) and Datavrn stores the prescribed spelling; the wording itself is not forgiven. PRESCRIBED ROWS BY STATEMENT FORMAT — a row label or sub-head that is not one of these is refused rather than stored (a value already saved on this statement stays editable): Division I (Non-Ind AS) ("schedule_iii_div1") — trAgeing: 6 buckets in this order [Not due | < 6 months | 6 months – 1 year | 1 – 2 years | 2 – 3 years | > 3 years], rows [Undisputed — considered good | Undisputed — considered doubtful | Disputed — considered good | Disputed — considered doubtful | Unbilled dues]; tpAgeing: 5 buckets in this order [Not due | < 1 year | 1 – 2 years | 2 – 3 years | > 3 years], rows [MSME | Others | Disputed dues — MSME | Disputed dues — Others | Unbilled dues]; cwipAgeing: 4 buckets in this order [< 1 year | 1 – 2 years | 2 – 3 years | > 3 years], rows [Projects in progress | Projects temporarily suspended]. contingent.subhead ∈ contingent liabilities [Claims against the company not acknowledged as debt | Guarantees | Other money for which the company is contingently liable] or commitments [Estimated amount of contracts remaining to be executed on capital account (net of advances) | Uncalled liability on shares and other investments partly paid | Other commitments] — set group to "contingent" or "commitment" to match, or the row is totalled under contingent liabilities (or leave subhead out and the row prints as its own line under its group). Division II (Ind AS) ("schedule_iii_div2") — trAgeing: 6 buckets in this order [Not due | < 6 months | 6 months – 1 year | 1 – 2 years | 2 – 3 years | > 3 years], rows [Undisputed — considered good | Undisputed — significant increase in credit risk | Undisputed — credit impaired | Disputed — considered good | Disputed — significant increase in credit risk | Disputed — credit impaired | Unbilled dues]; tpAgeing: 5 buckets in this order [Not due | < 1 year | 1 – 2 years | 2 – 3 years | > 3 years], rows [MSME | Others | Disputed dues — MSME | Disputed dues — Others | Unbilled dues]; cwipAgeing: 4 buckets in this order [< 1 year | 1 – 2 years | 2 – 3 years | > 3 years], rows [Projects in progress | Projects temporarily suspended]. contingent.subhead ∈ contingent liabilities [Claims against the company not acknowledged as debt | Guarantees excluding financial guarantees | Other money for which the company is contingently liable] or commitments [Estimated amount of contracts remaining to be executed on capital account (net of advances) | Uncalled liability on shares and other investments partly paid | Other commitments] — set group to "contingent" or "commitment" to match, or the row is totalled under contingent liabilities (or leave subhead out and the row prints as its own line under its group). Non-Corporate Entity ("icai_nce") and LLP ("icai_llp") — no ageing tables at all — trAgeing/tpAgeing/cwipAgeing are refused. contingent.subhead ∈ contingent liabilities [Claims against the entity not acknowledged as debt | Guarantees given on behalf of others | Other money for which the entity is contingently liable] or commitments [Estimated amount of contracts remaining to be executed on capital account (net of advances) | Uncalled liability on investments partly paid | Other commitments] — set group to "contingent" or "commitment" to match, or the row is totalled under contingent liabilities (or leave subhead out and the row prints as its own line under its group).

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadYes
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
removal_countNo
removal_tokenNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark destructiveHint=true, and the description goes far beyond by detailing the exact replace semantics (rows sent replace all saved rows), the clearing behavior with approval flow, and the fact that un-sent sections are left untouched. It even warns that clearing triggers an approval request and that tokens lapse after five minutes, adding substantial behavioral context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely long and dense, packed with essential rules but presented as a single wall of text without lists or tables. While every sentence adds value, the sheer length and lack of structure hurt readability; the most critical behavior (replace-not-merge) is front-loaded early, but the exhaustive prescribed-rows section could be better organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description covers every aspect an agent needs to call the tool correctly: sections to send, clearing semantics, the removal_token approval flow, prescribed rows per statement format, ratioReasons keying, msmeAmounts ordering, and the subsequent confirm_capture_review step. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17%, so the description carries the full burden of explaining parameters. It specifies how each section is structured, the msmeAmounts positional list of five items, the ratioReasons keyed by identifiers never display labels, and the exact clearing values for each section type—all information not present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it saves the notes/disclosures sections, a specific verb+resource that distinguishes it from other save_* tools like save_accounting_policies. It immediately specifies the scope and the non-merge replace behavior, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives detailed instructions on how to use the tool—such as sending only changed sections, using removal_token for clearing, and confirming disclosure review afterward. However, it does not explicitly contrast this tool against alternatives like save_accounting_policies or save_statement_settings, so choosing between siblings is left to inference rather than clearly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_partner_capitalSave partner or owner capital scheduleA
Destructive
Inspect

Save the partner or owner capital schedule for an LLP or other non-corporate entity — Note 3a (capital account) or Note 3b (current account), one section per call. Send the COMPLETE schedule for the section you name: anyone you leave out is removed, and a renamed partner reads as one removal plus one addition. If your request would remove anyone, would change the figures of a partner who stays — including their profit-sharing ratio — or would repeat a person’s name that is not already repeated on file, this returns an approval request first and changes NOTHING; tell your user exactly what would change and get their go-ahead before resending with the approval. Two rows with the same person name are both kept: Datavrn never merges them, because two partners may genuinely share a name. A repeated name is therefore saved as a separate row each time it appears, and every one of those rows adds to that person’s balance on the note, so check with your user that there really are that many people before you send a schedule that repeats a name. Adding a partner with a new, different name needs no approval. share_pct is the profit-sharing ratio stated in the LLP or partnership agreement at the reporting date. A ratio that changed part-way through the year cannot be represented here, and a partner who joined or left mid-year cannot be dated — record the position at the reporting date and tell your user the note cannot explain a mid-year change. These figures come from the agreement; Datavrn does not verify them. Datavrn warns about ratios that do not total 100% only when at least TWO capital-account rows carry a ratio, so a partly-filled schedule gets no warning at all — check the total yourself with get_partner_capital. This is your user’s data about named people, recorded as authorised by them. After a successful save, read the section back with get_partner_capital and show your user the new closing balances — this replaces the whole section, and a schedule they cannot see afterwards is a schedule they cannot check. Generate a fresh version after your last capture change — finalisation checks the version’s frozen capture state, not today’s.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsYes
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
account_kindYes
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
removal_countNo
removal_tokenNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (destructiveHint=true), the description discloses the full-replacement semantics ('anyone you leave out is removed'), the approval-request behavior that changes nothing, duplicate-name handling (kept as separate rows), and the fact that Datavrn does not verify figures. It also warns about the ratio-total warning limitation. All of this is essential behavioral context not present in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but nearly every sentence adds unique critical information, from approval flow to version finalisation. It is front-loaded with the primary purpose and then expands into necessary warnings, though a few redundant phrases (e.g., 'this is your user's data...') could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's destructive nature, 8 parameters, and no output schema, the description is exceptionally complete. It explains the overall process, success behavior, edge cases (mid-year changes, duplicate names), approval requirements, and post-save verification steps, leaving almost no operational ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25%, but the description compensates by explaining share_pct's meaning and the behavior of rows (complete schedule, duplicates, removals). It maps account_kind to Notes 3a/3b. However, it does not explicitly describe fields like opening/interest or the removal_token/count, leaving some parameters underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Save the partner or owner capital schedule for an LLP or other non-corporate entity — Note 3a (capital account) or Note 3b (current account).' It distinguishes itself from sibling tools like get_partner_capital and save_share_capital by specifying the entity type and the two account kinds.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the tool is for one section per call and emphasizes sending the complete schedule. It instructs the agent to verify totals and read back with get_partner_capital, and to generate a fresh version after capture changes, giving concrete guidance on when to use this tool in conjunction with others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_provision_movementsSave provision movementsA
Destructive
Inspect

Upsert provision movements (additions, amounts utilised) per provision line. Omitted saved lines stay unchanged. To remove selected saved lines, pass remove_leaf_codes; to remove the entire saved set, pass clear_all (never both). An actual removal first returns an approval request; nothing changes then. Review current movements in the Schedule III workspace, then resend the unchanged request with the approval details to proceed.

ParametersJSON Schema
NameRequiredDescriptionDefault
clear_allNo
client_idYesThe entity (client) id — from list_clients.
movementsYes
period_idYes
template_idYes
removal_countNo
removal_tokenNo
remove_leaf_codesNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotation destructiveHint=true, the description discloses critical behavioral traits: omitted saved lines stay unchanged (upsert semantics), removals first trigger an approval request and nothing changes until the request is resent with approval details, and the two removal modes are mutually exclusive. This provides significant operational context beyond the mere 'destructive' flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences long and front-loaded with the core action. Each sentence adds valuable detail: upsert behavior, omission semantics, removal modes, approval flow, and required workflow. It is slightly longer than ideal but appropriately so given the tool's complexity and destructive potential.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with destructive capabilities and an approval workflow, the description covers the essential semantics: upsert behavior, removal handling, and the resend-with-approval process. It does not explicitly state the success response for regular saves, but given the output schema is absent and the workflow is clearly outlined, this is a minor gap. Overall, it gives an agent enough context to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (13%), and the description compensates well by explaining key parameters: remove_leaf_codes, clear_all, and the movements array (additions, utilised). It also hints at approval tokens via 'approval details.' However, it does not explain template_id or period_id, nor the exact format constraints for additions, so the compensation is strong but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb 'Upsert' and clearly identifies the resource as 'provision movements' per provision line. It also differentiates from sibling save tools (e.g., save_asset_movements) by naming the target. The additional detailing of removal modes and the upsert semantics leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool (to save provision movements) and how to handle removals (via remove_leaf_codes or clear_all, never both). It also outlines a workflow: review current movements in the Schedule III workspace, then resend with approval details. While it does not explicitly name alternative tools for other movement types, the context is clear enough for an agent to select this tool appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_py_valuesSave prior-year comparativesA
Destructive
Inspect

Override the prior-year comparative for one or more statement lines with an audited figure. The prior-year column fills itself automatically from the previous year's Trial Balance read through the current groupings, so use this only when the audited financial statements differ from that figure (for example appropriations booked outside the ledger), or when there is no previous-year Trial Balance to derive from — ask your user for the audited figures in those cases.

ParametersJSON Schema
NameRequiredDescriptionDefault
valuesYes
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, and the description adds valuable behavioral context: the prior-year column auto-fills from the previous year's Trial Balance, so this tool is an exception override. It doesn't detail side effects beyond 'Override', but the annotation and verb together convey the destructive nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact: two sentences, the first stating purpose and the second providing usage conditions. Every sentence adds essential information with no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 parameters, no output schema) and the presence of annotations, the description covers purpose, usage conditions, and user interaction. However, it lacks explicit instructions on constructing the values array and what to expect after the save operation, leaving some gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only client_id has a description), and the description does not explain leaf_code, amount, period_id, or template_id semantics. The phrase 'audited figure' hints at the values parameter but does not clarify how leaf_code maps to statement lines or how amounts should be formatted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Override' and clearly identifies the resource as 'prior-year comparative for statement lines'. It also mentions the automatic fill behavior, which distinguishes this tool from siblings like save_adjustments or save_disclosures.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'use this only when' and lists two precise conditions (audited financial statements differ, or no previous-year Trial Balance). It also instructs the agent to ask the user for audited figures in those cases, providing clear decision guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_regulatory_affirmationsSave regulatory affirmationsA
Destructive
Inspect

Save the CARO / Other Regulatory Information affirmations your user has confirmed, one per title. SEND THE COMPLETE SET EVERY TIME: this replaces the whole Other Regulatory Information note, so any title you leave out of this call is removed — including one someone answered in the Datavrn app. Call list_statement_policy_choices first and send back every title. If your call would drop a saved affirmation, Datavrn saves nothing and returns an approval request naming how many would be dropped — show your user, and send the approval back only if they mean to drop them. A complete resend drops nothing and saves straight away. Use the exact affirmation headings this statement format carries; a heading Datavrn does not recognise is refused and nothing is saved. Some statement formats — the ICAI formats for LLPs and non-corporate entities — carry no Other Regulatory Information note at all, and this tool refuses for them. Each affirmation is a REGULATORY REPRESENTATION made in the entity’s name — for example whether any proceedings for benami property are pending, or whether the entity has been declared a wilful defaulter. Put each one to your user individually and record their answer. Never affirm one because it is the usual answer, never infer one from a template default, and never confirm a batch of them in one go. If an affirmation differs from last year’s answer, tell your user — call list_statement_policy_choices to see what was answered last year. Saving here re-opens the disclosure review — after your last change, confirm the disclosure review again with confirm_capture_review before generating. Recorded as authorised by the member you name. Generate a fresh version after your last capture change — finalisation checks the version’s frozen capture state, not today’s.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
affirmationsNo
on_behalf_ofNoThe email address your user signs in to Datavrn with. This records who authorised the change alongside the connection that made it. REQUIRED on an API-key connection — ask your user for it, do not guess. On an OAuth connection leave it out: the change is recorded as authorised by the member who connected; if you do supply it, it must be that member.
removal_countNo
removal_tokenNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true, but the description goes far beyond: it explains the replace-all behavior, the approval request for dropped affirmations, refusal on unrecognised headings, and side effects like reopening the disclosure review. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough but verbose, with repeated instructions (list_statement_policy_choices appears twice) and examples that could be condensed. While the length is somewhat justified by the destructive and complex nature of the tool, the redundancy prevents a higher score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the full workflow for a complex, destructive tool: preconditions, replacement semantics, approval and refusal paths, user consent rules, and post-save actions (reopen disclosure review, generate a fresh version). Despite no output schema, it gives enough context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description must compensate. It explains affirmations must be the complete set with exact headings and one per title, and it indirectly covers removal_count/removal_token through the approval flow, but does not explicitly mention these parameters. client_id, template_id, and period_id are not described beyond the schema, leaving some gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool saves CARO / Other Regulatory Information affirmations, one per title, and explicitly says it replaces the whole note. This specific verb+resource distinguishes it from sibling save_* tools like save_disclosures and save_accounting_policies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit pre-requisites: 'Call list_statement_policy_choices first and send back every title.' It also gives post-requisites: 'confirm the disclosure review again with confirm_capture_review before generating.' It names exclusions: 'the ICAI formats for LLPs and non-corporate entities — carry no Other Regulatory Information note at all, and this tool refuses for them.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_reserves_movementsSave reserves movementsA
Destructive
Inspect

Upsert reserves/equity movements (transfers in/out, dividends, other changes) per reserves line. Omitted saved lines stay unchanged. To remove selected saved lines, pass remove_leaf_codes; to remove the entire saved set, pass clear_all (never both). An actual removal first returns an approval request; nothing changes then. Review current movements in the Schedule III workspace, then resend the unchanged request with the approval details to proceed.

ParametersJSON Schema
NameRequiredDescriptionDefault
clear_allNo
client_idYesThe entity (client) id — from list_clients.
movementsYes
period_idYes
template_idYes
removal_countNo
removal_tokenNo
remove_leaf_codesNo
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true, but the description adds substantial behavioral detail: 'Omitted saved lines stay unchanged,' 'An actual removal first returns an approval request; nothing changes then,' and the need to resend with approval details. This goes well beyond the binary annotation flags and clarifies the side effects of removal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, with five sentences each serving a distinct purpose: primary action, update behavior, removal modes, approval flow, and procedural next step. It is front-loaded with the core verb and resource, and there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main behaviors (upsert, removal modes, approval workflow) and even gives a procedural hint ('Review current movements in the Schedule III workspace, then resend...'). However, with no output schema and low schema parameter explanation, it does not specify what the tool returns or the exact form of the 'approval details' (e.g., removal_token/removal_count), leaving some ambiguity for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%, so the description carries the burden. It explains the removal parameters (remove_leaf_codes, clear_all) and their exclusivity, and mentions the movement types (transfers in/out, dividends, other changes). However, it does not explain the core required parameters (client_id, template_id, period_id) in any detail, and the structure of 'movements' is left to the schema. It adds some meaning but does not fully compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Upsert reserves/equity movements (transfers in/out, dividends, other changes) per reserves line.' This specifies the action (upsert), the resource (reserves/equity movements), and even the type of movements, distinguishing it from sibling tools like save_asset_movements and save_provision_movements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (for reserves/equity movements) and gives explicit instructions for removal modes: 'pass remove_leaf_codes; to remove the entire saved set, pass clear_all (never both).' It also explains the approval resend workflow. However, it does not explicitly name alternatives or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_share_capitalSave share capitalA
Destructive
Inspect

Save the share-capital reconciliation for the period: the opening share count, the shares issued and bought back during the year, and the amount issued and the amount bought back.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesThe entity (client) id — from list_clients.
period_idYes
template_idYes
amount_issuedYes
shares_issuedYes
opening_sharesYes
amount_bought_backYes
shares_bought_backYes
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false. The description adds no behavioral context beyond that, such as overwriting existing data, permission requirements, or side effects. It merely lists the data fields, which is not behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the action and resource, then enumerates the key fields. Every word earns its place, and there is no wasteful repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given its complexity (8 required parameters), lack of output schema, and very low schema description coverage, the description does not provide enough context. It does not explain prerequisites, return values, or side effects. The listing of fields is helpful but insufficient for an agent to fully understand the tool's behavior without additional schema information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at only 13%, the description compensates partially by explaining the meaning of several fields (opening share count, shares issued/bought back, amounts). However, it does not cover client_id, template_id, or period_id explicitly, and it does not explain formats or nullability. This is an improvement over the bare schema but incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Save the share-capital reconciliation for the period' with a specific verb ('Save') and resource ('share-capital reconciliation'). This distinguishes it from sibling tools like save_partner_capital and save_reserves_movements by explicitly naming the share-capital domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context that this tool is for saving share-capital reconciliation data for a period, implying when to use it. It does not mention alternatives or exclusions, but the context is unambiguous enough to guide selection among similar save tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_statement_settingsSave statement settingsA
Destructive
Inspect

Save statement settings (rounding unit, signatory details, company information used on the statement face).

ParametersJSON Schema
NameRequiredDescriptionDefault
cinYes
client_idYesThe entity (client) id — from list_clients.
font_nameNoCalibri
period_idYes
unit_labelYes
include_cfsNo
template_idYes
entity_classNo
rounding_offNoWhole-rupee value rounding. When true (the default), every figure is rounded to the nearest rupee and the net rounding difference is posted as a computed "Rounding off" line under Other expenses so the balance sheet ties to the rupee. false renders exact two-decimal figures.
unit_divisorYes
equity_face_valueYes
registered_officeYes
suppress_nil_linesNo
entity_name_overrideYes
weighted_equity_sharesYes
weighted_equity_shares_pyYes
weighted_potential_equity_sharesYes
weighted_potential_equity_shares_pyYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation aspect is covered. The description adds context about the settings affected (rounding, signatory, company info), which is useful. However, it does not disclose overwrite behavior, required permissions, or side effects beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. The parenthetical adds concrete detail efficiently, making it one of the most concise yet informative descriptions among typical tools.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with 18 parameters, 13 required, and no output schema. The description only provides a high-level summary and does not explain required fields, save semantics, or what response to expect. An agent cannot reliably invoke this tool correctly without additional information beyond the sparse schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11% (only client_id and rounding_off have descriptions). The description names broad categories but fails to explain key required parameters like weighted_equity_shares, weighted_potential_equity_shares_py, or the distinction between unit_label and unit_divisor. With 18 parameters, the description does not compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Save') and resource ('statement settings') and lists concrete categories ('rounding unit, signatory details, company information used on the statement face'). This distinguishes it from sibling save tools such as save_accounting_policies or save_disclosures, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for saving statement-specific settings but provides no explicit when-to-use guidance, prerequisites, or comparison with alternatives like save_accounting_policies or save_share_capital. There are many save_* sibling tools, so an agent would need more direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_header_rowSet header rowA
Destructive
Inspect

Correct the detected header row of a staged upload (1-based). Only needed when get_upload/upload_trial_balance shows the wrong row was detected.

ParametersJSON Schema
NameRequiredDescriptionDefault
upload_idYesThe upload session id returned by upload_trial_balance.
header_rowYes
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutating nature is known. The description adds the 'staged upload' context and '1-based' indexing, but does not disclose side effects like whether the correction is reversible or affects downstream processing. It is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first states the purpose, the second gives the usage condition. No fluff, front-loaded with the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation tool with helpful annotations and a clear description, all necessary context is provided. The tool's use case, parameter semantics, and precondition are all apparent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers upload_id well but leaves header_row without a description. The description compensates by clarifying that header_row is 1-based and is the row to set as the header, adding key meaning not present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Correct' and the resource 'detected header row of a staged upload', making the tool's function immediately obvious. It also distinguishes itself from siblings by referencing get_upload/upload_trial_balance as the source of detection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'Only needed when get_upload/upload_trial_balance shows the wrong row was detected.' This gives a clear condition and implicitly names the alternative tools for checking detection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_trial_balanceUpload trial balanceAInspect

Stage a Trial Balance spreadsheet (xlsx or csv, max 4 MB) for an entity by INLINING its bytes as base64. This path is ONLY for programmatic callers (a script, Claude Code, an automation) that already have the raw file on disk. If a HUMAN has the file — e.g. they attached it to this chat — do NOT use this tool and do NOT base64-encode the file: call create_upload_link instead and give them the link to upload it in their browser. File size does not change this: even a small attached file goes through create_upload_link — inlining a human-supplied file is unreliable and its bytes routinely truncate. Even for a file you hold on disk, prefer create_upload_link once the file is larger than about 10 KB: base64 through a model context mutates a token often enough that the damage lands as a PLAUSIBLE trial balance, not as an obvious error. VERIFY THE HASH BEFORE YOU CONFIRM ANYTHING: this tool returns received_file_hash, the sha256 of the bytes the server actually received. Compute the sha256 of the file on your disk and compare the two. If they differ, the bytes changed in transit — do NOT call confirm_column_mapping on this session; re-send the file with create_upload_link instead. A mismatched file can still parse cleanly and still show sensible columns, so the hash is the only reliable check. Returns the upload session with detected columns and mapping SUGGESTIONS — nothing is ingested yet. Next: verify received_file_hash, then review the suggested column mapping with your user, then call confirm_column_mapping.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoLeave unset — the assistant handles Trial Balances only; any other format is refused (use the Datavrn web app).
sourceNoSet 'tally_file' when the file is a Tally xlsx export; omit otherwise.
client_idYesThe entity (client) id — from list_clients.
file_nameYesThe file's name, e.g. 'tb-2026-03.xlsx'.
file_base64YesThe file bytes, base64-encoded.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false, destructiveHint=false, leaving the description to carry the behavioral burden. It does so thoroughly: explains the risk of base64 mutation and truncation, that nothing is ingested yet, and that received_file_hash is the only reliable check. It even discloses that a corrupted file can appear plausible. This far exceeds minimal expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly packed with critical guidance. It opens with the core purpose, then flows through exclusions, size advisory, hash verification, and next steps. Every sentence adds operational value; none are filler. The structure mirrors the decision order an agent would follow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description fully explains what the tool returns (upload session with detected columns and mapping suggestions) and the required follow-up (verify hash, review mapping, call confirm_column_mapping). It covers all parameters, failure modes, and the alternative path. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond schema by explaining when to set source ('tally_file' for Tally exports), why format should be left unset, and that client_id comes from list_clients. It reinforces that file_base64 must be raw bytes encoded as base64. This extra context justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (stage), resource (Trial Balance spreadsheet), format (xlsx/csv), size limit (max 4 MB), and method (inlining base64). It clearly distinguishes itself from the sibling create_upload_link, so an agent cannot mistake which tool to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly reserves this tool for programmatic callers holding the file on disk, and outright forbids its use for human-supplied files, naming create_upload_link as the alternative. It also gives a concrete size threshold (prefer create_upload_link above ~10 KB) with a rationale, so usage conditions are unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_connectionVerify connectionA
Read-only
Inspect

Confirm the Datavrn connection is working and report what it can do. Call this first — or whenever the user asks whether Datavrn is connected — to get back the organization, the access profile (what this connection may see and do), and the next step. Running it successfully also marks the connection healthy in the user’s Datavrn settings.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states 'Running it successfully also marks the connection healthy in the user’s Datavrn settings', which is a side-effect write. However, annotations declare readOnlyHint=true, meaning no writes. This is a direct contradiction between description and annotation, making it impossible for the agent to trust either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the primary purpose, followed by usage guidance and side-effect disclosure. Every sentence earns its place, with no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and no output schema, the description comprehensively explains what the tool returns: organization, access profile, and next step. It also mentions the side effect, making it contextually complete for an agent to decide when and how to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema is empty, so there is nothing to explain. Per baseline for 0 params, a score of 4 is appropriate; the description adds no parameter details but doesn't need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a connection verification utility, using specific verbs 'Confirm' and 'report' on the 'Datavrn connection' resource. It distinguishes itself from sibling tools by focusing on connection status and capabilities, not data operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Call this first — or whenever the user asks whether Datavrn is connected'. This provides clear direct guidance and implies it is the appropriate tool for connection checks, with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Frequently Asked Questions

Discussions

No comments yet. Be the first to start the discussion!

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Deterministic verification for AI-generated analysis. Reconciliation, consistency and Excel-integrity checks that stop the line when the numbers don't add up.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local MCP server that connects Claude to real Excel .xlsx files for financial and operational analysis. Enables reading, comparing, cleaning, reconciling, and writing to Excel files without cloud dependency.
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    MCP server that automates audit risk assessment from Korean DART disclosures, providing deterministic risk signals, going-concern scoring, and memo draft generation without LLM judgment.
  • F
    license
    B
    quality
    B
    maintenance
    Enables AI assistants to perform financial reconciliation with a deterministic proof engine: intake files, match transactions, verify proofs, resolve exceptions, and sign off on balanced journals under the user's authority.
    21
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources