VeteranHQ
Server Details
VA disability rating and compensation calculations, condition lookup, and 38 CFR authority search
- Status
- Healthy
- Uptime
- 99.9% over 42 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 13 tools
Most tools have clearly distinct purposes keyed to distinct inputs (diagnostic code, condition name, rating percentage, slug, dates). There is real overlap between compare_rating_criteria and search_legal_authority, both of which return VASRD rating criteria for conditions/codes, and lookup_compensation_rate vs compute_retroactive_pay sit close together. The long descriptions do provide enough framing to separate them, but an agent could plausibly misselect on the criteria tools.
Every name is snake_case verb_noun, so the surface reads as a coherent family. The verb choice varies (find_ vs search_ vs lookup_ vs get_) in ways that are not fully predictable from the name alone, which is a minor deviation rather than a break in convention.
13 tools is squarely in the well-scoped range for a domain this broad (rating math, criteria lookup, eligibility, guides, pay). Each tool maps to a distinct veteran task and none appears redundant filler.
The surface covers ratings, criteria, combined ratings, retroactive pay, current rates, secondary conditions, presumptive eligibility, legal authority, guides, C&P prep and claim help — a strong lifecycle. Gaps remain around decision review/appeals and effective-date tooling, though some of that is an explicit out-of-scope boundary rather than a dead end.
Available Tools
13 toolsanalyze_rating_gapAnalyze a rating gapARead-onlyIdempotentInspect
Use this when a veteran gives a condition name and the percentage currently assigned for it, and wants that percentage compared against the VASRD criteria. Returns the tier that matches the current percentage, the next tier above it, and every higher tier on file for that condition, with the criteria each one requires. It reports when the condition is already at the highest documented tier, and when no criteria are on file for that condition name. documentedTiers lists the tier percentages on file for the condition, and currentPercentIsDocumentedTier says whether the percentage supplied is one of them: false means the figure does not correspond to a tier in the criteria on file, which happens when the rating came from a formula or a code this search did not resolve. For a veteran at zero percent, zeroPercentBasis is "documented" when a zero percent row for the code is on file and "unresolved" when the complete tier set could not be established, and zeroPercentUnderFourThirtyOne answers whether 38 CFR 4.31 is what assigns that zero: false where the schedule on file provides a zero percent evaluation or the rating is compensable, and null where the question is open. A null there is an unanswered question rather than a finding that 4.31 applies, because the rule turns on what the schedule provides for the code and this tool reads the criteria on file rather than the omissions in the schedule. It describes what the criteria require; only VA decides a rating.
| Name | Required | Description | Default |
|---|---|---|---|
| conditionName | Yes | Condition to analyze (e.g., "PTSD", "Lumbar strain"). | |
| currentPercent | Yes | Current VA rating percentage for this condition (e.g., 30, 50, 70). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior, and the description adds substantial context beyond them: it enumerates the return shape (current-tier match, next tier above, all higher tiers with criteria), and documents edge-case behavior for highest documented tier, missing criteria, undocumented percentages, and zero-percent handling including null semantics. It also adds the important scope caveat that "only VA decides a rating."
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The usage trigger is correctly front-loaded, but the body is a dense, run-on block that mixes return-field semantics into a single paragraph and re-explains the null/wording distinction in several clauses. Given no output schema this detail is defensible, but readability and structure suffer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of explaining return values and does so thoroughly: it covers the tier comparison output, the documentedTiers/currentPercentIsDocumentedTier fields, and the zeroPercentBasis and zeroPercentUnderFourThirtyOne branches with their interpretive caveats. Nothing essential to understanding the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both conditionName and currentPercent are documented with examples in the schema. The description adds only a paraphrase ("condition name and the percentage currently assigned"), with no extra syntax, format, or range guidance, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource+scope: it compares a currently assigned condition percentage against VASRD criteria and returns the matching tier plus all higher tiers. An agent can tell what it produces. It does not, however, differentiate itself from the sibling compare_rating_criteria, which sounds closely related, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger: use it "when a veteran gives a condition name and the percentage currently assigned for it, and wants that percentage compared against the VASRD criteria." That is a clear usage context, but it names no alternatives (e.g., compare_rating_criteria) and no when-not-to-use exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calculate_combined_ratingCalculate combined VA ratingARead-onlyIdempotentInspect
Use this when a veteran has two or more VA disability rating percentages and wants the combined rating. Applies the 38 CFR 4.25 Combined Ratings Table procedure: pairwise combination in severity order, whole-percent rounding after each step, then conversion to the nearest 10. Returns combinedRating as an object holding both figures under their own names: rounded is the rating after that conversion, raw is the Combined Ratings Table value before it. The arithmetic for each step comes back alongside them. The optional bilateral argument tags the ratings that affect both arms or both legs; when it is supplied and one pair has a compensable rating on each side, 38 CFR 4.26 applies to those ratings first: they combine with each other, 10 percent of that value is added rather than combined, and the result enters the 4.25 combination as one disability. Without that argument no bilateral factor is applied. The 38 CFR 4.26(d) comparison across eligible groupings is exhaustive here for up to 16 tagged compensable ratings; past that the response carries combinedRatingWithoutBilateralFactor in place of combinedRating, labelled as a figure the bilateral factor has not been applied to, with a note that VA and accredited representatives compute that factor. The result differs from adding the percentages and from a simple product. It reports what 38 CFR 4.25 and 4.26 yield for the percentages supplied, not what VA has assigned, and it does not evaluate special monthly compensation.
| Name | Required | Description | Default |
|---|---|---|---|
| ratings | Yes | Array of individual disability rating percentages (0-100), e.g., [70, 50]. | |
| bilateral | No | Optional. Tags the ratings that affect a paired extremity, so 38 CFR 4.26 can be applied: those ratings are combined with each other first, 10 percent of that value is added (not combined), and the result enters the 4.25 combination as one disability. The factor needs a compensable rating on both sides of one pair (both arms or both legs); tag every rating that affects an arm or a leg and the tool applies the factor only where the regulation allows it. Tagging more than 16 compensable ratings is accepted and answered, but past that the 38 CFR 4.26(d) comparison across eligible groupings is not exhaustive here, so the result carries the 4.25 combination of every rating supplied under combinedRatingWithoutBilateralFactor instead of a combined rating, and says where the complete calculation comes from. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only/idempotent/non-destructive, and the description adds substantial extra behavior: the step-by-step algorithm, the shape of the return (rounded vs raw figures plus per-step arithmetic), the bilateral-factor edge case, and the >16-rating fallback field. It also discloses limitations (regulatory output only, no SMC).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a dense block, but the domain is genuinely complex and the trigger is front-loaded before the mechanics. Nearly every sentence carries required information; only minor tightening (bulleting the 4.25/4.26/fallback branches) would improve scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully compensates by enumerating the returned fields (combinedRating.rounded, .raw, arithmetic, and the fallback combinedRatingWithoutBilateralFactor) and their meaning. Nothing an agent needs to invoke and interpret the call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description meaningfully expands on the optional bilateral argument beyond the schema: how tagged pairs combine, the 10-percent addition rule, the both-sides compensable requirement, and the >16-tagged cap behavior. This adds real interpretive value over the field text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (compute the combined VA disability rating) and names the governing regulation (38 CFR 4.25). It clearly distinguishes the result from naive addition or multiplication and from the separate compute_retroactive_pay / lookup_compensation_rate siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use this when a veteran has two or more VA disability rating percentages and wants the combined rating') and boundaries (does not report what VA assigned, does not evaluate special monthly compensation). It does not, however, route to a named alternative sibling when the input is a single rating or a pay question.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calculate_rolling_windowCheck a 12-month absence windowARead-onlyIdempotentInspect
Use this when a veteran supplies dates of incapacitating episodes and the question involves the 38 CFR 4.71a formula for rating intervertebral disc syndrome on incapacitating episodes, diagnostic code 5243. Totals the calendar duration of those episodes inside the 12 months ending on a reference date: the window opens the day after the same date one year earlier and closes on the reference date, with both ends counted, so an episode that began exactly 12 months before that date falls outside it by one day. That formula bands on total calendar duration in weeks, so weekend and holiday days inside an episode count, and only periods flagged as bed rest prescribed by a physician count toward a tier. Returns the window, a per-period breakdown, the tier the flagged episodes would support, and a separate employer-leave workday count that plays no part in that tier. Inside ivdsThresholdAnalysis, calendarDaysBelow60PercentThreshold and weeksBelow60PercentThreshold report how far the counted episodes fall below the 6-week duration that bands at 60 percent under that formula. They measure the reported history against a threshold in the regulation; they are not a target, since a veteran does not accumulate bed rest to reach a rating. The tier it reports is an estimate from the episodes supplied. It does not decide FMLA entitlement and does not apply to conditions rated outside diagnostic code 5243.
| Name | Required | Description | Default |
|---|---|---|---|
| referenceDate | No | The reference date for the 12-month rolling window (YYYY-MM-DD). Defaults to today. Typically the anticipated C&P exam date. Must be a real calendar date; an unparseable value is rejected. | |
| absencePeriods | Yes | Array of absence periods. Each must have at least a startDate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive behavior, so the safety profile is covered. The description adds substantial behavioral context beyond that: window semantics with a one-day exclusion edge case, weekend/holiday inclusion, the physician-prescribed bed rest requirement, and the tier being an estimate rather than a decision. This earns a 3 rather than higher only because some of this overlaps with what the schema already specifies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long and dense but front-loaded with the usage condition and organized around window logic, outputs, and exclusions. Slight redundancy (restating the bed rest rule that the schema also carries) keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the burden of describing returns (window, per-period breakdown, tier, employer-leave workday count) and the threshold fields. It also covers legal scope limits, so nothing an agent needs to call or interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds interpretive meaning beyond the schema by clarifying that only physician-prescribed bed rest periods count toward a tier and that weekend/holiday days inside an episode count toward calendar duration, reinforcing why the boolean and date fields behave as they do.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action (totals the calendar duration of incapacitating episodes inside a 12-month window) tied to a specific legal basis (38 CFR 4.71a, diagnostic code 5243). An agent can distinguish this from generic rating calculators in the sibling list without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Opens with an explicit trigger (veteran supplies episode dates and the question involves the 4.71a IVDS formula/DC 5243) and closes with explicit exclusions (does not decide FMLA entitlement, does not apply outside DC 5243). When-to-use and when-not are both stated, not inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_presumptive_eligibilityCheck presumptive eligibilityARead-onlyIdempotentInspect
Use this when a veteran asks whether a condition is presumptively service connected, or which conditions are presumptive for a given exposure or service era. Returns matching presumptive conditions with the service era, exposure type, required service, legal authority and evidence needed for each, and repeats that legal authority in a provenance block on every row, with the date the presumption took effect where one date is true of the whole row; that date can be null for a row that groups conditions added on different dates. At least one of condition, serviceEra or exposureType is required. Filters combine with AND: condition plus exposureType or serviceEra narrows to their intersection, and each filter needs at least one word of three or more characters or the query is refused. A condition description that names a qualifying circumstance, such as a diagnosis that has to come after the qualifying service, states a limit of the presumption rather than a detail beside it. An empty result means no entry satisfies that exact combination, not that the condition is non-presumptive. Whether a particular veteran meets the service requirement depends on service records this tool does not read.
| Name | Required | Description | Default |
|---|---|---|---|
| condition | No | Condition to check for presumptive status (e.g., "Parkinson's disease", "Hypertension"). | |
| serviceEra | No | Service era or conflict (e.g., "Vietnam", "Gulf War", "Post-9/11") for era-specific presumptives. | |
| exposureType | No | Known toxic/environmental exposure (e.g., "Agent Orange", "burn pits", "Camp Lejeune water"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even with readOnlyHint, idempotentHint, and openWorldHint already provided, the description adds substantial behavioral context: empty results mean the exact combination has no entry, not that the condition is non-presumptive; the legal authority repeats in a provenance block on every row; a date can be null for grouped rows; and a condition description naming a qualifying circumstance is a limit, not a detail. It also discloses that the tool does not read service records, managing expectations for end-to-end eligibility determination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries critical information: the query trigger, return contents, filter constraints, empty-result semantics, and a limitation. It front-loads the core purpose and use case, then layers necessary caveats. There is no filler or redundant restatement of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three optional parameters locked down by the schema and rich annotations, the description is complete: it covers purpose, filter mechanics, required input, empty-result interpretation, provenance/date behavior, and a key limitation about service records. Nothing needed to select and invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents condition, serviceEra, and exposureType. The description goes beyond the schema by explaining that at least one is required, that they combine with AND as a narrowing intersection, and that each filter is refused unless it contains a word of three or more characters. This adds meaningful semantic constraints beyond the property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('check') and resource ('presumptive eligibility') and precisely defines the queries it answers: whether a condition is presumptively service connected, or which conditions are presumptive for an exposure or era. It clearly distinguishes itself from sibling rating and claims tools by focusing on presumptive service connection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this tool 'when a veteran asks whether a condition is presumptively service connected, or which conditions are presumptive for a given exposure or service era.' It also gives important usage constraints: at least one filter is required, filters combine with AND, and every filter needs a word of three or more characters. It does not explicitly name alternatives or state when not to use the tool, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_rating_criteriaCompare rating criteriaARead-onlyIdempotentInspect
Use this when a veteran names a VA diagnostic code and wants the rating criteria for it. Returns the VASRD criteria on file for that code at every tier, highest percentage first, with the condition name, body system and 38 CFR reference, so the findings each tier requires can be read side by side. It takes a diagnostic code rather than a condition name. A code that 38 CFR 4.71a, 4.73 or 4.124a rates in two columns, one for the dominant arm (Major) and one for the non-dominant arm (Minor), returns the regulation's own table and an empty tiers list: schedule is that table as published, with schemaVersion, variant and provenance (the eCFR edition date and the SHA-256 of the source text), every row carrying both published percentages, and scheduleText renders it with its headings, footnotes and the notes whose stated scope includes the code. Notes printed in the same table whose applicability is not established are quoted in a separate labelled section; placement alone does not establish that they apply to this code. For a neuritis or neuralgia code rated on one of those nerve scales, schedule lists each maximum with its qualifying text and ratedOnScale is the scale. The response lists every row and every maximum without choosing one, and when that table cannot be read it carries a message and a link to the eCFR section instead, even with no criteria on file. Other codes with no criteria on file return an empty result. Every response carries a source block: the 38 CFR reference for the code, a link to the current rating schedule at the publisher, and a scheduleSnapshot whose retrievedAt is the date this text was read from eCFR. That date is null for the criteria on file, with a note saying it is not recorded, so the criteria are usable as the schedule on file and the current text governs where the two differ; for a table under schedule it is the date the snapshot was read. It is not a lookup of what a rating pays and it does not combine ratings.
| Name | Required | Description | Default |
|---|---|---|---|
| diagnosticCode | Yes | VA diagnostic code (e.g., "8100" for migraines, "5260" for knee). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, idempotentHint=true, and destructiveHint=false. The description goes far beyond these by detailing edge cases (two-column codes for dominant/non-dominant arms, neuritis/neuralgia scales, no criteria on file), the source block with provenance (eCFR edition date, SHA-256), and the distinction between scheduleSnapshot dates. It also explains when tiers are empty and how notes are handled. This is rich behavioral disclosure that exceeds what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, very long paragraph with dense information. While every sentence carries relevant detail, the lack of structure (no bullets, headings, or paragraph breaks) makes it harder to scan. It is front-loaded with the primary use case, but the extensive edge-case explanations could be organized more clearly. It earns a 3 because it is comprehensive but not concise or well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of an output schema, the description must explain return values and behavior thoroughly. It covers the structure of responses (schedule, tiers, source block, scheduleSnapshot), edge cases (two-column tables, nerve scales, no criteria), and the provenance details. It also clarifies what the tool does not do. This is complete for an agent to invoke correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (the single parameter diagnosticCode is described with examples). The description adds meaning beyond the schema by explicitly stating 'It takes a diagnostic code rather than a condition name,' which is a crucial semantic distinction. It also gives examples (e.g., '8100' for migraines) in the schema, but the description reinforces the code-vs-name point and explains the implications for lookup behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear directive: 'Use this when a veteran names a VA diagnostic code and wants the rating criteria for it.' It specifies the resource (VA diagnostic code) and the action (returns rating criteria). It also differentiates from siblings by explicitly stating 'It is not a lookup of what a rating pays and it does not combine ratings,' which rules out lookup_compensation_rate and calculate_combined_rating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance in the first sentence and adds exclusions: 'It is not a lookup of what a rating pays and it does not combine ratings.' It also clarifies that it takes a diagnostic code rather than a condition name, preventing misuse. Though it doesn't name sibling tools directly, the 'not' statements effectively signal when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compute_retroactive_payEstimate retroactive back payARead-onlyIdempotentInspect
Use this when a veteran gives an effective date and a rating increase and asks what the retroactive amount would be. Returns an estimate of the difference in monthly compensation for each month of the period, with a per-month breakdown, the assumptions behind it, and its own disclaimer. Under 38 CFR 3.31 payment starts on the first day of the month FOLLOWING the month the award became effective, with no exception when the effective date is itself the first of a month, and the period runs through the month of adjudication, which is counted in full. VA rate tables take effect on December 1, so a payment year runs from December 1 through November 30 and December of one year is paid at the next payment year rates. Each month of the breakdown reports the payment year it was paid from (rateYear), that table effective date (rateTableEffectiveDate) and where the table came from (rateSource), and the tables the period drew on are listed once each in rateTables; a published table carries the VA page it was transcribed from, and a reconstructed table carries none, because it is derived rather than transcribed. Monthly rates for payment years 2000 onward are read from VA published rate tables; earlier years are reconstructed from the current published table back-adjusted by the SSA cost-of-living chain, and the result states which of the two the requested period used. A month past the newest published table is served from the newest one available, and the result then reports totalIsExact false with incompleteReason naming those months. The total is an estimate either way, because dependent status is taken as constant across the period. It handles an increase only: the new rating has to be higher than the previous one, and effective dates before 1990 are refused. It does not set an effective date and it does not report an amount VA has authorized.
| Name | Required | Description | Default |
|---|---|---|---|
| isTdiu | No | Whether the new rating includes TDIU (treats as 100% for $). | |
| toRating | Yes | New VA rating percentage (0-100). | |
| hasSpouse | Yes | Whether the veteran was married during the period. | |
| fromRating | Yes | Previous VA rating percentage (0-100). | |
| effectiveDate | Yes | ISO date YYYY-MM-DD of rating change effective date. | |
| dependentCount | Yes | Number of dependent children under 18. | |
| adjudicationDate | No | ISO date the rating decision was issued. Defaults to today. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe, read-only, idempotent, closed-world operation, and the description goes well beyond them: it discloses the estimate's assumptions (dependent status held constant), the completeness flag totalIsExact=false with incompleteReason, how reconstructed vs published rate tables are handled, and the 38 CFR 3.31 payment-start and payment-year rules. This is unusually rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single-sentence trigger is front-loaded and the flow is logical, but the body is dense with regulatory mechanics and per-field output detail (rateYear, rateTableEffectiveDate, rateSource, rateTables) that could be tightened. Nearly every sentence is informative, though the length is near the limit of what an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining returns, and it does: per-month breakdown, assumptions, disclaimer, rate-table provenance, and the exactness flag. Combined with the refusal conditions and rate-source rules, an agent has everything needed to call and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (effectiveDate, fromRating, toRating, hasSpouse, dependentCount, isTdiu, adjudicationDate) is already documented in the schema. The description adds usage-level context for effectiveDate (pre-1990 refused, payment timing) but no new syntax or constraints for the remaining inputs, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('estimate retroactive back pay') and scopes it tightly: a veteran-supplied effective date plus a rating increase. It explicitly distinguishes itself from rate-lookup work by clarifying it 'does not set an effective date and it does not report an amount VA has authorized,' so an agent can route it apart from lookup_compensation_rate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Opens with an explicit trigger ('Use this when a veteran gives an effective date and a rating increase and asks what the retroactive amount would be') and supplies exclusion conditions: it handles increases only, requires the new rating to be higher, and refuses effective dates before 1990. The applicability envelope is fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_claim_helpFind free and accredited claim helpARead-onlyIdempotentInspect
Use this when a veteran asks who can help them with a VA disability claim, or when they describe a crisis. The Veterans Crisis Line comes first, in full, when need is crisis, and no other need returns crisis content. Setting urgent true puts the veteran’s first action at the top of routeNow and does not change the crisis routing. Every result carries accredited representation routes: VA accredited Veterans Service Organizations, whose services on VA benefit claims are free, and accredited attorneys and claims agents, who may charge fees only for representation provided after VA issues notice of an initial decision on the claim (38 CFR 14.636(c) and 38 U.S.C. 5904(c)(1)). It also returns the VA Office of General Counsel accreditation search and the VA.gov find a representative tool, the state or territory veterans agency matching a two-letter state code, VA phone numbers suited to the need, the documents worth bringing, and links to the VA forms that need points at. It names no individual, firm or private directory. It does not prepare, submit or file any form, and it reads no account data, so the answer is the same for every caller with the same inputs. A two-letter code with no agency on file returns stateAgency null and coverage.stateResolved false, with the national routes unchanged. The askedFor argument carries the veteran’s own words. It does not change the routing and does not withhold it: a request that this tool or the model calling it act as the veteran’s attorney, agent or representative before VA, a request to supply words for a veteran or a clinician to use, which covers asking for a personal, buddy or lay statement, a statement in support, or a nexus letter the veteran already has to be rewritten, polished, tightened or cleaned up, to choose a decision review lane for the veteran, or to prepare or file a claim returns status ok with the same findings, plus a boundary object whose kind names the primary boundary crossed and whose kinds lists every boundary the request crossed, in the fixed precedence representation_request, then words_for_testimony, then review_lane_choice, then preparation_request, so a request that crosses two keeps both. Its opening, rule and route state what VeteranHQ does not do, one sentence per kind in kinds, the accreditation rule (38 CFR 14.629) and the fee rule (38 CFR 14.636, 38 U.S.C. 5904(c)(1)), and where the work belongs; those instruments are merged into sources and the same paragraph leads message, behind the Veterans Crisis Line when need is crisis. Provenance is indexed rather than repeated: sources is the registry of full references, each with a stable id, and every guidance line and the boundary carry sourceIds into it. A line’s citation is the sources entry whose id it names, and its passage is line.quote when that field is present and that entry’s quote otherwise. An id is the citation for a regulation or statute (38 CFR 14.636(b)), a short slug for a VA page or form (va-find-accredited-rep, va-form-21-526ez), and product-boundary for the one entry that states what VeteranHQ does and does not do, which carries no url because there is no document to open. A form is cited by its number, with its name in title. One kind of provenance is absent from sources and stays on the line as authority: the state veterans agency page, which is one state’s own contact page and changes with the state argument.
| Name | Required | Description | Default |
|---|---|---|---|
| need | Yes | What the veteran is trying to do. Use general when the veteran has not said, and crisis when they describe a mental health emergency. | |
| state | No | Two-letter US state or territory code, for example CA, TX or PR. Adds that state veterans agency to the answer. Omit it when the veteran has not said where they live. | |
| urgent | No | True when the request is time sensitive, for example a letter, an exam notice or a date the veteran is worried about. It puts their first action at the top of the answer. It does NOT change the crisis routing: use need crisis when the veteran describes a mental health emergency. | |
| askedFor | No | What the veteran asked for, in their own words, when they are available. It changes nothing about the routing and does not withhold it. A request that this tool act as the veteran's representative, attorney or agent before VA, to prepare or file a claim, to choose a decision review lane for the veteran, or to supply words for a veteran or a clinician to use returns the same answer with a boundary object added, naming what VeteranHQ does not do, the accreditation and fee rules and where the work belongs; its kinds array lists every one of those classes the request crossed, in the precedence representation_request, words_for_testimony, review_lane_choice, preparation_request, and kind is the first of them. Ordinary requests for help are answered with no boundary object. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far exceeds the annotations (readOnly/idempotent/closed-world). It discloses determinism ('it reads no account data, so the answer is the same for every caller with the same inputs'), boundary-object behavior for out-of-scope requests with fixed precedence ordering, the null/false edge case for an unmatched state code, the provenance/sourceIds citation model, and the one authority that stays on the line rather than in sources. This is unusually rich behavioral disclosure beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The usage trigger is correctly front-loaded in the first sentence, but the rest is a single enormous run-on paragraph with repeated phrasing ('does not change the routing and does not withhold it' appears twice, echoing the schema). Much of the provenance/citation detail is justified by the absence of an output schema, but the lack of breaks or ordering makes the routing rules hard to extract.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden and does so: it explains the routeNow/message shape, the boundary object and its kinds precedence, sources with stable ids and sourceIds back-references, stateAgency/coverage.stateResolved edge cases, and the authority exception for state agency pages. An agent has enough to interpret both normal and boundary responses without a return schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: urgent controls ordering of routeNow without changing crisis routing; askedFor never changes or withholds routing but conditionally adds a boundary object whose kinds follow a fixed precedence; omitting state suppresses the state agency line, and an unmatched code yields stateAgency null with coverage.stateResolved false. This goes beyond the schema text, though much of it duplicates the parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific function: given a veteran's need, it returns accredited representation routes (VSOs, attorneys/claims agents), the VA OGC accreditation search, the VA.gov find-a-rep tool, state agency, phone numbers, documents and forms, and it handles crisis routing. It also draws explicit self-boundaries ('names no individual, firm or private directory', 'does not prepare, submit or file any form'), which separates it from adjacent tools like search_legal_authority. The core purpose is clear, though the dense enumeration of return payloads partially buries it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Opens with an explicit trigger: 'Use this when a veteran asks who can help them with a VA disability claim, or when they describe a crisis.' It also states a routing rule (crisis content only when need is crisis) and clarifies that urgent does not override crisis routing. What is missing is explicit naming of alternative siblings or when NOT to call this tool versus, e.g., search_legal_authority for the underlying regulations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_secondary_conditionsFind secondary conditionsARead-onlyIdempotentInspect
Use this when a veteran names a service-connected condition and asks what other conditions can follow from it. Returns the records whose own primary condition is the one searched for, each with the strength of the supporting evidence, the medical rationale, key studies, filing notes, the diagnostic codes involved, and the primary condition the record itself names, so a row can be read against the query it answers. Other primary conditions that the search terms reached are summarised separately under relatedPrimaries, with the term that matched and a count, rather than being returned among the rows. Each row carries its provenance: the paragraph of 38 CFR 3.310 the link rests on, the rating schedule section and diagnostic code its evaluation comes from, the eCFR URLs for those sections, and the date the codes and percentages were checked against the rating schedule. That check establishes that a code exists in the schedule and that an evaluation is one the schedule offers for it, which is a question of validity rather than of whether a code fits a particular veteran. The key studies were drafted from published research and have not been verified citation by citation, which every row states in its provenance note. The limit parameter sets how many records one response carries, default 5 and maximum 15, and offset is how many records of the same result are skipped before the first one returned, default 0 and maximum 100000. Records are ordered by strength of evidence and then by their primary and secondary condition names, so the order is the same on every call. A response is also held to a size budget of about 12 KB: when the records within the limit would exceed it, the last of them are left out whole rather than shortened, a response with any record to serve carries at least one, and responseTrimmedForSize is true. totalExactCount is every record on file for the search, count and exactCount are the records in the response, hasMore says whether any records follow them, and nextOffset is the offset that returns the next records and is null when none follow. relatedPrimaries is returned whole on every response. A totalExactCount of 0 means nothing is on file for that search term, not that no link exists. An offset past the last record returns no records, while totalExactCount still counts the records on file. It does not diagnose, it does not establish that a particular veteran's condition is secondary, and it does not supply the medical nexus opinion a secondary claim needs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many secondary-condition records one response carries, 1 to 15, default 5. A response is also held to a size budget of about 12 KB, so it can carry fewer records than the limit, whole ones only; responseTrimmedForSize is then true and nextOffset gives the offset of the first record it left out. | |
| offset | No | How many records of the same result to skip before the first one returned, 0 to 100000, default 0. Records keep the same order on every call, so the nextOffset one response reports returns the records that follow it. | |
| primaryCondition | Yes | Primary service-connected condition to search against (e.g., "PTSD", "Lumbar strain"). Partial matches supported. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world, but the description adds real behavioral context beyond them: the ~12 KB per-response size budget that trims whole records and sets responseTrimmedForSize, deterministic ordering by evidence strength then names, offset-past-end behavior, and the caveat that key-study citations are unverified. This is exactly the extra context the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded correctly, opening with the usage condition, but the single dense paragraph repeats limit/offset defaults and maximums verbatim from the schema and packs many clauses into one block. Given no output schema, much of the length is justified, but the parameter reiteration is not.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining return values and does so thoroughly: row contents, provenance fields, relatedPrimaries summarization, totalExactCount/count/exactCount/hasMore/nextOffset semantics, and the 0-result meaning. An agent has everything needed to interpret and paginate the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates limit (default 5, max 15) and offset (default 0, max 100000) and describes their interaction with the size budget, but that same explanation is already present in the schema's own parameter descriptions, so it adds little beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action and resource: it returns secondary-condition records whose own primary condition matches the query, and it names the distinct handling of relatedPrimaries. An agent can tell this apart from siblings like analyze_rating_gap or search_legal_authority without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It opens with an explicit trigger ('Use this when a veteran names a service-connected condition and asks what other conditions can follow from it') and closes with explicit non-uses (does not diagnose, does not establish a veteran's condition is secondary, does not supply a nexus opinion). Both the when and the when-not are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_guideRead a VeteranHQ guideARead-onlyIdempotentInspect
Use this when a slug from search_guides names the guide a veteran needs and the answer should carry what the guide says. Returns the slug, title and summary, the date the facts in it were last verified against its sources, the public URL of the guide on veteranhq.app, the body split into the sections the guide itself defines, each with a heading and its text, and the sources the guide cites, each with a title and a URL. A slug this library does not hold returns a tool error carrying GUIDE_NOT_FOUND, whose message says to call search_guides and use a slug from its results. A guide longer than one result can carry returns the sections that fit, plus omittedSections counting the sections left out and a note stating that count and the guide URL. A single section larger than the whole budget is served shortened at a paragraph or sentence boundary, with truncatedSection naming its heading and the note repeating it, so a section that was cut is distinguishable from one served whole. Those three fields are absent when the whole guide is returned, so a shortened answer is distinguishable from a complete one. A guide states the rules as VeteranHQ read them from the cited sources on the date it reports, and VA decides a claim. It reads no account data, so the answer is the same for every caller with the same inputs.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | The guide to read, as returned in the slug field by search_guides. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations. It discloses truncation behavior (omittedSections, truncatedSection, note), how to distinguish a shortened answer from a complete one, that the tool reads no account data, and that the guide states rules as read from sources on a verification date. It also explains the GUIDE_NOT_FOUND error behavior. This is rich behavioral context that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense, covering return fields, truncation behavior, error handling, and scope limitations. It is front-loaded with the primary use case and then details edge cases. It could be slightly more concise, but every sentence earns its place by conveying behavior an agent needs to know.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with one parameter, no output schema, and no nested objects, the description is remarkably complete. It covers the return payload, error behavior, truncation semantics, and the tool's scope (no account data). An agent has everything needed to invoke it correctly and interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the single parameter (slug) with 100% coverage, so the baseline is 3. The description adds value by explaining that the slug must come from search_guides results and that an invalid slug produces a GUIDE_NOT_FOUND error. This is meaningful context beyond the schema's field description, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('read') and resource ('a VeteranHQ guide'), and immediately ties it to a sibling ('search_guides') by stating the slug comes from there. It clearly distinguishes this tool from the other siblings, which are about calculations, searches, and legal authority, not reading a guide's content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: when a slug from search_guides names the guide a veteran needs and the answer should carry what the guide says. It also names the alternative (search_guides) and explains the error path when a slug is not found, telling the agent to call search_guides and use a slug from its results. This is explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_compensation_rateLook up monthly compensation rateARead-onlyIdempotentInspect
Use this when a veteran wants the monthly VA disability compensation amount for a given rating percentage and dependent situation. Returns the monthly and annual amounts from the VA compensation rate table in effect, with the rate year and effective date. The rating is read to the nearest 10 percent, and dependents change the amount only at 30 percent and above. Children are counted as children under 18; school-age children 18 to 23, dependent parents and aid and attendance are outside what it models. It gives the published rate for that combination, not what a specific veteran is paid, and it is not a back-pay calculation for a past period.
| Name | Required | Description | Default |
|---|---|---|---|
| rating | Yes | Disability rating percentage (0–100, rounded to nearest 10). | |
| hasSpouse | No | Whether veteran has a spouse. Affects rates at 30%+. Defaults to false. | |
| childrenCount | No | Dependent children under 18. Affects rates at 30%+. Defaults to 0. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description does not contradict them. It adds substantial behavioral context: rating is rounded to nearest 10 percent, dependents affect rates only at 30% and above, children are limited to under 18, and the result is the published table rate rather than an individualized payment or retroactive amount. This goes well beyond the annotation metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary use case and then provides dense, non-redundant details about return values, rounding behavior, dependency rules, and exclusions. Every sentence earns its place; there is no filler or repetition of the title or tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, three parameters, and absence of an output schema, the description is remarkably complete. It explains what is returned (monthly and annual amounts, rate year, effective date), what the model includes and excludes, and how edge cases like dependents and rounding are handled. This is sufficient for an agent to select and invoke the tool correctly without additional documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides full descriptions for all three parameters (100% coverage), and the description reinforces and extends their meaning: it clarifies that the rating is read to the nearest 10 percent, that hasSpouse and childrenCount only matter at 30% and above, and that childrenCount refers to children under 18. This adds semantic value beyond the schema field names and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific use case: 'Use this when a veteran wants the monthly VA disability compensation amount for a given rating percentage and dependent situation.' It clearly identifies the resource (VA compensation rate table), the action (look up monthly and annual amounts), and the key inputs (rating and dependents). It also distinguishes itself from back-pay calculations, separating it from the sibling tool compute_retroactive_pay.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool and what is excluded: it is not a back-pay calculation, not what a specific veteran is paid, and it does not model school-age children 18–23, dependent parents, or aid and attendance. However, it does not name alternative sibling tools that should be used for those excluded cases, so it stops short of full alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_cp_examPrepare for a C&P examARead-onlyIdempotentInspect
Use this when a veteran has a VA compensation and pension examination scheduled for a condition and wants to know what it involves. Returns the examination type and typical length, what the examiner does, the questions examiners commonly ask, the Disability Benefits Questionnaire sections that apply, and records worth bringing. It does not return advice about what to describe, emphasize or omit at the examination. For a condition with a guide on file it also returns the diagnostic code for that condition and where its rating criteria are published, which are what the examiner's findings are scored against; any other condition returns a general exam guide. Educational reference about the examination process. It does not schedule, reschedule or contact VA, and the examiner's findings and VA decide the rating.
| Name | Required | Description | Default |
|---|---|---|---|
| condition | Yes | Condition being examined (e.g., "PTSD", "knee", "sleep apnea", "tinnitus"). | |
| currentRating | No | Current rating if already rated (for increase exams). Optional. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, idempotent, and non-destructive, and the description adds meaningful behavior: it is educational only, returns a general exam guide for conditions without a file guide, includes diagnostic code and rating criteria only when a guide exists, and cannot take actions like scheduling or contacting VA. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the trigger and a compact list of return contents, then adds boundary statements that prevent misuse. Every sentence contributes useful information, and the length is proportionate to the tool's educational scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description enumerates what the tool returns, explains the conditional behavior for guides, and clarifies what it will not do. For a read-only educational reference tool, this provides enough context for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description indirectly reinforces the 'condition' parameter by explaining condition-specific behavior, but it does not add any detail about 'currentRating' beyond the schema's own explanation. The schema carries the parameter documentation burden adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific trigger ('a veteran has a VA compensation and pension examination scheduled') and states exactly what the tool returns: exam type, length, examiner actions, common questions, DBQ sections, and records to bring. It clearly distinguishes itself from sibling tools by focusing on exam preparation rather than ratings, eligibility, or legal authority.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use it ('Use this when a veteran has... scheduled') and provides important exclusions: it does not give advice about what to describe, does not schedule/reschedule/contact VA, and does not decide ratings. It does not name a specific sibling alternative, so the when-not-to-use guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_guidesSearch the VeteranHQ guidesARead-onlyIdempotentInspect
Use this when a veteran asks about a VA topic VeteranHQ has written a guide on. Returns the guides whose text matches the query, highest scoring first, each carrying the guide slug, title, summary and public URL, and no body text. The ordering is lexical and deterministic, computed from where each word of the query appears: a match in the guide title scores above one in the summary, a summary match above one in a section heading, and a heading match above one in the body, with an added score when every word of the query appears somewhere and a further one when the whole query appears in a title. Guides tied on score are ordered by slug, so the same query returns the same list in the same order on every call. No model runs and no network call is made. The limit parameter caps how many guides come back, from 1 to 10, default 5. An empty results array means no guide in this library contains any word of the query, which is a fact about the library and not about VA rules or about what VA holds. Reading one is a second call: pass a returned slug to get_guide. It reads no account data, so the answer is the same for every caller with the same inputs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many guides to return, 1 to 10. Default 5. | |
| query | Yes | What the veteran is asking about, in keywords or a short phrase, for example "C&P exam" or "how do I get my military medical records". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, and the description still adds substantive behavior: no model runs, no network call, deterministic lexical ordering, tie-breaking by slug, and no account data read. It also explains the empty-array semantics precisely, which is exactly the context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the when-to-use sentence and then the return shape, with the long scoring-explanation sentence mid-block. Every sentence carries information, though the scoring algorithm paragraph is dense and slightly over-specified for a selection-time decision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry return semantics — and it does: slug, title, summary, public URL, no body text, ordering guarantees. Combined with the empty-result interpretation and get_guide handoff, nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the limit bounds and default are already documented; the description restates them but adds the interpretive meaning of an empty results array and the keyword/phrase framing of query. It goes beyond the schema without duplicating it much.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (VeteranHQ guides) plus matching scope, and clearly distinguishes itself from get_guide ('Reading one is a second call') and from search_legal_authority by sourcing only its own guide library. An agent can tell what it does and what it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when a veteran asks about a VA topic VeteranHQ has written a guide on'), names the follow-up alternative get_guide with the exact handoff (pass a returned slug), and frames the empty-result case as a library fact rather than a VA rule. When-to-use and what-to-do-next are both explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_legal_authoritySearch case law and regulationsARead-onlyIdempotentInspect
Use this when a question calls for the text of a VA rating regulation, the wording of a legal standard, or case law, searched by keyword, by exact citation, or by diagnostic code. Returns 38 CFR rating criteria for matching conditions; doctrine entries quoted verbatim from the regulation, statute or decision they come from, each with pinpoint citations, the date the text was captured, when that text took effect, and the reviewed holdings that construe it; and excerpts from Court of Appeals for Veterans Claims decisions with docket number, case name and relevance score. CAVC decisions are binding precedent for the Board of Veterans' Appeals. The citation parameter takes a section such as "3.310", "38 CFR 3.310(b)" or "38 U.S.C. 5107(b)", and a citation that cannot be read as a section matches nothing rather than being guessed at. A citation on its own is a complete call: query is optional when citation is supplied, and a call carries query, citation, or both. The limit parameter caps results per source, default 5 and maximum 10; the doctrine block is capped at four entries per response and reports the true total. The status field describes the answer in one of four words. "ok" means every component answered and this response carries each matched result it selected. "partial" means a matched result is missing from it, whether shed for the size budget, withheld as a criterion text too long to fit, or unreachable because a component could not complete its search. "no_match" means every component answered and none of them matched. "unavailable" means a component did not answer, so nothing in that response is readable as an absence of authority. statusReason gives the reason for that word in one sentence, derived from the same values as the word itself. A truncation object reports what was shortened, whatever the status: doctrinePropositionsTrimmed counts entries serving their first propositions while reporting their own true total, textBounded says a criterion text was withheld whole rather than cut, and resultsDropped counts the matched results this response does not carry, under regulatory, caseLaw and doctrine. Shortening that keeps every selected result arrives as status "ok" with truncation populated. Docket numbers and case names come only from the returned results, and when the case-law corpus is empty the response says so and returns the regulatory and doctrine results alone. When a response would exceed its size budget it sheds results rather than overrunning, reports responseTrimmedForSize alongside the true totals, and withholds a criterion text too long to fit whole rather than cutting it: that row carries criteriaOmitted and the length of what was withheld, so a partial rule does not arrive as a complete one. citationResolved says whether any source carries the citation supplied: true on positive evidence, false only where the corpus-wide membership scan completed and nothing named the section, and null where that scan did not complete, which is an open question rather than a finding of absence. citationStatus renders the same answer as resolved, unresolved, unknown, unparsed or not_supplied. The coverage block reports how far the search reached: citationScanComplete for that membership question, candidateScansComplete with incompleteCandidateScans for components whose scan window filled before every candidate was examined, truncatedComponents for the sources whose results were shortened, doctrineMatchedOn for whether the doctrine entries came from the citation, the query or the standing authorities, and searchComplete for the conjunction of all of it. totalsAreMeasuredCounts says whether the totals count everything that matched; where it is false a zero is bounded by this search rather than by the corpus. It does not read an individual claim, it does not predict how a claim will be decided, and it is not legal advice.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 5, max 10). | |
| query | No | Search query (e.g., "PTSD secondary to MST", "DC 8100 migraine 50 percent criteria", "CUE in combined rating calculation"). Optional when citation is supplied; supply query, citation, or both. | |
| citation | No | Exact citation to look up (e.g., "3.310", "38 CFR 3.310(b)", "4.16", "38 U.S.C. 5107(b)"). Sufficient on its own: a call with only a citation is answered as a citation lookup. A citation that cannot be read as a CFR or USC section matches nothing and is reported. | |
| diagnosticCode | No | Filter by diagnostic code (e.g., "8100"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds substantial behavior beyond them: size-budget shedding, whole-criterion withholding vs truncation, the four status values, citationResolved null-vs-false semantics, and the coverage/truncation blocks. This is exactly the extra context annotations cannot supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the usage trigger and the response-contract detail is justified by the absence of an output schema, but the body is roughly 500 words and repeats citation-matching behavior more than once ('matches nothing' / citationResolved / citationStatus / citationScanComplete restate the same fact). A tighter pass could cut length without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a fairly complex multi-source response, the description carries the entire return-value burden and does so thoroughly (status, statusReason, truncation, coverage, citation resolution, and the empty-case-law behavior). An agent has enough to interpret any response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: a bare citation is a complete call and query is optional alongside it, an unparseable citation 'matches nothing rather than being guessed at', and limit is per-source with a hard max of 10. These clarify ambiguity the schema alone leaves open.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('search ... VA rating regulation ... legal standard ... case law') and enumerates the three lookup modes (keyword, exact citation, diagnostic code). It is unmistakably distinct from the sibling tools, which are calculators, eligibility checks and payment tools rather than textual authority lookups.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It opens with an explicit trigger condition ('Use this when a question calls for the text of a VA rating regulation, the wording of a legal standard, or case law') and closes with scope exclusions ('does not read an individual claim, does not predict how a claim will be decided, is not legal advice'). It stops short of naming a sibling tool as the alternative for adjacent needs, but the exclusions effectively route the agent away from non-search tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
find_secondary_conditions2 fields changed- added
Input schema / properties / limitAdded value: +{ + "description": "How many secondary-condition records one response carries, 1 to 15, default 5. A response is also held to a size budget of about 12 KB, so it can carry fewer records than the limit, whole ones only; responseTrimmedForSize is then true and nextOffset gives the offset of the first record it left out.", + "maximum": 15, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "description": "How many records of the same result to skip before the first one returned, 0 to 100000, default 0. Records keep the same order on every call, so the nextOffset one response reports returns the records that follow it.", + "maximum": 100000, + "minimum": 0, + "type": "integer" +}
2 tool updates
- Changed
calculate_rolling_window3 fields changed- changed
Input schema / properties / absencePeriods / items / properties / dayCount / descriptionPrevious value: -"Total CALENDAR days absent in this period, weekends included. Use this when the veteran reports \"20 days between March and December\" instead of exact dates. If provided, startDate/endDate define the outer span and dayCount overrides the calendar days counted within it. Do NOT convert a work-day figure: ask the veteran for calendar days."New value: +"Total CALENDAR days absent in this period, weekends included, for a period reported as a count (for example \"20 days between March and December\") rather than as exact dates. If provided, startDate/endDate define the outer span and dayCount overrides the calendar days counted within it. A work-day total alone does not establish calendar duration: this value is a reported calendar-day total, not an estimate converted from work days." - changed
Input schema / properties / absencePeriods / items / properties / endDate / descriptionPrevious value: -"End date of the absence period in YYYY-MM-DD format. If omitted, assumed same as startDate (single day). Must be on or after startDate: a reversed range is rejected rather than counted, so ask the veteran again instead of guessing the order."New value: +"End date of the absence period in YYYY-MM-DD format. If omitted, assumed same as startDate (single day). Must be on or after startDate; a reversed range is rejected rather than counted." - changed
Input schema / properties / absencePeriods / items / properties / physicianPrescribedBedRest / descriptionPrevious value: -"True only when a physician PRESCRIBED bed rest for this period and treated the veteran, which is what 38 CFR 4.71a Note (1) requires of an incapacitating episode. Optional; when omitted or false the period is reported as a plain absence and contributes NOTHING to the IVDS tier. Never set it true from a work absence, a sick day, or self-directed rest."New value: +"True only when a physician PRESCRIBED bed rest for this period and treated the veteran, which is what 38 CFR 4.71a Note (1) requires of an incapacitating episode. Optional; when omitted or false the period is reported as a plain absence and contributes NOTHING to the IVDS tier. A work absence, a sick day, or self-directed rest without a physician's prescription does not meet that definition."
- Changed
find_claim_help1 field changed- changed
Input schema / properties / askedFor / descriptionPrevious value: -"What the veteran asked for, in their own words, when they are available. It changes nothing about the routing and never withholds it. A request that this tool act as the veteran's representative, attorney or agent before VA, to prepare or file a claim, to choose a decision review lane for the veteran, or to supply words for a veteran or a clinician to use returns the same answer with a boundary object added, naming what VeteranHQ does not do, the accreditation and fee rules and where the work belongs; its kinds array lists every one of those classes the request crossed, in the precedence representation_request, words_for_testimony, review_lane_choice, preparation_request, and kind is the first of them. Ordinary requests for help are answered with no boundary object."New value: +"What the veteran asked for, in their own words, when they are available. It changes nothing about the routing and does not withhold it. A request that this tool act as the veteran's representative, attorney or agent before VA, to prepare or file a claim, to choose a decision review lane for the veteran, or to supply words for a veteran or a clinician to use returns the same answer with a boundary object added, naming what VeteranHQ does not do, the accreditation and fee rules and where the work belongs; its kinds array lists every one of those classes the request crossed, in the precedence representation_request, words_for_testimony, review_lane_choice, preparation_request, and kind is the first of them. Ordinary requests for help are answered with no boundary object."
2 tool updates
- Added
get_guide - Added
search_guides
1 tool update
- Changed
find_claim_help1 field changed- changed
Input schema / properties / askedFor / descriptionPrevious value: -"What the veteran asked for, in their own words, when they are available. It changes nothing about the routing and never withholds it. A request to prepare or file a claim, to choose a decision review lane for the veteran, or to supply words for a veteran or a clinician to use returns the same answer with a boundary object added, naming what VeteranHQ does not do, the accreditation and fee rules and where the work belongs; ordinary requests for help are answered with no boundary object."New value: +"What the veteran asked for, in their own words, when they are available. It changes nothing about the routing and never withholds it. A request that this tool act as the veteran's representative, attorney or agent before VA, to prepare or file a claim, to choose a decision review lane for the veteran, or to supply words for a veteran or a clinician to use returns the same answer with a boundary object added, naming what VeteranHQ does not do, the accreditation and fee rules and where the work belongs; its kinds array lists every one of those classes the request crossed, in the precedence representation_request, words_for_testimony, review_lane_choice, preparation_request, and kind is the first of them. Ordinary requests for help are answered with no boundary object."
1 tool update
- Changed
search_legal_authority3 fields changed- changed
Input schema / properties / citation / descriptionPrevious value: -"Exact citation to look up (e.g., \"3.310\", \"38 CFR 3.310(b)\", \"4.16\", \"38 U.S.C. 5107(b)\"). A citation that cannot be read as a CFR or USC section matches nothing and is reported."New value: +"Exact citation to look up (e.g., \"3.310\", \"38 CFR 3.310(b)\", \"4.16\", \"38 U.S.C. 5107(b)\"). Sufficient on its own: a call with only a citation is answered as a citation lookup. A citation that cannot be read as a CFR or USC section matches nothing and is reported." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query (e.g., \"PTSD secondary to MST\", \"DC 8100 migraine 50 percent criteria\", \"CUE in combined rating calculation\")."New value: +"Search query (e.g., \"PTSD secondary to MST\", \"DC 8100 migraine 50 percent criteria\", \"CUE in combined rating calculation\"). Optional when citation is supplied; supply query, citation, or both." - removed
Input schema / requiredRemoved value: -[ - "query" -]
1 tool update
- Added
find_claim_help
1 tool update
- Changed
search_legal_authority1 field changed- added
Input schema / properties / citationAdded value: +{ + "description": "Exact citation to look up (e.g., \"3.310\", \"38 CFR 3.310(b)\", \"4.16\", \"38 U.S.C. 5107(b)\"). A citation that cannot be read as a CFR or USC section matches nothing and is reported.", + "type": "string" +}
1 tool update
- Changed
calculate_combined_rating5 fields changed- added
Input schema / properties / bilateralAdded value: +{ + "description": "Optional. Tags the ratings that affect a paired extremity, so 38 CFR 4.26 can be applied: those ratings are combined with each other first, 10 percent of that value is added (not combined), and the result enters the 4.25 combination as one disability. The factor needs a compensable rating on both sides of one pair (both arms or both legs); tag every rating that affects an arm or a leg and the tool applies the factor only where the regulation allows it. Tagging more than 16 compensable ratings is accepted and answered, but past that the 38 CFR 4.26(d) comparison across eligible groupings is not exhaustive here, so the result carries the 4.25 combination of every rating supplied under combinedRatingWithoutBilateralFactor instead of a combined rating, and says where the complete calculation comes from.", + "items": { + "additionalProperties": false, + "properties": { + "extremity": { + "description": "The extremity that rating affects.", + "enum": [ + "left-arm", + "right-arm", + "left-leg", + "right-leg" + ], + "type": "string" + }, + "index": { + "description": "Position in the ratings array of the rating this entry describes (0-based).", + "maximum": 19, + "minimum": 0, + "type": "integer" + } + }, + "required": [ + "index", + "extremity" + ], + "type": "object" + }, + "maxItems": 20, + "type": "array" +} - changed
Input schema / properties / ratings / descriptionPrevious value: -"Array of individual disability rating percentages (0–100), e.g., [70, 50]."New value: +"Array of individual disability rating percentages (0-100), e.g., [70, 50]." - added
Input schema / properties / ratings / items / maximumAdded value: +100 - added
Input schema / properties / ratings / items / minimumAdded value: +0 - added
Input schema / properties / ratings / maxItemsAdded value: +20
10 tool updates
- First observed
analyze_rating_gap - First observed
calculate_combined_rating - First observed
calculate_rolling_window - First observed
check_presumptive_eligibility - First observed
compare_rating_criteria - First observed
compute_retroactive_pay - First observed
find_secondary_conditions - First observed
lookup_compensation_rate - First observed
prepare_cp_exam - First observed
search_legal_authority
Related MCP Connectors
900,000+ Board of Veterans' Appeals decisions: VA outcomes, grant rates, PACT Act, ratings.
Real Board of Veterans' Appeals outcome data for VA disability claims. No key needed.
Search 13,000+ US vaccine court (VICP) decisions: cases, court text, statistics, attorneys. Free.
VA facility locations, services, wait times, and satisfaction scores
Related MCP Servers
AlicenseAqualityDmaintenanceDisability insurance quote intake and coverage guidance for high-income professionals. Provides a quote_request action that files a lead with a licensed brokerage, plus read-only tools for specialty guidance, carrier comparison, benefit-cap math, and rider definitions.63MIT- AlicenseNot gradedqualityBmaintenanceEnables querying the Code of Virginia by citation to retrieve state statutes, verifying which citations are valid.275 npmMIT
- AlicenseAqualityCmaintenanceVerified ICD-10-CM code lookup & validation for AI agents — official descriptions, not guesses.395 npm1Apache 2.0
- AlicenseNot gradedqualityDmaintenanceProvides 59 clinical medical calculators and scoring tools for healthcare professionals and AI assistants, covering renal, cardiovascular, pulmonary, critical care, and other specialties.5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.