databutler
Server Details
Verified facts for AI agents: rates, exams, recalls, holidays, and what changed since your cutoff.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 9 tools
Each listed tool has a distinct high-level purpose (catalogue, coverage, provenance, specific lookups, changelog). The catch-all `call` tool overlaps conceptually with the directly listed lookup tools, but its description clarifies its role for the long tail; this is a minor ambiguity.
All tool names use snake_case, but the naming patterns are mixed: bare verbs (`call`), nouns (`catalogue`, `coverage`), noun phrases (`exam_topic_paper`, `domain_provenance`), and verb phrases (`what_changed_since`). Still readable but not a consistent verb_noun convention.
Nine tools is a well-scoped front-door surface for a multi-domain data service. Each tool earns its place by covering a distinct meta or lookup role, and the catch-all `call` avoids needing dozens of direct tools.
The combination of `catalogue`, `coverage`, `call`, and `what_changed_since` gives access to the full underlying tool surface and changelog, so no major operations are missing. However, many common lookups (e.g. US rates, software EOL) require the two-step catalogue→call pattern rather than direct tools, a minor friction.
Available Tools
9 toolscallCall any toolARead-onlyIdempotentInspect
Run any Data Butler tool by name — the long tail not listed here: exam_spec_lookup, exam_paper_index; exact statistics (distribution, hypothesis_test, confidence_interval, bayes_update, linear_regression, descriptive_stats); the vehicles hub (uk_mot_reliability, uk_recalls_for_model, uk_recalls_search, fr_recalls_for_model, fr_recalls_search, jp_recalls_for_model, jp_complaints_summary); software end-of-life (support_status {product, version}, support_timeline {product} for nodejs, python, ubuntu, debian, android, ios, macos, windows, oracle-jdk, go, ruby, php, postgresql, react); policy_rate (Bank of England Bank Rate, Fed federal funds target range, ECB key rates — committed data verified against the banks' own pages); us_rates_lookup (US federal rates and thresholds for the current tax year: income-tax, social-security, medicare, retirement, hsa, estate-gift); de_rates_lookup (German tax and social-insurance figures for the current Veranlagungszeitraum: income-tax, social-insurance, minimum-wage, minijob, kindergeld, kinderfreibetrag, sparer-pauschbetrag, arbeitnehmer-pauschbetrag, entfernungspauschale); uk_vehicle_rules_lookup (UK car tax: first-year CO2 bands, standard rate, expensive car supplement, 2001–2017 bands, pre-2001 rates, electric cars, historic vehicles; MOT due dates, fees, retests, penalties — verified against gov.uk); UK exam dates (uk_exam_dates {kind: results|timetable|deadlines|spec-changes, board?, qualification?}); calendar facts (calendar_facts {country, kind: holidays|tax-year|dst, year?, region?} and next_holiday {country, from?} — holidays for uk, us, de, fr; tax years and DST for uk, us, de, fr, au, in, ca, br, pl; 2026–2028); uk_gov_process (UK government services — passport and ETA fees, driving licence renewal, SORN, register to vote, birth/death registration, Self Assessment deadlines, NI number, Universal Credit, EU Settlement Scheme, eVisa — fees, processing times, start URLs, dated changes, verified against gov.uk). uk_rates_lookup, domain_provenance, package_provenance and take_home_pay are listed directly. See catalogue for schemas. The result carries the inner tool's freshness and cite.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | tool name from catalogue | |
| arguments | No | that tool's arguments |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the result 'carries the inner tool's freshness and cite,' and several families are flagged as 'committed data verified against the banks' own pages' or 'verified against gov.uk,' which tells the agent about data provenance and trustworthiness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the key caveat are correctly front-loaded in the first clause, but the body is a long unstructured run-on enumeration mixing tool names, parameter snippets, supported countries, and provenance notes. For a dispatcher the list has value, yet it is far from tight and buries useful groupings in a wall of text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates partially by stating the result carries the inner tool's freshness and cite, and by pointing to the catalogue for per-tool schemas. It does not explain how nested 'arguments' should be shaped or what happens for an invalid tool name, which keeps it short of fully complete for a 2-param dispatcher with a nested object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description goes further by effectively enumerating valid values for the required 'tool' parameter (the long-tail names, the support_status/support_timeline products, the calendar_facts kinds, etc.), which the schema's generic 'tool name from catalogue' does not. It still defers argument schemas to the catalogue rather than describing them here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause 'Run any Data Butler tool by name' gives a specific verb (run) plus resource (any Data Butler tool) and immediately distinguishes this dispatcher from the sibling 'catalogue' by framing itself as the execution path for 'the long tail not listed here.' It even names which siblings are listed directly (uk_rates_lookup, domain_provenance, package_provenance, take_home_pay), so an agent can separate them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly scopes usage to tools 'not listed here' and explicitly notes the four tools that are listed directly, which implies those should be called on their own rather than through this dispatcher. What is missing is an explicit 'do not use this for X' statement or error-handling guidance for an unknown tool name, so the routing is strong but inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
catalogueTool catalogueBRead-onlyIdempotentInspect
Every tool across all Data Butler servers (UK rates; UK exams; provenance; exact statistics; vehicles: UK MOT, UK, French and Japanese recalls; take-home pay for ten countries; software support status and end-of-life timelines for 14 runtimes and operating systems; central-bank policy rates: Bank of England, Fed, ECB; US federal rates and thresholds (us_rates_lookup: income tax, Social Security, Medicare, retirement, HSA, estate and gift); German tax and social-insurance figures (de_rates_lookup: Einkommensteuertarif, Sozialversicherung, Mindestlohn, Minijob, Kindergeld, Kinderfreibetrag, Pauschbeträge, Entfernungspauschale); UK vehicle tax (VED) and MOT rules; UK exam dates and spec changes: results days, the summer 2027 timetable, entry deadlines, new specifications; calendar facts: public holidays, tax-year starts and daylight-saving switches for 2026–2028; UK government services: passport, driving licence, vehicle tax/SORN, voting, birth/death registration, Blue Badge, Self Assessment, NI number, Universal Credit, EU Settlement Scheme, eVisa, ETA), with input schemas — including tools not listed at this front door. Optional query filters by keyword. Run any of them with call, or connect to serverUrl directly.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | keyword filter, e.g. regression, exam, npm |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds two useful facts beyond the annotations: the catalogue is exhaustive ('including tools not listed at this front door') and entries carry input schemas. It says nothing about result size, ordering, or whether the keyword matches names or descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but it is followed by a single sprawling parenthetical enumerating every server, dataset and sub-tool the catalogue returns — content the tool itself will hand back on invocation. This is padding, not orientation, and it buries the one actionable sentence at the very end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-param, read-only discovery tool with no output schema, the description conveys the essentials: it covers all servers, it is exhaustive beyond the front door, entries include input schemas, and results feed into call. Missing only the return shape (fields, ordering) and the filter's matching rule.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema itself gives an example ('regression, exam, npm'). The description only restates 'Optional query filters by keyword,' adding no matching semantics (name vs description vs server) or behavior on zero matches. Baseline 3 is correct when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause states a specific resource and scope: 'Every tool across all Data Butler servers ... with input schemas.' That is a clear discovery/catalogue purpose, and the closing line ('Run any of them with call') implicitly separates it from the execution sibling call. The mass of enumerated content dilutes the signal but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'Run any of them with call, or connect to serverUrl directly' tells the agent what to do with the results, not when to reach for this tool versus coverage, what_changed_since, or the domain-specific lookups. No exclusions or alternatives are named for the discovery step itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
coverageCoverageARead-onlyIdempotentInspect
Exactly what Data Butler covers and does not, per area — call this before assuming something is unsupported. Areas: uk-rates, uk-exams, provenance, stats, vehicles, tax, software-eol, policy-rates, us-rates, de-rates, uk-vehicle-rules, uk-exam-dates, calendar, uk-gov-process (omit for all). Returns rate categories and tax years, boards/subjects, package ecosystems, stats operations, vehicle datasets, take-home-pay countries, the 14 software products with end-of-life data, central-bank policy rates (boe, fed, ecb), US federal rate categories and tax year, German rate categories and Veranlagungszeitraum, UK vehicle tax (VED) and MOT rule categories, UK exam dates (results, timetable, deadlines, spec-changes for jcq, aqa, pearson, ocr, wjec-eduqas; sqa results day), calendar facts: public holidays (uk, us, de, fr), tax years and daylight saving (uk, us, de, fr, au, in, ca, br, pl) for 2026–2028, UK government services (passports, DVLA, voting, registrations, Self Assessment, benefits, eVisa, ETA), and an explicit notCovered list.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | one of uk-rates, uk-exams, provenance, stats, vehicles, tax, software-eol, policy-rates, us-rates, de-rates, uk-vehicle-rules, uk-exam-dates, calendar, uk-gov-process |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds genuine behavioral value beyond that: the response includes an explicit notCovered list, and it enumerates the actual coverage scope (tax years, country sets, the 14 EOL products, etc.), which tells the agent what to expect back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the "call before assuming" instruction are correctly front-loaded, but the description then unfolds into a single enormous run-on sentence enumerating the return contents of every area. Much of that enumeration restates what the tool itself will return, which is bloat rather than necessary pre-call information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of describing returns, and it does so thoroughly: area names, the shape of the coverage data, and the presence of a notCovered list. An agent has everything needed to decide whether to call it and with which area.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum values are already listed in the schema, so baseline is 3. The description goes further by mapping each area name to its contents (e.g. calendar = holidays/tax years/DST for named countries, uk-gov-process = passports/DVLA/voting), which adds meaning the bare enum does not carry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific, unusual verb+resource: it reports exactly what Data Butler covers and does not, per area, which an agent can distinguish from data-fetching siblings like uk_rates_lookup or take_home_pay. It does not explicitly name or contrast with the closest sibling, catalogue, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"call this before assuming something is unsupported" gives explicit, actionable timing for invocation, and "omit for all" clarifies the no-argument case. There is no stated when-not condition or comparison against the catalogue sibling, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_provenanceDomain provenanceARead-onlyIdempotentInspect
Who runs this domain? Live RDAP lookup: registration date and age, registrar, nameservers, status, whether the registrant is redacted, plus a typosquat check (edit-distance / brand-substring) against high-value brands. A very recently registered domain resembling a bank or big brand is a classic phishing signal — but report it as a signal, not a conclusion.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | e.g. example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent and open-world behavior, so the description is not carrying the safety burden. It adds real value beyond them by disclosing the returned data set (registration age, registrar, nameservers, status, registrant redaction) and the edit-distance/brand-substring typosquat logic, plus an interpretation caveat. It omits failure behavior (e.g. no RDAP record, rate limiting).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core question, then mechanism, then outputs, then the interpretive caveat. Three sentences, each carrying distinct information with no repetition of the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating what the lookup returns, and the single parameter is trivial. For a one-argument read-only lookup, an agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single required 'domain' parameter with 100% schema description coverage, so the schema already documents the input. The description adds no format or normalization guidance beyond the schema's 'example.com' example — baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific question ('Who runs this domain?') and a concrete mechanism ('Live RDAP lookup'), then enumerates the resource attributes returned (registration date/age, registrar, nameservers, status, redaction) plus a typosquat check. The domain-vs-package distinction from the sibling package_provenance is obvious from the resource itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context — recently registered domains resembling banks or big brands as a phishing signal — and instructs how to report the result ('as a signal, not a conclusion'). It does not name alternatives or state exclusions, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exam_topic_paperExam topic → paperARead-onlyIdempotentInspect
Which exam paper is a topic on? The canonical student revision question ("is electricity paper 2 AQA?" — answer: no, Paper 1). Input a topic plus optional board/subject/level; returns the paper, sibling topics on that paper, spec code, and the official spec URL. Knows quirks like AQA Combined Science Trilogy Physics Paper 2 having no Space physics. Coverage: AQA GCSE Physics (8463); AQA GCSE Biology (8461); AQA GCSE Chemistry (8462); AQA GCSE Combined Science: Trilogy (8464); AQA GCSE Mathematics (8300).
| Name | Required | Description | Default |
|---|---|---|---|
| board | No | e.g. aqa | |
| level | No | e.g. gcse | |
| topic | Yes | e.g. electricity, forces, ecology, organic chemistry | |
| subject | No | e.g. physics, combined science |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, closed-world behavior, so the bar is lower. The description adds genuine value beyond that: it discloses known data quirks (AQA Combined Science Trilogy Physics Paper 2 has no Space physics) and an explicit board/subject coverage boundary, which an agent needs to avoid over-trusting results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core answer is front-loaded as a one-line question, then the return shape, then the coverage list. Every sentence earns its place except the coverage enumeration, which is dense but genuinely useful for scoping. Slightly long, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming exactly what is returned (paper, sibling topics, spec code, official spec URL) and disclosing board/subject coverage limits. For a 4-parameter lookup tool with full schema coverage and read-only annotations, this is nearly complete; only ambiguity handling (e.g. unknown topics, wrong-board inputs) is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and all four parameters carry their own examples in the schema, so the baseline is 3. The description restates the same fields ("topic plus optional board/subject/level") without adding syntax, casing normalization, or matching behavior (e.g. whether 'aqa' vs 'AQA' matters), so it adds no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete question-answer framing ("Which exam paper is a topic on?") and enumerates the returned fields (paper, sibling topics, spec code, spec URL), so the resource and output are unambiguous. It does not explicitly distinguish itself from siblings exam_paper_index or exam_spec_lookup, whose names suggest overlapping territory (spec codes/URLs), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a canonical call example ("is electricity paper 2 AQA?" → no, Paper 1) and states the input shape (topic plus optional board/subject/level), which makes the usage context clear. The explicit coverage list also tells the agent when the tool applies and when it will not. No alternative tool or exclusion condition is named, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
package_provenancePackage provenanceARead-onlyIdempotentInspect
Who publishes this package, and is it a typosquat? Live npm or PyPI lookup: first-publish date and age, release count, latest version, maintainers/author, linked repo, plus a typosquat check against popular package names. Use when an agent is about to install or recommend an unfamiliar dependency.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | package name, e.g. express or requests | |
| ecosystem | Yes | npm or pypi |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: this is a live network lookup (freshness/latency implication) and it enumerates the fields returned, including a typosquat heuristic. It does not mention failure modes for unknown packages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the two questions an agent actually has, then the data returned, then the usage condition. Dense but every clause carries information; the interrogative opener is slightly less efficient than a declarative verb phrase.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned fields (first-publish date/age, release count, latest version, maintainers, linked repo, typosquat result) and the scenario that warrants a call. Nothing an agent needs to decide or invoke is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two parameters, so the schema already documents name and the npm/pypi enum. The description reinforces the ecosystem scope but adds no syntax, format, or edge-case detail beyond it; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (publishers of a package, typosquat status) plus the exact lookup source (live npm or PyPI). It is clearly distinguishable from the sibling domain_provenance because it names the artifact type it operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use when an agent is about to install or recommend an unfamiliar dependency" gives a concrete triggering condition. There are no explicit exclusions or alternatives named, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_home_payTake-home payARead-onlyIdempotentInspect
Net salary after income tax and social contributions for an employee in one of ten countries (uk, us, ca, au, in, br, pl, fr, de, es), computed live by the LifeByNumbers salary calculators (same operator). Takes a gross salary in local currency (annual by default, or monthly with period) and a country code, plus optional state and filingStatus for the US, province for Canada, region (england|scotland) for the UK. Returns gross, net and total deductions (each annual and monthly), a line-by-line breakdown typed as tax / social / contribution / repayment, effectiveRate, the tax year the figures apply to, currency, the assumptions behind the estimate and the upstream calculator page as source. Spain (es) is listed but answers covered:false because lifebynumbers.net computes it client-side only.
| Name | Required | Description | Default |
|---|---|---|---|
| gross | Yes | gross salary in local currency, annual unless period is monthly, e.g. 50000 | |
| state | No | US only: two-letter state code, e.g. TX (default CA) | |
| period | No | period of gross (default annual); results always give both annual and monthly | |
| region | No | UK only: scotland applies Scottish income tax bands (default england) | |
| country | Yes | ISO 3166-1 alpha-2, lower case: uk, us, ca, au, in, br, pl, fr, de, es | |
| province | No | Canada only: two-letter province/territory code, e.g. BC (default ON) | |
| filingStatus | No | US only (default single) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, open-world), so the bar is lower, yet the description still adds real value: live computation by the same-operator calculators, the upstream page as source, the returned assumptions, and the critical limitation that Spain (es) answers covered:false. It stops short of describing latency, failure behavior, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and supported countries are front-loaded, then parameters, then return shape, with the Spain caveat last. It is a dense, somewhat run-on single paragraph, but almost every clause carries information and nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden well, enumerating gross/net/deductions, annual+monthly figures, the breakdown typing (tax/social/contribution/repayment), effectiveRate, tax year, currency, assumptions, and source. The es covered:false caveat closes the main coverage ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all seven parameters, defaults, and enums; the description largely restates these (gross annual-by-default, US state/filingStatus, CA province, UK region). Baseline 3 is correct because the schema does the heavy lifting and the description adds little param-level detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Net salary after income tax and social contributions') and scopes it precisely with the ten supported country codes and the source calculators. An agent can immediately distinguish this from a sibling like uk_rates_lookup (rate lookup) versus a full take-home calculation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (need net pay given a gross salary) and documents period/country options, but never names an alternative tool or states when NOT to use it. No explicit routing guidance against the sibling tools, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uk_rates_lookupUK rates lookupARead-onlyIdempotentInspect
Current UK (2026-27) rates and thresholds, verified against gov.uk with source URLs and verification dates. Categories: income-tax, national-insurance, minimum-wage, student-loans, vat, isa, state-pension, state-pension-age, stamp-duty, capital-gains, pensions, dividends, savings, marriage-allowance, statutory-sick-pay, statutory-parental-pay, child-benefit. Use whenever a user asks about UK income tax, National Insurance, minimum wage, student loans, VAT, ISAs, pensions, stamp duty, capital gains, dividends, savings, the State Pension or State Pension age, statutory sick/maternity/paternity pay or Child Benefit — models reliably misremember these numbers.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | one of: income-tax, national-insurance, minimum-wage, student-loans, vat, isa, state-pension, state-pension-age, stamp-duty, capital-gains, pensions, dividends, savings, marriage-allowance, statutory-sick-pay, statutory-parental-pay, child-benefit (omit to list all) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, closed-world behavior, so the description's real contribution is data provenance: verified against gov.uk, with source URLs, verification dates and a stated tax year (2026-27). That freshness and sourcing context is genuinely beyond the structured fields. It stops short of describing result shape or staleness handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool returns and its verification status, followed by the trigger conditions. The one inefficiency is that the 17-category enumeration duplicates the schema's enum verbatim rather than just summarizing it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only lookup with no output schema, the description carries the needed burden: it says what the data is, what version it is, that it is gov.uk-verified with URLs and dates, and when to reach for it. Nothing required to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single category parameter already enumerates all valid values with an 'omit to list all' default. The description repeats the same category list without adding format or behavioral detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (current UK rates and thresholds) plus the payload returned: source URLs and verification dates. The category enumeration and the 'models reliably misremember these numbers' framing make it clearly distinct from computational siblings like take_home_pay.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit and unusually thorough trigger list ('Use whenever a user asks about UK income tax, National Insurance...'). It does not, however, name when-not to use it or point to an alternative such as take_home_pay for net-pay calculations, so the routing guidance is one-sided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_changed_sinceWhat changed sinceARead-onlyIdempotentInspect
What has changed in Data Butler's verified datasets since a date — pass your training-cutoff date to learn which UK tax/NI/benefit thresholds, exam-spec facts, official vehicle datasets (UK DVSA, French RappelConso and Japanese MLIT recalls), software end-of-life dates (Node.js, Python, Ubuntu, Debian, Android, iOS, macOS, Windows, Java, Go, Ruby, PHP, PostgreSQL, React — cycles that reached end of life or whose dates moved), central-bank policy rates (Bank of England, Fed, ECB), US federal tax thresholds (IRS inflation adjustments, Social Security wage base, retirement and HSA limits), German tax and social-insurance thresholds (Grundfreibetrag, Beitragsbemessungsgrenzen, Zusatzbeitrag, Mindestlohn, Kindergeld), UK vehicle tax (VED) rates and MOT rules, UK exam dates (results days, summer timetable paper dates, entry deadlines and late-fee dates, announced specification changes), calendar facts (public holidays for the UK, US, Germany and France; tax-year and daylight-saving rules for nine countries — new years published, holidays added or moved) and UK government service fees and rules (passports, ETA, Universal Credit, eVisa) changed after your knowledge ends. Each entry: from → to, effectiveFrom, official source, verifiedDate, and a maintainer note (e.g. which Budget). Areas: uk-rates, uk-exams, vehicles, software-eol, policy-rates, us-rates, de-rates, uk-vehicle-rules, uk-exam-dates, calendar, uk-gov-process, or all. Coverage starts 2026-08-31. The same changelog is published for humans at https://databutler.dev/changes, with feeds at https://databutler.dev/changes.rss and https://databutler.dev/changes.atom. Output is a Verified Changes protocol document (verified-changes/0.1, spec https://databutler.dev/protocol).
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | default all | |
| since | Yes | ISO date YYYY-MM-DD, e.g. your training cutoff |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds substantial context on top: the coverage boundary ('Coverage starts 2026-08-31'), the shape of each returned entry (from → to, effectiveFrom, official source, verifiedDate, maintainer note), and that output is a versioned verified-changes/0.1 protocol document. These are exactly the constraints an agent needs and cannot get from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded correctly, but the body is a single ~250-word run-on sentence that re-lists all twelve enum values and then dumps an exhaustive catalogue of every domain, dataset and jurisdiction covered. Much of that enumeration duplicates the schema enum and belongs in docs, not the tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so — entry field breakdown, protocol identifier, spec URL, and the human-facing changelog/feed URLs. Combined with the disclosed coverage start date, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, which sets a baseline of 3, and the description usefully reinforces the semantic intent of `since` as 'your training-cutoff date' rather than an arbitrary interval. The area list is repeated verbatim from the enum and adds no new meaning, but the `since` guidance lifts it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause states a specific verb and resource — 'what has changed in Data Butler's verified datasets since a date' — and immediately scopes it to post-cutoff knowledge. It also names the concrete data domains, which distinguishes it cleanly from point-lookup siblings like uk_rates_lookup and take_home_pay.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear triggering condition — 'pass your training-cutoff date to learn which ... thresholds ... changed after your knowledge ends' — which tells the agent exactly when this tool applies. It does not, however, name sibling alternatives or state when NOT to use it (e.g., use uk_rates_lookup for a single current value), so routing is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- Changed
coverage1 field changed- changed
Input schema / properties / area / descriptionPrevious value: -"one of uk-rates, uk-exams, provenance, stats, vehicles, tax, software-eol, policy-rates, us-rates, de-rates, uk-vehicle-rules, uk-exam-dates, calendar"New value: +"one of uk-rates, uk-exams, provenance, stats, vehicles, tax, software-eol, policy-rates, us-rates, de-rates, uk-vehicle-rules, uk-exam-dates, calendar, uk-gov-process"
- Changed
uk_rates_lookup1 field changed- changed
Input schema / properties / category / descriptionPrevious value: -"one of: income-tax, national-insurance, minimum-wage, student-loans, vat, isa, state-pension, stamp-duty, capital-gains, pensions (omit to list all)"New value: +"one of: income-tax, national-insurance, minimum-wage, student-loans, vat, isa, state-pension, state-pension-age, stamp-duty, capital-gains, pensions, dividends, savings, marriage-allowance, statutory-sick-pay, statutory-parental-pay, child-benefit (omit to list all)"
- Changed
what_changed_since1 field changed- changed
Input schema / properties / area / enumPrevious value: -[ - "uk-rates", - "uk-exams", - "vehicles", - "software-eol", - "policy-rates", - "us-rates", - "de-rates", - "uk-vehicle-rules", - "uk-exam-dates", - "calendar", - "all" -]New value: +[ + "uk-rates", + "uk-exams", + "vehicles", + "software-eol", + "policy-rates", + "us-rates", + "de-rates", + "uk-vehicle-rules", + "uk-exam-dates", + "calendar", + "uk-gov-process", + "all" +]
9 tool updates
- First observed
call - First observed
catalogue - First observed
coverage - First observed
domain_provenance - First observed
exam_topic_paper - First observed
package_provenance - First observed
take_home_pay - First observed
uk_rates_lookup - First observed
what_changed_since
Related MCP Connectors
Cited US recall lookup for AI agents: CPSC and FDA data, nothing invented.
- TalarionOAuthcom.talarion
A knowledge base of verified, dated facts your LLM would otherwise get wrong.
Free fact-checks, papers, source vetting, plus verified AI pricing, comparisons, guides, and tools.
Live AI data for agents: model releases, regulations (EU AI Act), GenAI glossary, daily news.
Related MCP Servers
AlicenseAqualityDmaintenanceThe deterministic fact-verification layer for AI agents. Validates the structured facts an agent emits — IBANs, payment cards, VAT and national tax IDs, crypto and bank addresses, domains, emails, phone numbers, securities and academic identifiers, plus dates, currencies and holidays — against checksums and curated authoritative data, not guesses.561Apache 2.0- AlicenseNot gradedqualityDmaintenanceVerified knowledge base for AI agents. Stop hallucinations with certified, source-backed facts. Covers Swiss law, health, finance, climate, AI/ML, and more. 8 tools, no API key needed, public and free.MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to retrieve and compare verified country facts such as VAT/GST, public holidays, visas, tipping, and minimum wages across 70 countries, with automatic x402 micropayments in USDC on Base per call.96 npmMIT
- FlicenseNot gradedqualityDmaintenanceReturns verified financial data (SEC EDGAR & FRED) with machine-readable citations. Guaranteed zero hallucinations for AI agents.-
Glama MCP Gateway
Add one secure layer between your agents and this server.