databutler-vehicles
Server Details
UK MOT failure data and official UK, French, Japanese and EU vehicle recalls.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 9 tools
Each tool is scoped by jurisdiction plus data source (fr/jp/uk recalls, UK MOT, JP complaints), and descriptions give explicit routing guidance (use uk_mot_vehicle_search before uk_mot_reliability; use *_search for cross-brand keyword queries vs *_for_model for structured lookups). The only real fuzziness is eu_safety_gate_recalls being a stub that overlaps with fr_recalls_*, plus two similarly named 'search' tools in different datasets, but descriptions resolve both.
Names follow a predictable region_source_operation pattern throughout (fr_recalls_for_model, jp_recalls_for_model, uk_recalls_for_model, uk_mot_vehicle_search, uk_mot_reliability), all snake_case with no style mixing. The eu_safety_gate_recalls and jp_complaints_summary names deviate slightly in suffix but remain readable and consistent with the region-prefix convention.
Nine tools is well within the ideal range and each earns its place: one jurisdiction gets a structured lookup and a keyword search, plus a dedicated MOT reliability pair. No redundant or filler tools given the multi-region scope.
Coverage is solid for FR/JP/UK recalls, JP complaints and UK MOT, but the EU Safety Gate tool is an explicit non-functional stub and there is no US (NHTSA) surface, no per-recall detail fetch, no VIN-level lookup, and no unified cross-jurisdiction query. These are notable gaps for a general vehicle-safety server that agents will hit on common US or EU questions.
Available Tools
9 toolseu_safety_gate_recallsEU Safety Gate recalls (stub)ARead-onlyIdempotentInspect
STUB — not yet live. EU Safety Gate (RAPEX) vehicle alerts are being added; until then this returns covered:false with the official Safety Gate URL and points to fr_recalls_* (RappelConso republishes the Safety Gate automobile alerts that concern France).
| Name | Required | Description | Default |
|---|---|---|---|
| brand | No | ||
| model | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly/idempotent/non-destructive/openWorld=false), the description discloses the actual runtime behavior an agent cannot infer: the call returns covered:false and an official Safety Gate URL, i.e. it is a deliberate no-op placeholder. That is genuinely useful and consistent with the annotations; no auth, rate-limit or side-effect detail is needed for a stub.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with 'STUB — not yet live', which is the single most decision-relevant fact, followed immediately by the alternative. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stub with no output schema, the description covers the essentials: it is not live, what it returns, and where to go instead. The remaining gap is the unaddressed brand/model parameters and any hint of when the stub will become live, but for a placeholder this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both brand and model are undocumented in the schema (0% coverage) and the description never mentions them, so an agent gets no signal whether they filter, are ignored, or are reserved for the live version. The stub framing mitigates this slightly, but the description does not compensate for the coverage gap as required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool currently is and does: a not-yet-live stub for EU Safety Gate (RAPEX) vehicle alerts that returns covered:false plus the official URL. It names the resource precisely and distinguishes itself from siblings by naming fr_recalls_*. It stops short of 5 only because the eventual retrieval behavior is described prospectively ('are being added') rather than fully defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-not guidance ('not yet live') and routes the agent to an alternative family (fr_recalls_*) with the reason that RappelConso republishes the Safety Gate automobile alerts concerning France. The routing is conditional on the agent needing the data rather than a blanket 'always use X', which keeps it just under a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fr_recalls_for_modelFrench recalls for a modelARead-onlyIdempotentInspect
Official French vehicle recall notices (RappelConso, DGCCRF open data; 1663 automobile fiches 2018-01-19 → 2026-08-07) for a brand and model: fiche number, date, risks, defect description (motif, in French), commercialisation periods and the official fiche link. Omit model to list a brand. Matches by a reviewed alias table plus whole-word text match; each fiche says which.
| Name | Required | Description | Default |
|---|---|---|---|
| brand | Yes | e.g. peugeot, Mercedes-Benz, VW | |
| model | No | e.g. 208, Classe A, Golf (optional: omit to list the brand) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (read-only, idempotent, not open-world), so the description earns credit for adding non-obvious behavior: matching is done via a reviewed alias table plus whole-word text match, and each fiche reports which method applied. It does not mention result caps or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource and scope, then the optional-parameter rule, then matching semantics — a logical order with no filler. The parenthetical date/count clause is dense but earns its place as provenance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields (fiche number, date, risks, motif in French, commercialisation periods, fiche link) and notes the French-language content. Minor gaps remain around result volume/ordering, but an agent can call this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description goes beyond it by explaining the semantic consequence of omitting model and by disclosing the alias/whole-word matching strategy that governs how brand and model strings are interpreted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (recall notices for a brand/model) with the exact data source (RappelConso, DGCCRF), coverage window, and the concrete fields returned. The name/title pair 'recalls_for_model' vs. sibling 'fr_recalls_search' is clearly differentiated by the 'for a brand and model' scoping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear operational rule — 'Omit model to list a brand' — which is the main usage fork for this tool. It does not, however, tell the agent when to prefer this over fr_recalls_search or eu_safety_gate_recalls, so the routing decision against siblings is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fr_recalls_searchFrench recalls searchARead-onlyIdempotentInspect
Keyword search across all RappelConso automobile fiches (brand, model text, label, defect description, risks), newest first, optionally since a date. Text is French and lowercase ("airbag", "incendie", "takata", "batterie haute tension"). Use for cross-brand questions ("recent EV battery fire recalls in France").
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | words that must all appear, e.g. "airbag takata" | |
| since | No | optional ISO date YYYY-MM-DD — only fiches published on/after it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish safe read-only, idempotent, non-destructive behavior, so the description is free to add value: it discloses that results are ordered newest-first, that matching requires all words to appear, and that indexed text is French and lowercase. Pagination/result-count behavior is not mentioned, which is the only real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: scope and searchable fields come first, then ordering and the optional date filter, then the intended use case. Every clause carries information an agent would otherwise have to guess.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only search with no output schema, the description covers scope, fields indexed, language, ordering, date filtering, and intended use. Only result volume/pagination and the returned fiche shape are unaddressed, which is a minor gap given the annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented and the baseline is 3. The description adds genuinely new semantics beyond the schema: the indexed text is French and lowercase, with concrete vocabulary examples ("airbag", "incendie", "takata") that tell the agent what tokens will actually match.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (keyword search) and a precisely scoped resource (all RappelConso automobile fiches, with the exact fields searched), and it distinguishes itself from the model-specific sibling by framing its use case as cross-brand questions. An agent can tell this apart from fr_recalls_for_model or uk_recalls_search without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear positive trigger ("Use for cross-brand questions") with a concrete example, which implicitly routes model-specific queries to fr_recalls_for_model. However, it never explicitly names the alternative tool or states an exclusion, so the routing is inferred rather than declared.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jp_complaints_summaryJapanese owner complaints summaryARead-onlyIdempotentInspect
Owner defect reports filed with MLIT's 不具合情報ホットライン (19659 across 75 models): total reports for a model and the breakdown by defective device (brakes, lights, engine, body…) from the newest sample. Reports are unverified owner submissions — compare shapes, not raw counts. Accepts English or Japanese names.
| Name | Required | Description | Default |
|---|---|---|---|
| make | Yes | e.g. honda | |
| model | Yes | e.g. n-box |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and openWorldHint=false, so safety is covered. The description adds real context beyond that: the upstream source (MLIT hotline), the corpus size (19659 reports across 75 models), the 'newest sample' scoping, and the caveat that submissions are unverified owner reports.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the resource and return shape come first, caveats second. The parentheticals carry real information (dataset size, model coverage, device categories) rather than filler, though the sentence is long enough that a reader must parse three separate ideas in one pass.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with full schema coverage and annotations, the description covers the essentials, including what is returned (total plus device breakdown) despite there being no output schema. Minor gaps remain around pagination or how large the breakdown response can get.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both make/model carry examples, so baseline is 3. The description adds meaning the schema does not: 'Accepts English or Japanese names,' which resolves an ambiguity the lowercase English examples ('honda', 'n-box') do not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource: total owner defect reports for a model plus a breakdown by defective device (brakes, lights, engine, body). It also pins the data source (MLIT's 不具合情報ホットライン) and labels the data as owner-submitted, which cleanly separates it from the recall-oriented siblings like jp_recalls_for_model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives interpretive guidance ('compare shapes, not raw counts') and warns the reports are unverified, which is genuinely useful context for using the output. However, it never says when to pick this tool over jp_recalls_for_model or the other recall/reliability siblings, leaving the alternative-selection decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jp_recalls_for_modelJapanese recalls for a modelARead-onlyIdempotentInspect
Official Japanese recall filings (国土交通省 MLIT; 919 filings matched to 75 popular JDM models) for a make and model: notification number, date, defective device, affected vehicles, situation and remedy (Japanese originals) and the official filing PDF. Accepts English (Toyota Prius, Honda N-BOX) or Japanese (トヨタ プリウス) names.
| Name | Required | Description | Default |
|---|---|---|---|
| make | Yes | e.g. toyota, Honda, スズキ | |
| model | Yes | e.g. prius, N-BOX, ジムニー |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld=false, and the description adds meaningful context beyond them: the data source (MLIT), the curated coverage limit (919 filings across 75 popular JDM models), that returned text is Japanese-original, and that a filing PDF is included. The coverage cap is exactly the kind of constraint an agent needs to judge result completeness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resource and source, then the input-language flexibility. Dense but each clause carries information; the parenthetical statistics are slightly heavy but earn their place as coverage constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing the returned fields (notification number, date, defective device, affected vehicles, situation, remedy, PDF). Minor gaps remain: no indication of result volume, pagination, or behaviour when a make/model has no matched filings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds real value by stating that both English ('Toyota Prius') and Japanese ('トヨタ プリウス') inputs are accepted, which the schema only hints at via examples. It does not clarify matching behavior (e.g. partial vs exact, case handling).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — retrieving official Japanese recall filings for a given make and model — and enumerates the returned fields. The geographic and domain scope ('国土交通省 MLIT') cleanly separates it from the fr_*, uk_* and eu_* recall siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'for a make and model' framing and the JP-specific scope make it obvious when to pick this over fr_recalls_for_model or uk_recalls_for_model. It does not, however, give explicit when-not guidance or mention the related jp_complaints_summary alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uk_mot_reliabilityUK MOT reliabilityARead-onlyIdempotentInspect
Real UK MOT failure data for a used car: per-category failure rates (overall and interpolated at a given mileage) and the riskiest components with typical failure mileages, for one model generation. Live proxy to whatbreaks.uk (same operator; aggregated from DfT anonymised MOT results). Use when a user is buying, comparing or maintaining a car sold in the UK. If several generations match, the result lists candidates — retry with year or the exact model name.
| Name | Required | Description | Default |
|---|---|---|---|
| make | Yes | e.g. Volkswagen | |
| year | No | optional registration year | |
| model | Yes | e.g. Golf — generation resolved via year | |
| mileage | No | optional current mileage in miles |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/openWorld/non-destructive, so the bar is lower, and the description still adds provenance (live proxy to whatbreaks.uk, aggregated from DfT anonymised MOT results) plus the multi-match fallback behavior. It does not mention latency, rate limits, or data recency beyond the source attribution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool returns, then usage context, then the retry hint — three dense sentences with no filler. Slightly packed, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by sketching the return shape (per-category rates, riskiest components, typical failure mileages, candidate list on ambiguous generations). Combined with annotations covering the safety profile, an agent has enough to call it correctly, though response format and pagination are unmentioned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: mileage drives interpolation of failure rates, and year resolves the model generation. That clarifies how the two optional parameters interact with the core result rather than just restating field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: real UK MOT failure data, broken out as per-category failure rates and riskiest components with typical failure mileages, scoped to one model generation. This is clearly distinguishable from sibling tools like uk_mot_vehicle_search and the various recalls tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage context ('buying, comparing or maintaining a car sold in the UK') and a concrete recovery path when several generations match (retry with year or exact model name). It does not, however, state when NOT to use it or name an alternative sibling, so routing between this and uk_mot_vehicle_search is still partly inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uk_mot_vehicle_searchUK MOT vehicle searchARead-onlyIdempotentInspect
Search the 530 UK model generations covered by the MOT dataset by free text (e.g. "focus", "golf"). Use to check coverage or disambiguate a generation before uk_mot_reliability. Live proxy to whatbreaks.uk.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | free-text query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful behavior beyond that: the dataset's fixed scope (530 generations) and the fact that it is a live proxy to whatbreaks.uk, implying an external call with potential latency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the scope and examples are front-loaded, and the downstream-tool relationship comes last where it reads naturally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only search this covers purpose, scope, usage and external dependency adequately, and no output schema means return values needn't be described. It would be stronger if it hinted at what the result contains (generation names/IDs needed to feed uk_mot_reliability).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage the baseline is 3, and the description adds real meaning by giving example queries ("focus", "golf") and clarifying that the free text matches against model generations rather than arbitrary fields. It still doesn't note matching rules (partial match, case sensitivity, aliases).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (530 UK model generations in the MOT dataset) with concrete query examples. It also distinguishes itself from the sibling uk_mot_reliability by framing this as the coverage/disambiguation step that precedes it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use to check coverage or disambiguate a generation before uk_mot_reliability" explicitly gives the condition for choosing this tool and names the downstream tool. It does not spell out when not to use it or contrast it with the other search siblings (e.g. uk_recalls_search), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uk_recalls_for_modelUK recalls for a modelARead-onlyIdempotentInspect
Official UK vehicle safety recalls (DVSA, Open Government Licence; 4876 recalls launched 2016-01-01 → 2027-02-03, 272 makes) for a make and model: recall number, launch date, concern, defect and remedy text (DVSA's English originals), build-date range, vehicles affected and a GOV.UK recall-checker link. Omitting model lists the make's recalls and model names. "Vauxhall Corsa", "vauxhall corsa" and "VAUXHALL CORSA" resolve identically; "Mercedes-Benz", "VW" and "Land Rover" are understood. Covers cars, vans, HGVs, buses, motorcycles, trailers and components, not individual VINs.
| Name | Required | Description | Default |
|---|---|---|---|
| make | Yes | e.g. Vauxhall, Mercedes-Benz, VW, Land Rover | |
| model | No | e.g. Corsa, A Class, Golf (optional: omit to list the make) | |
| since | No | optional ISO date YYYY-MM-DD — only recalls launched on/after it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, yet the description adds substantial extra context: data provenance and licence (DVSA, Open Government Licence), dataset scale and coverage window (4876 recalls, 2016→2027, 272 makes), case-insensitive and alias normalization behaviour, and the explicit exclusion of VIN-level lookups. This is rich behavioural disclosure beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded — the core purpose and return fields come first, then normalization and scope notes. Slightly overpacked with parenthetical detail, but nearly every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned fields and linking them to a GOV.UK checker. Combined with the coverage window and scope exclusions, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it documents normalization ('Vauxhall Corsa', 'vauxhall corsa', 'VAUXHALL CORSA' resolve identically; 'Mercedes-Benz', 'VW', 'Land Rover' understood) and the effect of omitting model. The 'since' parameter is only indirectly covered via the launch-date window.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Official UK vehicle safety recalls (DVSA...)' for a make and model, and enumerates the exact fields returned (recall number, launch date, concern, defect, remedy, build-date range, vehicles affected, GOV.UK link). It is immediately distinguishable from sibling tools like uk_mot_reliability and uk_mot_vehicle_search by domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance: 'Omitting model lists the make's recalls and model names' and bounds the tool's scope ('not individual VINs'). However, it never names or contrasts the closest sibling, uk_recalls_search, so the agent must infer which recall tool to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uk_recalls_searchUK recalls searchARead-onlyIdempotentInspect
Keyword search across all served DVSA recalls (make, model names, concern, defect and remedy text), newest first, optionally since a date. Text is English ("airbag", "fire", "steering rack", "takata", "high voltage battery"); every word must appear. Suited to cross-make questions such as recent EV battery fire recalls in the UK.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | words that must all appear, e.g. "airbag takata" | |
| since | No | optional ISO date YYYY-MM-DD — only recalls launched on/after it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds real operational context the annotations lack: results are ordered newest first, matching is conjunctive (every word must appear), and the corpus is English-language DVSA text. Pagination/result limits are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core behavior before the examples. The example keywords (
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter search tool with no output schema, the description covers scope, matching semantics, language, ordering and a usage scenario. It omits result-count limits, pagination and how matches on different fields are weighted, which an agent would want before relying on it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented ('words that must all appear', 'only recalls launched on/after it'). The description reinforces the AND semantics and date scoping but adds no syntax beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (keyword search) over a named resource (all served DVSA recalls) and enumerates the searched fields (make, model, concern, defect, remedy). The 'all served / cross-make' framing implicitly distinguishes it from the sibling uk_recalls_for_model, which is model-scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use case ('cross-make questions such as recent EV battery fire recalls in the UK'), which tells the agent when this tool fits. It does not explicitly name uk_recalls_for_model as the narrower alternative or state exclusions, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
- First observed
eu_safety_gate_recalls - First observed
fr_recalls_for_model - First observed
fr_recalls_search - First observed
jp_complaints_summary - First observed
jp_recalls_for_model - First observed
uk_mot_reliability - First observed
uk_mot_vehicle_search - First observed
uk_recalls_for_model - First observed
uk_recalls_search
Related MCP Connectors
Real UK MOT failure data by car model and mileage, aggregated from official DfT test results.
UK vehicle MOT status and full MOT history from official DVSA data, by registration.
UK used cars: road tax (VED), ULEZ charges, MOT dates, DVSA reliability, live dealer stock.
US vehicle recalls, complaints, EPA figures, VIN decode and trouble codes, with sources.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides comprehensive vehicle reports by aggregating data from multiple public sources to decode VINs, check recalls, and view safety ratings. It enables users to validate VINs locally and retrieve technical specifications, fuel economy, and vehicle photos without requiring API keys.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables natural-language queries over U.S. vehicle records, including recalls, service bulletins, diagnostic trouble codes, VIN decoding, fuel economy, and crash ratings.1MIT
- FlicenseNot gradedqualityBmaintenanceSearch, localized specs (180 spec types across 19 categories), compare, and structured filters over 102k+ vehicle variants in 19 languages, from cars-data.com.-
- AlicenseAqualityAmaintenanceDecode VINs, look up specs, history, recalls, market value, and OBD codes. Recognize license plates and VINs from images. Access comprehensive vehicle data by year, make, and model to power automotive workflows.12MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.