ghed-mcp
This server gives AI assistants direct MCP access to WHO Global Health Expenditure Database (GHED) data for health-financing research, comparison, and analysis.
Fetch headline health expenditure indicators (e.g., CHE, GGHE-D, OOP) and detailed SHA 2011 codebook variables.
Build country health-financing profiles by country and year, using the latest available values.
Compare countries or country groups (LAC, OECD, LDC, SSA, WHO regions, World Bank income groups) on one or more indicators.
Analyze trends, rank countries by absolute change, percent change, or CAGR, and compute trend summaries with period guards.
Summarize country-group statistics: medians, top/bottom countries, coverage, and mixed-year warnings.
Search and explore indicators, variables, methodology, curated topics, and research use cases.
Resolve country names, aliases, ISO3 codes, regions, income labels, and curated group memberships.
Build tidy long research panels or export-ready research packages with data, codebook, availability CSV, and README.
Validate accounting decompositions with additive hierarchies and balance checks (e.g., CHE by financing scheme).
Assess data availability and quality, including metadata completeness, data-type mix, and caution flags.
Export results as CSV for use in pandas, Excel, R, Stata, or visualization tools.
Manage local caching of the WHO workbook, check for updates, refresh the cache, and inspect workbook version/provenance.
Provides tools for accessing and querying the World Health Organization's Global Health Expenditure Database (GHED), enabling retrieval of health expenditure indicators, country profiles, trend analyses, and research-ready panel data across countries, regions, and income groups.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ghed-mcpCompare health expenditure trends in Brazil and Argentina from 2010 to 2020"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ghed-mcp
A Model Context Protocol (MCP) server that gives AI assistants like Claude direct access to the World Health Organization's Global Health Expenditure Database (GHED), built for comparative health-financing research.
An independent open-source project by Decilion. It is not an official WHO product and does not imply WHO endorsement. WHO provides the underlying data and methodology.
What it does
ghed-mcp wraps the WHO GHED all-data workbook in task-oriented MCP tools so an AI assistant can answer questions like:
"Build me a health-financing profile for Colombia."
"Compare out-of-pocket spending as a share of health expenditure across LAC countries since 2000."
"What's the government priority gradient by World Bank income group?"
"Decompose Peru's current health expenditure by financing scheme for 2023."
Country names, ISO3 codes, WHO region codes (AFR, AMR, EMR, EUR, SEAR, WPR) and World Bank income labels (Low, Lower-middle, Upper-middle, High) are all accepted, with aliases; region="Americas" and income="UMIC" work the same as the canonical values. Collective aliases match the academic global-health convention: income="LMIC" expands to the union of Low + Lower-middle + Upper-middle (not the World Bank's narrower lower-middle-only definition), and income="MIC" expands to Lower-middle + Upper-middle. CSV export is built in.
Related MCP server: WealthGuard MCP
Why this exists
Raw access to GHED is technically possible from an AI assistant with tool access, but in practice it's painful: there is no stable documented API, the all-data workbook contains thousands of variables across the SHA 2011 accounting framework, indicator codes are cryptic (gghed_che is "domestic general government health expenditure as a share of current health expenditure"), and not every variable is additive (you can't sum percentages or PPP values as accounting identities). ghed-mcp collapses the friction:
Discovers the latest GHED all data workbook automatically from WHO's Documentation Centre.
Caches the XLSX locally and builds a derived SQLite database for fast queries.
Steers the model toward headline
INDICATORSfirst, with detailed SHA series available on demand.Knows the additive hierarchies (CHE = HF1+HF2+HF3+HF4+HFnec, GGHE-D = FS1+FS3, …) and validates breakdowns with a sum-vs-parent balance check.
The tool design reflects how health-financing researchers actually work: country profiles, regional and income-group benchmarks, financing-mix decompositions, and metadata for citation.
First visit? Start with the researcher quickstart for a short route through the tools, example research prompts, interpretation checks, and what to include when reporting a problem.
Install
Requires Python 3.11 or newer and an MCP client that can launch local
stdio servers. Check python3 --version (Windows: py -3 --version) and use
a supported interpreter before creating the environment. No WHO API key is required.
The shell examples below use macOS/Linux.
Install 0.6.1 from PyPI in a virtual environment:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install mcp-server-ghed==0.6.1On Windows PowerShell, use:
py -3 -m venv .venv
.\.venv\Scripts\python.exe -m pip install mcp-server-ghed==0.6.1The same wheel and source archive are also available in the GitHub release.
Or install from a source checkout:
git clone https://github.com/Decilion/ghed-mcp.git
cd ghed-mcp
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .Connect your MCP client
Use the absolute path to the executable in the environment where you installed
this package: .venv/bin/ghed-mcp on macOS/Linux, or
.venv\Scripts\ghed-mcp.exe on Windows. In JSON, escape Windows backslashes,
for example C:\\Users\\you\\project\\.venv\\Scripts\\ghed-mcp.exe.
Merge the entry into any existing mcpServers object instead of replacing it.
Claude Code:
claude mcp add --transport stdio --scope user ghed -- /absolute/path/to/.venv/bin/ghed-mcp--scope user makes the server available across Claude Code projects.
Use --scope local if you want it only in the current project.
Claude Desktop: add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"ghed": {
"command": "/absolute/path/to/.venv/bin/ghed-mcp"
}
}
}Codex CLI:
codex mcp add ghed -- /absolute/path/to/.venv/bin/ghed-mcpThis writes an entry to ~/.codex/config.toml. List or remove with codex mcp list / codex mcp remove ghed.
Other MCP-compatible clients (Cursor, Cline, Continue, etc.): point them at the ghed-mcp console script in your venv. The MCP protocol is the same across clients; only the registration UI differs.
Restart your client. The ghed server should appear with all tools, the ghed://indicator/{indicator_code}, ghed://methodology, ghed://topics/{topic_id}, and ghed://research-use-cases/{use_case} resources, and the compare_health_expenditure prompt available.
Verify the installation
python -m pip show mcp-server-ghedOn Windows, use .\.venv\Scripts\python.exe -m pip show mcp-server-ghed.
After registering and restarting your client, ask it to call ghed's
topics_index tool. This checks the connection without fetching WHO data.
The server exposes 35 tools. Running ghed-mcp alone in a terminal
starts a stdio process that waits for an MCP client; it is not a web server.
If the server is missing, verify the absolute executable path, install into that same environment, and restart the client. If WHO is temporarily unavailable, retain the error and retry later; an upstream failure does not mean no data exists.
The first workbook-backed call downloads the GHED workbook and builds a
SQLite cache under ~/.cache/ghed-mcp/. Even cache_status and version can
trigger this on an empty cache; topics_index and research_use_cases do not.
Initial setup can take several minutes depending on the workbook, network and
machine. Later calls reuse the cache. Set GHED_MCP_CACHE_DIR in the MCP server
process environment to relocate it (use the client configuration when launched
from a desktop app). check_for_updates checks WHO metadata without downloading
the workbook; call refresh_cache explicitly to adopt a newer workbook.
To prepare the cache before connecting a client with a short tool timeout, run this once in the installation environment and wait for it to finish:
ghed-mcp --warm-cacheOn Windows, run .\.venv\Scripts\ghed-mcp.exe --warm-cache. This downloads
and indexes the workbook only if needed, prints the status, and exits.
It does not refresh an existing workbook. Use the same GHED_MCP_CACHE_DIR
for this command and the MCP client if you customized that setting.
Research correctness and cache behavior
country_profile(year=...) uses the latest available value at or before the
requested year. Inspect actual reference years and mixed_reference_years.
Curated group members are matched against exact workbook ISO3 codes. Responses
include country_group_resolution with supported and unsupported members; an
empty resolved group returns no data. Explicit country lists, region and income
filters use the same intersection for extracts and quality assessments. Exact
country names take precedence over partial-name matching, and three-letter
inputs are treated as ISO3 codes regardless of case.
Accounting breakdowns report complete, missing_children, and
balance_status (balanced, unbalanced, or incomplete). balanced is null
when a parent or component is missing. Missing values are not assumed to be
zero; a fully observed zero parent and zero components can balance. Research
panels, trends and rankings include workbook provenance and query parameters.
The source.workbook_version lines come from the same SQLite snapshot as the
response data. workbook_modified_at is a local cache timestamp, not WHO's
release date. When an income filter is provided, source.income_resolved
lists the matching workbook classes, including collective alias expansion.
Downloads are serialized across processes, checked for archive integrity and
required worksheet layouts, then atomically replace the workbook. Invalid
downloads preserve the existing cache. Workbook validation and SQLite rebuilds
run off the MCP event loop. Rebuilds are serialized separately
and detect source changes; an interrupted or incompatible workbook produces an
actionable error. Data requests wait for an active rebuild while the MCP event
loop remains responsive. Each store keeps one SQLite snapshot, and response
provenance identifies the workbook signature behind that snapshot. Later requests
detect a changed workbook and open the refreshed cache.
An update during a rebuild can require retrying the query. refresh_cache remains
an explicit operation; ordinary queries do not check WHO for new releases.
The server uses the FastMCP API from MCP SDK 1.x (mcp>=1.15.0,<2). SDK 2.x is
excluded because it changes that server API. Cache locking uses filelock.
Updating
Version 0.6.1 includes the September 2026 correctness and reliability fixes.
See the changelog for details.
Upgrade in your existing virtual environment with
python -m pip install --upgrade mcp-server-ghed.
The unpinned upgrade command selects the newest compatible PyPI release.
Use mcp-server-ghed==0.6.1 to reproduce the version documented here.
From your existing clone, with its virtual environment active:
git pull --ff-only
python -m pip install -e .Restart the MCP client so its server process loads the updated code and tool
schemas. Installation resolves declared dependencies: both servers require
mcp>=1.15.0,<2 and httpx>=0.27.0; GHED additionally requires
openpyxl>=3.1.0 and filelock>=3.16,<4.
Compatibility changes in 0.6.1
In 0.6.1, single-observation trends return null change metrics and
change_status="insufficient_observations"; change rankings exclude them.
Two or more observations are needed to establish change. Relative change
and CAGR retain their existing fractional numeric convention, now labeled.
Compatibility changes in 0.6.0
Consumers should handle balanced=null when accounting components are missing,
inspect balance_status and complete, and use country_group_resolution to
understand excluded economies. Empty selections no longer expand to global data.
Workbook-derived responses include source.dataset_signature, which identifies
exactly the cached data used by that response.
Tool reference
Tables show defaults; capped limits are clamped to at least 1 and at most the listed maximum. Raising a limit beyond that maximum does not fetch more data.
Tool signatures below show the canonical parameter names, omitting deprecated
aliases. Unknown arguments are rejected, so getting the names right matters. In particular: country is singular, countries is the list form, year filters are year_start / year_end (not year_from / year_to).
find_country_code also accepts the deprecated country_name alias. Pass
either country or country_name, not both.
Cache and version
Tool | Signature | Purpose |
|
| Download or re-download the public GHED workbook and rebuild SQLite |
|
| Workbook, SQLite cache, source document, and row counts |
|
| Compare the cached source document with the current all-data workbook metadata |
|
| Workbook version lines and cache provenance |
Methodology and discovery
Tool | Signature | Purpose |
|
| GHED variable classes, categories, cautions, and curated topics |
|
| Curated topic map for common health-expenditure questions |
|
| Literature-inspired GHED research workflows and recommended variables |
|
| Map a natural-language research question to likely GHED variables and cautions |
|
| Counts by GHED Codebook category |
|
| Paginated headline indicators only ( |
|
| Paginated full GHED codebook variables; |
|
| Search headline indicators by default; |
|
| Search all variables, including detailed SHA series; |
|
| Codebook metadata for one variable |
Country resolution
Tool | Signature | Purpose |
|
| Countries and territories in the workbook, optionally by group |
|
| Available GHED region and income group values |
|
| Bundled group definitions, member counts and verification date |
|
| Expand a bundled group to its ISO3 list |
|
| Resolve a country name fragment or alias to ISO3 |
|
| Source, data-type, and estimation notes from the Metadata sheet; |
|
| Latest headline health expenditure values for one country |
Data extraction
Tool | Signature | Purpose |
|
| One indicator with optional country/group/year filters; |
|
| One indicator across countries, returned as tidy rows or CSV; |
|
| One indicator across a country group (curated, regional, or income-based); |
|
| Group stats, coverage, top/bottom countries, and mixed-year warnings; |
|
| First/latest country trends for one indicator; |
|
| First/latest country trends with period guards and per-indicator warnings; |
|
| Rank countries by absolute change, percent change, or CAGR; |
Research workflows
Tool | Signature | Purpose |
|
| Availability summary for one or more variables before panel construction |
|
| Tidy long panel for multiple variables, countries, and years; |
|
| Export-ready data CSV, codebook CSV, availability CSV, and README text; |
Quality and accounting checks
Tool | Signature | Purpose |
|
| Known additive parent-child relationships for a variable |
|
| Classify a variable as total, component, ratio/share, amount, or context series |
|
| Country-year breakdown with child sum, shares, and balance check |
|
| Availability, metadata completeness, data-type mix, and caution flags; |
Resources and prompts
Resource
ghed://indicator/{indicator_code}: readable view of one indicator's metadataResource
ghed://methodology: readable methodology guide for variable selectionResource
ghed://topics/{topic_id}: readable view of a curated topic and its indicator codesResource
ghed://research-use-cases/{use_case}: readable view of one research use casePrompt
compare_health_expenditure(countries, indicator): guided template for cross-country health-financing analysis
summarize_country_group(year=...) restricts observations to that exact year
and takes precedence over latest_only. This differs from the at-or-before
reference year used by country_profile.
Trend results label percent_change and cagr as fractions: 0.10 means
10% relative change or 10% per year, respectively. absolute_change is in the
indicator's units, or percentage points for percentage indicators. Read
change_units, actual first/latest years and any possibly_truncated flag.
Multi-indicator trends apply the same period guards as single-indicator trends.
Regional analysis
Both gho-mcp and ghed-mcp expose the same curated country groupings beyond what WHO and the World Bank publish as built-in dimensions. Pass country_group="LAC" (or any of the codes below) on the data tools and the server resolves to the right ISO3 list without you having to enumerate codes by hand.
Available groups
These are the definitions bundled with this release, last verified in the source on 2026-05-06. The World Bank region lists use FY2026 definitions; they are not a live classification service. Member counts describe the curated lists, not the number of economies with observations in either WHO database.
Code | Definition | Members |
| 33 sovereign Latin American & Caribbean states (PAHO/Decilion convention) | 33 |
| World Bank's 42-economy LAC region: sovereign states plus territories (Aruba, Cayman Islands, Curaçao, Puerto Rico, etc.) | 42 |
| World Bank East Asia & Pacific (FY2026) | 38 |
| World Bank Europe & Central Asia | 58 |
| World Bank "Middle East, North Africa, Afghanistan and Pakistan" (FY2026) | 23 |
| MENA without Israel and Malta | 21 |
| World Bank North America (Bermuda, Canada, USA) | 3 |
| World Bank South Asia (FY2026; without AFG and PAK, now in MENA) | 6 |
| World Bank Sub-Saharan Africa | 48 |
| UN Least Developed Countries | 44 |
| OECD member countries | 38 |
Aliases include natural-language ("Latin America and Caribbean", "Sub-Saharan Africa", "Least Developed Countries") and official codes (LCN, SSF, etc.). Two read-only tools, list_curated_country_groups and resolve_country_group_membership, let an assistant inspect or expand the lists at runtime.
Using country_group= on the data tools
country_group= merges (deduplicated) with any explicit countries= list and composes with region / income via AND semantics. The comments below identify the server for each call; indicator codes and
response schemas differ between GHO and GHED:
compare_countries(indicator_code="oops_che", country_group="LAC",
latest_only=True) # GHED
compare_countries(indicator_code="WHOSIS_000001", country_group="OECD",
year_start=2010, year_end=2023) # GHO
list_curated_country_groups() # both
resolve_country_group_membership("LAC") # both; returns 33 ISO3 codes
# ghed-mcp also exposes:
build_research_panel(indicator_codes=["che_gdp", "gghed_che"],
country_group="OECD", year_start=2000, year_end=2024)
summarize_country_group(indicator_code="ext_che", country_group="LDC",
latest_only=True)
list_countries(country_group="LAC", income="High") # LAC HICsCurated-group members use exact workbook ISO3 membership. Unsupported members are omitted and reported in country_group_resolution; an empty selection returns no observations. User-supplied countries= remain strict, so unsupported codes and ambiguous names raise errors.
Membership cadence
The lists are static Python data bundled with each package. Membership changes require a new package release and an upgrade; reinstalling the same version does not refresh them. WHO region/income values come from the respective data source and should not be assumed to represent historical classifications for every year.
The LAST_VERIFIED constant in country_groups.py records the bundled review
date. Consult the authoritative sources for subsequent changes:
World Bank country and lending groups: https://datahelpdesk.worldbank.org/knowledgebase/articles/906519
UN Least Developed Countries: https://policy.desa.un.org/least-developed-countries
OECD members: https://www.oecd.org/en/about/members-partners.html
Re-check before time-sensitive group comparisons and at least annually. See the UN graduation updates for scheduled changes; the package does not automatically remove graduating LDCs.
The same country_groups.py file lives in both ghed-mcp and gho-mcp (canonical source: ghed-mcp), so both servers start from the same definitions. Actual supported membership
can differ by database; inspect the reported unsupported codes before combining extracts.
Examples
Each block below shows a natural-language prompt and a sketch of the underlying MCP tool calls. These calls illustrate arguments for your assistant; they are not standalone Python scripts. Actual values and coverage depend on the WHO source.
Country profile
"Give me a Colombia health-financing profile."
The assistant calls country_profile(country="Colombia") and returns CHE as % GDP, CHE per capita (USD), government share of CHE, OOP share of CHE, external share, GGHE-D as % GDP, and GGHE-D as % GGE, each for its latest available year, with a mixed_reference_years warning if reference years differ.
Comparative LAC analysis
"Compare out-of-pocket burden across the Andean countries since 2000."
The assistant calls:
compare_countries(
indicator_code="oops_che",
countries=["Colombia", "Ecuador", "Peru", "Bolivia", "Venezuela"],
year_start=2000,
latest_only=False,
format="csv",
)and gets a CSV ready to drop into any spreadsheet, statistical package, or charting tool.
Income-group gradient
"What's the public health-spending priority gradient by income group in 2022?"
The assistant calls summarize_country_group("gghed_gge", income="upper middle income", year=2022) (and parallel calls for Low / Lower-middle / High), returning median, top-five, and bottom-five for each group with coverage ratios.
Accounting-identity decomposition
"Decompose Peru's CHE in 2023 by financing scheme."
The assistant calls:
explain_indicator_relationship(indicator_code="che")
build_additive_breakdown(
indicator_code="che",
country="Peru",
year=2023,
relationship_id="che_by_financing_scheme",
)and returns each child component (HF.1, HF.2, HF.3, HF.4, HF.nec) with its value,
share of the parent where defined, and balance_status. Only fully observed
components that reconcile within tolerance yield balanced=true; incomplete
breakdowns return balanced=null, and complete non-reconciling ones return false.
This example does not assume Peru has a complete breakdown for that year.
From CSV output to analysis tools
compare_countries(..., format="csv") and build_research_package(...) return CSV strings under the csv, data_csv, codebook_csv, or availability_csv keys. Two common downstream paths:
To pandas, for time-series analysis or modelling (optional: install
pandas and matplotlib in your analysis environment):
import io, pandas as pd
# csv_text is the value of result["csv"] from compare_countries
df = pd.read_csv(io.StringIO(csv_text))
df["year"] = df["year"].astype(int)
# For a multi-indicator research package, select one indicator first.
if df["indicator_code"].nunique() != 1:
raise ValueError("Select one indicator before pivoting.")
if df.duplicated(["country_code", "year"]).any():
raise ValueError("Resolve duplicate country-year observations before pivoting.")
df = df.pivot(index="year", columns="country_code", values="value")
df.plot(title="Selected indicator by country")To any external tool (Excel, Google Sheets, R, Stata, Tableau, charting platforms, etc.):
with open("data.csv", "w", encoding="utf-8", newline="") as f:
f.write(csv_text)The columns indicator_code, indicator_name, country_code, country_name, region, income, year, value, unit, currency are tidy-format-friendly and map cleanly into most analysis or visualization workflows. For wide-format / per-country columns, pivot first (pandas snippet above).
Advanced queries
The friendly tools cover headline indicators, country/region/income-group filtering, year ranges, and the most common additive decompositions. For detailed SHA series by function, provider, disease/condition, cross-tabs, capital, age or COVID-19 reporting items, explore the full codebook with list_variables and search_variables:
list_variables(category_1="HEALTH EXPENDITURE DATA", category_2="HEALTH CARE FUNCTIONS")
search_variables(query="diabetes", category_1="HEALTH EXPENDITURE DATA")For long-code SHA hierarchies (e.g. sha11.HC, sha11.HP, sha11.HF), use additive_hierarchy(indicator_code=...); it returns curated codebook formulas first, then inferred direct children from the SHA long-code tree for current-NCU amount variables. Pair with build_additive_breakdown to validate any decomposition for a country-year.
Inspect what each variable actually is before pulling; explain_indicator_relationship(indicator_code) classifies it as additive_parent, component, derived_ratio_or_share, amount_series, or context_or_conversion_series and surfaces interpretation cautions.
Topics covered by topics_index
core_spending: CHE level and scale (CHE/GDP, CHE per capita USD/PPP)government_spending: GGHE-D level, share of CHE, share of GDP, fiscal priority (GGE)out_of_pocket: household burden, OOP share of CHE, OOP per capitaexternal_aid: external funding for health and donor dependenceprivate_spending: private domestic spending and voluntary prepaymentcapital: capital health expenditure (HK)primary_health_care: PHC level and share of CHEmacro_context: GDP, population, exchange rates, PPP conversion factors
research_use_cases adds literature-inspired patterns:
health_financing_transition: financing-mix change with income, time, or reformfinancial_protection_oop: OOP indicators as macro context for UHC researchgovernment_priority: government health-spending effort and priority in the public budgetdonor_dependence: external funding dependence and its trajectoryprivate_and_voluntary_insurance: private, voluntary, and prepaid arrangementsservices_providers_sha: detailed SHA series by function, provider, scheme, sourceprimary_health_care: PHC spending levels and shares
Development
The 2026-09-15 regression suite contains 91 tests. GitHub Actions runs it on Python 3.11, 3.12, 3.13 and 3.14 against the minimum and latest compatible MCP SDK, then builds both distributions and checks wheel imports outside the checkout.
python -m pip install -e ".[dev]"
python -m pytestTests use a synthetic GHED workbook fixture; no network access required for the standard suite.
Release procedure: RELEASING.md.
Limitations
GHED does not expose a stable documented API. This server discovers the current all-data workbook from WHO's Documentation Centre and caches it locally; if WHO changes the directory structure or naming, the discovery logic will need a release.
A workbook-backed cold start downloads and indexes the data. Download size, cache size and duration vary by WHO release and machine; allow several minutes and sufficient disk space.
The all-data workbook contains thousands of variables; use
list_variable_categoriesfor current counts. Detailed SHA series can be sparse for recent years; usedata_availabilityandassess_data_qualitybefore strong claims.Variant series (current NCU, constant NCU, current USD, constant USD, PPP, per-capita, %CHE, %GDP, %GGE) are not interchangeable as accounting identities.
additive_hierarchyandbuild_additive_breakdownonly validate current-NCU amount variables.This server does not produce population-weighted regional, income-group or global estimates.
summarize_country_groupcomputes unweighted descriptive statistics over returned observations. Withlatest_only=True, country years can differ; use a fixedyearand inspect coverage for comparable summaries.Latest years can be preliminary; inspect Version sheet and Metadata notes via
versionandget_country_metadata.
About the WHO Global Health Expenditure Database
The WHO Global Health Expenditure Database (GHED) is the World Health Organization's central platform for internationally comparable data on health spending. It is the authoritative source for indicators on:
Levels and trends of health expenditure across countries and territories, with most series running from 2000 onward; use
list_countriesfor the cached workbook coverageSystem of Health Accounts 2011 (SHA 2011): health expenditure decomposed by financing arrangements, revenues, providers, functions, diseases and conditions, capital formation, and primary health care
Universal Health Coverage financing context: government share, out-of-pocket burden, external funding, voluntary insurance
Macro denominators and conversion variables: GDP, population, exchange rates, price indexes, to support per-capita, %GDP, constant-price, and PPP-adjusted analysis
Country-level metadata: sources, data type (Documented / Estimated / Imputed), methods of estimation, country footnotes, for transparent citation
GHED underpins WHO's Global Spending on Health annual report, Health Accounts country profiles, and the financing chapter of World Health Statistics. The data is free and openly published through the all-data workbook this server wraps, and through the official GHED portal with its own visualizations and downloads.
This MCP server is plumbing. The data, the indicator definitions, the SHA 2011 methodological work, and the country-level data validation are all WHO's. If you use values retrieved through this server, please:
Cite WHO as the source. The
sourceblock on every data response includes the workbook path, modification time, parameters, and retrieval timestamp to make this straightforward.Visit the GHED portal for indicator metadata, methodology notes, and the official visualizations. The MCP exposes the data: the portal provides the canonical context.
Read the global health expenditure reports for WHO's curated narrative analysis of what the data shows.
GHED is a public good. The most valuable contribution any user can make is to support and reference WHO's underlying data work.
Built by
Decilion provides global health consulting across Latin America and the Caribbean, including applied AI for global health.
This server is one of Decilion's open-source contributions to the global health data community. It pairs naturally with gho-mcp for combined GHO + GHED workflows. If you use it in research, a brief acknowledgment is appreciated but not required.
License
MIT. See LICENSE.
Available Tools
35 toolsadditive_hierarchyBRead-only
Return known additive child relationships for a GHED variable.
| Name | Required | Description | Default |
|---|---|---|---|
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and openWorldHint=true, so the description does not need to repeat those. It adds the concept of 'known' relationships, implying non-exhaustiveness, which aligns with openWorldHint. No contradictions, but little added insight beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 9 words, front-loading the key action. It is efficient, but could benefit from a brief second sentence clarifying the parameter or output scope. Nonetheless, it achieves conciseness without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple lookup tool with one parameter and an existing output schema, the description is minimally adequate. It misses details like error handling or data scope, but for a straightforward retrieval, it passes the bar. Improvement would add context about the return format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description mentions 'GHED variable' but the parameter is named 'indicator_code', with no explanation of their relationship. The agent must infer that indicator_code identifies a GHED variable. This is insufficient for low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'return' and the resource 'additive child relationships for a GHED variable'. It distinguishes from siblings like 'explain_indicator_relationship' by focusing on returning existing relationships rather than explaining them. However, 'GHED' is not explained, which assumes domain knowledge.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. There is no mention of context, prerequisites, or when not to use it. This leaves the agent without decision support.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_data_qualityCRead-only
Summarize metadata, availability, and cautions for an indicator/filter.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| income | No | ||
| region | No | ||
| country | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| country_group | No | ||
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so no contradiction. The description adds minimal behavioral context beyond the purpose—it does not explain what 'cautions' entails or how the summary is structured. For a read-only tool, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise with no wasted words. However, it is too brief given the tool's complexity, sacrificing necessary details for brevity. A few more sentences would improve completeness without harming conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks context about the output (which has an output schema) and the interplay of filters. Despite having many parameters, no explanation of how they are used together is provided. The tool's complexity demands more descriptive context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning no parameter descriptions exist in the schema. The description only mentions 'indicator/filter' but does not explain any of the 9 parameters, such as geography filters or year range. This is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Summarize') and the resource ('metadata, availability, and cautions for an indicator/filter'). It distinguishes itself from siblings like 'data_availability' and 'get_indicator_metadata' by combining these aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide guidance on when to use this tool versus alternatives. For instance, if only metadata is needed, 'get_indicator_metadata' might be more appropriate; if only availability, use 'data_availability'. No such context is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_additive_breakdownBRead-only
Check a current-NCU accounting total for one country-year.
Use additive_hierarchy to choose a relationship. If complete is false, the identity cannot be validated even when observed components nearly sum to the total. Missing components are not zero or implicitly included elsewhere.
| Name | Required | Description | Default |
|---|---|---|---|
| year | Yes | ||
| country | Yes | ||
| indicator_code | Yes | ||
| relationship_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint. The description adds specific behavioral context: if complete is false, the identity cannot be validated even when components sum nearly to the total, and missing components are not zero or implicitly included. This aligns with the openWorldHint and provides useful nuance beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loaded with the primary purpose. It avoids repetition and each sentence earns its place by adding either the core action or critical behavioral caveats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists (so return format is covered), the description is incomplete for successful invocation. It fails to explain the parameters, making it unclear what indicator_code, country, year, and relationship_id should contain. The reference to additive_hierarchy is vague without guidance on how to use it. The tool's name implies building a breakdown, but the description only says 'check', creating uncertainty about expected behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning no parameter descriptions exist in the schema. The description does not explain any of the four parameters (year, country, indicator_code, relationship_id) or their formats, meanings, or defaults. It only mentions 'country-year' and references additive_hierarchy, leaving the agent without guidance on how to populate the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: checking a current-NCU accounting total for one country-year. It names the resource and the action, and references the sibling additive_hierarchy for selecting a relationship, which helps differentiate it. However, the name suggests 'building' a breakdown while the description says 'check', creating slight ambiguity about the primary action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises using additive_hierarchy to choose a relationship, providing some usage direction. It also explains a condition (complete=false) that affects validation interpretation. However, it does not explicitly state when to use this tool versus alternatives, nor does it give exclusions or alternative tools for similar tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_research_packageBRead-only
Build export-ready CSV data, codebook, and README text for a research extract.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| income | No | ||
| region | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| country_group | No | ||
| indicator_codes | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true and openWorldHint=true, so the description's mention of 'build' might imply creation, but the readOnlyHint indicates no state change. The description adds context about the output format (CSV, codebook, README) beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence covers the tool's outputs efficiently. However, it could be slightly more structured (e.g., listing outputs) without adding length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters with zero description coverage, the description is incomplete. It does not explain parameter usage, valid combinations, or how to configure the export. The output schema exists but does not compensate for missing parameter guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and 8 parameters, the description offers no parameter-level explanation. For example, it does not clarify the difference between 'countries' and 'country_group' or how 'top' limits rows. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool produces export-ready CSV data, codebook, and README text for a research extract. This specific verb-output combination distinguishes it from sibling tools like build_additive_breakdown or build_research_panel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when a researcher needs an exportable package of data and documentation, but it lacks explicit guidance on when not to use it or alternatives among many siblings (e.g., get_indicator_data for raw data).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_research_panelARead-only
Build a tidy long panel for multiple GHED variables across countries and years.
country_group accepts curated codes ("LAC", "OECD", "LDC", "SSA", …)
that resolve to ISO3 lists; combine freely with countries, region,
and income (filters apply via SQL AND).
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| format | No | rows | |
| income | No | ||
| region | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| country_group | No | ||
| indicator_codes | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint, openWorldHint), the description discloses that country_group resolves to ISO3 lists and filters combine via SQL AND, offering behavioral insight into parameter interaction. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using two sentences to convey purpose and key parameter behavior without extraneous detail. It is well-structured with a clear main statement followed by a specific note on country_group.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 9 parameters and an output schema, the description lacks completeness by not explaining most parameters or the output format. While annotations cover safety, the description does not fully enable an agent to use the tool correctly, though the existence of an output schema mitigates return value ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage of 0%, the description must compensate. It only explains the country_group parameter and the filter combination logic, leaving 8 parameters (indicator_codes, countries, region, income, year_start, year_end, top, format) unaddressed. This is insufficient for an agent to correctly invoke the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool builds a tidy long panel for multiple GHED variables across countries and years, specifying the resource (GHED variables), action (build panel), and scope (countries/years). This distinguishes it from sibling tools like get_indicator_data which retrieve single indicators, or compare_countries which focus on comparisons.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how country_group works with curated codes and that filters (countries, region, income) combine via SQL AND, providing context for usage. However, it does not explicitly state when to use this tool vs alternatives, such as when a multi-variable panel is needed versus simple indicator retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cache_statusARead-only
Return cache status and WHO source-document details; initializes an empty cache.
The first call may take minutes. Run ghed-mcp --warm-cache in a terminal first if the MCP client has a short tool timeout. Use check_for_updates to check for newer WHO data without downloading the workbook.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint. The description adds valuable behavioral context: it discloses the cache initialization side effect (though read-only) and the performance implication of the first call taking minutes. This is beyond what annotations convey and helps the agent anticipate behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no waste. The primary purpose is front-loaded, followed by essential guidance and an alternative. Every sentence earns its place, and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the return format is covered. The description explains the side effect (cache init), performance characteristics, and provides an alternative for update checks. For a parameterless tool, this is fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the baseline is 4. The description correctly omits parameter details since none exist, and there is no additional parameter-related meaning needed. The description adds value through behavioral context, not parameter explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns cache status and WHO source-document details, and initializes an empty cache. This clearly distinguishes it from siblings like refresh_cache (which refreshes) and check_for_updates (which checks updates without downloading). The purpose is explicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it warns that the first call may take minutes and suggests pre-warming with a terminal command, and it directs users to check_for_updates as an alternative for checking newer data without downloading. This is a clear when-to-use and when-not-to-use with an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_for_updatesARead-only
Compare the local cache source document with the current GHED all-data file.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only and open-world behavior. The description adds context about what is being compared but does not introduce behavioral traits beyond those already captured.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that is concise and contains no unnecessary words. Every part contributes to the meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of parameters and the presence of an output schema, the description is sufficiently complete. It could mention what the output indicates, but the core function is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so schema coverage is 100%. The description does not need to add parameter information; a baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('compare') and the specific resources ('local cache source document' vs 'current GHED all-data file'), distinguishing it from sibling tools like cache_status and refresh_cache.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case (checking cache freshness) but does not explicitly state when to use this tool versus alternatives like cache_status or refresh_cache.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_countriesARead-only
One GHED indicator across countries, returned as tidy rows or CSV.
Pass either an explicit countries list or a curated country_group
(e.g. "LAC", "OECD", "LDC", "SSA") — or both, in which case they are
merged. See list_curated_country_groups for available groups.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| format | No | rows | |
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| latest_only | No | ||
| country_group | No | ||
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds context beyond annotations: it mentions output format options (rows vs CSV) and parameter merging behavior. Annotations already declare readOnlyHint and openWorldHint, so the description supplements without contradicting. It does not detail rate limits or other potential behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: two sentences that front-load the purpose, then provide parameter guidance. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters (1 required) and 0% schema description coverage, the description covers the core parameters but misses several optional ones. Output schema exists, so return values are covered. It is adequate for basic use but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain all parameters. It covers indicator_code, countries, country_group, and format, but not year_start, year_end, latest_only, or top. These omissions leave significant gaps for an 8-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns 'One GHED indicator across countries' in tidy rows or CSV format. It specifies the core resource (GHED indicator) and action (compare across countries), distinguishing it from siblings like 'compare_country_group' or 'get_indicator_data'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to use countries vs country_group parameters, including that they can be merged, and references list_curated_country_groups for available groups. It does not explicitly state when not to use this tool, but the guidance is sufficient for typical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_country_groupARead-only
Compare one GHED indicator across a country group (curated, regional, or income-based).
Pass at least one of country_group (e.g. "LAC", "OECD", "LDC"),
region, or income. Multiple filters compose via SQL AND, so
country_group="LAC", income="High" returns LAC HICs.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| format | No | rows | |
| income | No | ||
| region | No | ||
| year_end | No | ||
| year_start | No | ||
| latest_only | No | ||
| country_group | No | ||
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds filtering behavior (AND composition) but does not detail response format or pagination. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise paragraphs: first sentence defines purpose, second explains filter logic. No redundant words, front-loaded with key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters and an output schema, the description lacks details on what 'compare' returns (e.g., summary statistics, differences), and omits explanation for year range, latest_only, top, and format parameters. Incomplete for the complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%. The description explains the key grouping parameters (country_group, region, income) and their composition, but ignores other parameters like year_start, year_end, latest_only, top, and format. Partial compensation for low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'compare' and the resource 'GHED indicator across a country group'. It distinguishes from siblings like compare_countries and summarize_country_group by specifying the grouping dimension (curated, regional, income-based).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance: 'Pass at least one of country_group, region, or income' and explains filter composition via SQL AND. However, it does not explicitly state when to avoid this tool or suggest alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_trendsARead-only
Compute first/latest summaries for several indicators; use indicator_trend for one.
percent_change and cagr are fractions (0.10 = 10%); absolute changes in shares are percentage points. Check each item's years, warnings and limits.
| Name | Required | Description | Default |
|---|---|---|---|
| income | No | ||
| region | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| country_group | No | ||
| min_year_count | No | ||
| indicator_codes | Yes | ||
| min_period_years | No | ||
| top_per_indicator | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the read-only and open-world annotations, the description adds important output semantics: percent_change and cagr are fractions, absolute share changes are percentage points, and the agent should inspect years, warnings, and limits. This helps the agent correctly interpret results and anticipate caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded; the main action is stated first, followed by crucial numeric interpretation notes and a caveat. Every sentence carries useful information, though some phrasing (e.g., 'first/latest summaries') is slightly terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, zero schema descriptions, and no parameter-level guidance in the prose, the description is not complete enough for an agent to confidently set filters and thresholds. The output schema exists and the numeric caveats help, but input semantics remain largely undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain any of the 10 parameters. It references indicators and output metrics, but nothing about country/region filters, year ranges, min_year_count, min_period_years, or top_per_indicator. The description fails to compensate for the complete lack of schema-level parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation ('Compute first/latest summaries') on a specific resource ('several indicators') and explicitly contrasts itself with indicator_trend for a single indicator. This makes its scope clear and distinguishes it from a key sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent to use indicator_trend when only one indicator is needed, giving a clear routing rule. It does not discuss other siblings like compare_countries or compare_country_group, but the single-vs-several indicator distinction provides clear enough context for the main alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
country_profileBRead-only
Latest headline health-expenditure values for one country.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | ||
| country | Yes | ||
| indicator_codes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds that the tool returns 'latest' values, implying currency but not high transparency. No mention of data limits, pagination, or potential warnings, but the annotations cover the safety profile adequately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no unnecessary words. It is front-loaded and efficient, conveying the core purpose without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, no schema descriptions, and a tool that likely has nuanced optional filtering (year, indicator codes), the description is too sparse. While an output schema exists, the lack of parameter guidance makes the tool significantly incomplete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters. It only hints at 'country' implicitly, but does not explain 'year' or 'indicator_codes' meanings, formats, or relationships. This leaves the agent without critical information to use optional parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as providing 'latest headline health-expenditure values for one country.' This succinctly states the verb (provides), resource (headline health-expenditure values), and scope (one country), distinguishing it from sibling tools like compare_countries that handle multiple countries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives. It does not mention when not to use it or point to siblings like compare_countries or get_indicator_data for different needs, leaving the agent to infer usage without explicit direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
data_availabilityARead-only
Summarize availability for indicators before building a research panel.
| Name | Required | Description | Default |
|---|---|---|---|
| income | No | ||
| region | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| country_group | No | ||
| indicator_codes | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is consistent with annotations (readOnlyHint=true), indicating a read-only summary operation. However, it adds no behavioral details beyond what the annotations already convey, such as data freshness or pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the key verb and resource. It is appropriately brief but could benefit from slight elaboration on parameter usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, return values need not be described. However, with 7 parameters and 0% schema description coverage, the description is incomplete for guiding parameter selection. Annotations partially compensate, but parameter semantics are lacking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, leaving 7 parameters completely undocumented. The description mentions 'indicators' but does not elaborate on the meaning or usage of the fields like 'countries', 'region', 'year_start', etc. This is insufficient for an agent to correctly fill parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb 'summarize' and resource 'availability for indicators', and the context 'before building a research panel' distinguishes it from sibling tools like 'build_research_panel' and 'get_indicator_metadata'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear timing guidance ('before building a research panel'), but does not explicitly state when not to use the tool or mention alternatives among the 34 sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_indicator_relationshipBRead-only
Explain whether a variable is a total, component, ratio/share, or context series.
| Name | Required | Description | Default |
|---|---|---|---|
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include 'readOnlyHint: true' and 'openWorldHint: true', providing safety and scope context. The description only says 'explain', which is consistent but adds no behavioral details like whether it requires specific permissions or how it handles missing data. Annotations carry most of the burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys the core functionality with no wasted words. It is concise and structured optimally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, read-only, output schema present), the description is minimally viable. It explains what it does but lacks usage context. With an output schema, return values don't need explanation, but some guidance on when to use it would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (parameters have no descriptions). The description mentions 'variable' but the parameter is 'indicator_code' with no explanation of what should be provided, format, or examples. It does not compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Explain whether a variable is a total, component, ratio/share, or context series.' It uses a specific verb ('explain') and resource ('indicator relationship'), distinguishing it from sibling tools like 'additive_hierarchy' and 'build_additive_breakdown'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for understanding how a variable relates to others, but it does not explicitly state when to use this tool vs alternatives like 'additive_hierarchy' or 'build_additive_breakdown'. No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_country_codeCRead-only
Find ISO3 country codes by name, alias, or code fragment.
| Name | Required | Description | Default |
|---|---|---|---|
| country | No | ||
| country_name | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint. The description adds no extra behavioral details such as case sensitivity, partial match behavior, or handling of multiple matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, waste-free. Could benefit from brief structure or examples, but efficient for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple lookup tool with output schema and annotations. Missing information on behavior when both parameters are null or conflicting, and no fallback guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should clarify parameter roles. It mentions search by name, alias, or code fragment but does not map these to the two parameters (country and country_name), leaving ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool finds ISO3 country codes using name, alias, or code fragment. It is specific and distinguishes from siblings that deal with country metadata or listings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like get_country_metadata or list_countries. No when-not instructions or context for optimal use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_country_metadataARead-only
Return source, data-type and estimation notes from the Metadata sheet.
If no indicator-specific notes exist, query country alone for broader notes. Country notes are context, not proof of the selected indicator's data type or the cause of a trend. Empty metadata does not establish data quality.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| country | No | ||
| indicator_code | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only and open-world, and the description adds non-obvious interpretive caveats: country notes are context, not proof of indicator data type or trend causation, and empty metadata does not establish data quality. This is useful behavioral guidance beyond what the annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and contains no filler. Each sentence earns its place, including the caveats about metadata interpretation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema and read-only annotations, so return-value detail is not required. However, the top parameter remains unexplained, and the relationship to get_indicator_metadata is not clarified, leaving minor gaps for a tool that otherwise relies on clear context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It indirectly maps indicator_code and country via 'indicator-specific notes' and 'query country alone', but it gives no explanation of the top parameter, its default, or its effect on results. This is only partial parameter coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: return source, data-type, and estimation notes from the Metadata sheet. It is clear enough, but it does not explicitly contrast itself with the sibling get_indicator_metadata, so differentiation is more implicit than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on parameter fallback: if no indicator-specific notes exist, query the country alone for broader notes. It does not explicitly name alternatives like get_indicator_metadata or give exclusion criteria, but it gives enough context for an agent to know how to adjust its query.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_indicator_dataARead-only
Fetch one GHED indicator with optional country and year filters.
Spatial filters compose: country (singular), countries (list), and
country_group (curated, e.g. "LAC") merge into a single country list,
and region / income further constrain via SQL AND.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| income | No | ||
| region | No | ||
| country | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| latest_only | No | ||
| country_group | No | ||
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and open-world hints. The description adds behavioral transparency by detailing how spatial filters interact, which goes beyond annotations. However, it does not mention other behaviors like pagination (top parameter) or data freshness, but this is acceptable given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two paragraphs: a clear function statement and a focused explanation of spatial filter composition. It avoids redundancy and front-loads the main purpose. Minor improvement could further tighten the second paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (10 parameters, output schema exists), the description covers purpose and spatial filter logic but omits details on temporal filters, pagination, and behavior of the top parameter. The output schema covers return values, but the agent still needs more behavioral context for complete usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains spatial filter composition (country, countries, country_group, region, income) but fails to describe temporal filters (year_start, year_end), latest_only, or top. This partial coverage leaves gaps for a significant number of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Fetch one GHED indicator with optional country and year filters', specifying the verb (Fetch), resource (GHED indicator), and available filters. However, it does not differentiate from sibling tools like compare_countries or indicator_trend, which may cause confusion about when to use this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how spatial filters compose (merge into single country list, AND constraints), providing some guidance on filter usage. But it lacks explicit context on when to use this tool versus alternatives, such as search_indicators or compare_trends.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_indicator_metadataBRead-only
Return Codebook metadata for one GHED indicator.
| Name | Required | Description | Default |
|---|---|---|---|
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, so the description adds no new behavioral context. It is consistent with the read-only nature, but does not disclose any additional traits like pagination, authentication, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no fluff, efficiently conveying the tool's purpose. It is appropriately front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is simple (one required parameter, output schema exists), the description is adequate. However, it could mention what 'codebook metadata' includes (e.g., definition, source) to improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the tool description does not explain the parameter 'indicator_code' beyond its existence. The description fails to compensate for the missing schema descriptions, leaving the agent with minimal guidance on parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns 'Codebook metadata for one GHED indicator', which is a specific verb-resource combination. It distinguishes from siblings like 'get_indicator_data' (returns data) and 'list_indicators' (lists all indicators), but does not explicitly differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, such as 'search_indicators' or 'list_indicators', or when not to use it. The description simply states what it does without context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
indicator_trendARead-only
Compute first/latest trends for one indicator; use compare_trends for several.
percent_change and cagr are fractions (0.10 = 10%). absolute_change is in the indicator's units, or percentage points for shares. Inspect actual first/latest years and change_units before interpreting or ranking results.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| income | No | ||
| region | No | ||
| year_end | No | ||
| countries | No | ||
| year_start | No | ||
| country_group | No | ||
| indicator_code | Yes | ||
| min_year_count | No | ||
| min_period_years | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the description's job is to add context. It explains the units of percent_change and cagr (fractions) and absolute_change (indicator units or percentage points), which is valuable behavioral info. It doesn't cover every edge case but adds meaningful detail beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no fluff. The first sentence states purpose and alternative; the second covers units and interpretation advice. Information is front-loaded and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, a complex schema, and zero parameter explanation in the description, the tool is severely under-described. The description explains only the core operation and units, but does not address filtering options, output structure, or how the parameters interact. An agent would struggle to use this tool correctly beyond the simplest case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no parameter-specific guidance. The only implied parameter is indicator_code from the main verb, but all 9 optional parameters (top, income, region, year_start, etc.) are left undocumented, forcing the agent to infer their meaning from schema titles alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Compute' with a specific resource ('first/latest trends for one indicator') and explicitly differentiates from the sibling tool compare_trends for multiple indicators. This makes the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names the alternative tool (compare_trends) and the condition for using it ('for several'), and gives practical guidance on interpreting results by inspecting actual years and change_units before ranking. This is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_countriesARead-only
List countries and territories available in GHED, optionally by group.
country_group accepts curated codes (LAC, OECD, LDC, SSA, …) and
intersects with region / income when more than one is set, so
country_group="LAC", income="High" returns LAC HICs.
| Name | Required | Description | Default |
|---|---|---|---|
| income | No | ||
| region | No | ||
| country_group | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint and openWorldHint. The description adds value by explaining that country_group accepts curated codes and intersects with region/income, clarifying the filtering logic beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences. The first states the core purpose, the second adds crucial behavioral detail. No extraneous or redundant text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and annotations, the description covers the essential purpose and key filtering behavior. However, it does not define all parameters, which is a gap for a tool with 0% schema description coverage. It is minimally adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage. The description only elaborates on country_group with examples, but does not define region or income parameters or their accepted values. This leaves significant ambiguity for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List countries and territories available in GHED', using a specific verb and resource. It distinguishes from sibling tools like get_country_metadata (which provides details) and find_country_code (which maps codes), and adds context about optional grouping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (to list countries, optionally filtered by group) and describes the intersection behavior, but does not explicitly state when not to use it or mention alternatives among the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_country_groupsBRead-only
List GHED country grouping values by region and World Bank income class.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint, openWorldHint) indicate basic safety, and the description confirms a listing operation. However, the mention of filtering by region/income is inconsistent with the parameterless schema, potentially confusing behavior. No additional behavioral traits disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single-sentence description is concise, but the phrase 'by region and World Bank income class' is misleading given the empty parameter set. It earns its place but introduces inaccuracy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and an output schema present, the description minimally covers purpose. It lacks context about how the output relates to sibling tools or how to use the returned groups. Some detail is missing for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist (schema coverage 100%), so the description adds value by explaining the tool's purpose (listing groupings by region and income class). This contextualizes what the tool returns beyond the name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists GHED country grouping values, with a specific scope ('by region and World Bank income class'). However, this implies filtering parameters that do not exist in the input schema, creating a slight mismatch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings like list_curated_country_groups or summarize_country_group. The description does not specify context or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_curated_country_groupsARead-only
List Decilion's curated country groupings (WB regions, LDCs, OECD).
Returns groups beyond GHED's built-in WHO regions and World Bank income
classes — World Bank geographic regions, the UN Least Developed Countries
list, and OECD membership. Use the returned members lists with the
countries parameter on data tools, or pass the group code to
resolve_country_group_membership to get just the ISO3 list.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate safe read (readOnlyHint) and no side effects (openWorldHint). The description adds value by specifying the types of groups returned and how to leverage them, but does not detail response format beyond what the output schema likely provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and immediate usage guidance. Every sentence adds unique value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters, existing annotations, and presence of an output schema, the description is complete: it explains what the tool does, what it returns, and how to use the results in other tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the description cannot add parameter semantics. Baseline 4 is appropriate as the schema coverage is 100% and no parameter info is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists Decilion's curated country groupings (WB regions, LDCs, OECD) and explicitly contrasts them with built-in groups, distinguishing it from siblings like list_country_groups and resolve_country_group_membership.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use the tool (for groups beyond built-in) and how to use the returned data (with `members` lists or via resolve_country_group_membership). Lacks explicit 'when not to use' but the contrast with built-in groups implies it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_indicatorsBRead-only
List headline GHED indicators only (category_1 = INDICATORS).
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| skip | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true. The description adds that the tool returns only headline indicators (category_1 = INDICATORS), providing useful filtering context. However, it does not disclose pagination behavior or result limits beyond schema defaults.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-front-loaded sentence that captures the core functionality without unnecessary words. Every word serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description sufficiently explains the tool's purpose but lacks detail on parameter semantics. Given the presence of an output schema and the simplicity of the tool, it is minimally adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the meaning or usage of 'skip' and 'top' parameters. It only provides the titles 'Skip' and 'Top', which are not fully explanatory for an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List headline GHED indicators only (category_1 = INDICATORS)', specifying the verb 'list', the resource 'headline GHED indicators', and a filtering condition. This distinguishes it from sibling tools like 'list_countries' or 'list_variables'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives such as 'search_indicators' or 'get_indicator_metadata'. There is no explicit when-to-use or when-not-to context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_variable_categoriesARead-only
List GHED variable category counts from the Codebook.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint. The description adds minimal context: it lists counts from the Codebook. It does not elaborate on potential changes, frequency, or source limitations beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no unnecessary words. It efficiently conveys the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters, an existing output schema, and annotations, the description is complete enough. It identifies the data source ('Codebook') and the nature of the output ('counts'), which is sufficient for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so baseline is 4. The description need not add parameter details; the schema coverage is 100% by default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and the specific resource ('GHED variable category counts from the Codebook'). It distinguishes itself from siblings like 'list_variables' by focusing on categories and providing counts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, such as 'list_variables' or 'search_variables'. The description does not specify any context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_variablesCRead-only
List all GHED Codebook variables, optionally filtered by category.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| skip | No | ||
| category_1 | No | ||
| category_2 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds the source (GHED Codebook) and filtering capability, but does not elaborate on behavioral traits like pagination or listing scope. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, efficiently stating the core functionality. However, it is slightly too brief, sacrificing clarity on parameters and usage context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of 4 parameters, many sibling tools, and no schema descriptions, the description is incomplete. It fails to explain parameter usage, provide guidance on when to use this vs. search_variables, or clarify the role of categories.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters. It only mentions optional filtering by category, leaving category_1, category_2, skip, and top undefined. The skip and top parameters for pagination are entirely unmentioned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists GHED Codebook variables with optional filtering by category. It uses a specific verb 'List' and resource 'variables', but does not explicitly differentiate from similar sibling tools like search_variables or list_variable_categories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as search_variables or list_variable_categories. It does not mention prerequisites, constraints, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
methodology_guideARead-only
Explain how GHED variables are organized and how to choose the right series.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint and openWorldHint as true. The description adds value beyond those by specifying the explanatory nature of the tool—it conveys knowledge rather than performing data mutations. No contradictions exist, and the description appropriately complements the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that immediately conveys the tool's purpose. Every word earns its place; there is no redundancy or unnecessary detail. It is front-loaded and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, explanatory role), the description is complete enough. An output schema exists, so return values do not need description. The description adequately covers what the tool does and why an agent would use it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so baseline score is 4 as per instructions. The description does not need to elaborate on parameters. Schema coverage is effectively 100% since there are no parameters to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool explains how GHED variables are organized and how to choose the right series. It uses a specific verb 'Explain' and specifies the resource 'GHED variables'. This purpose is distinct from all sibling tools, which do not offer explanatory guidance on variable organization.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to use the tool: when needing to understand variable organization and series selection. While it lacks explicit 'when not to use' statements, the context is clear enough that an agent can infer appropriate usage, especially given the unique explanatory role among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rank_country_changesARead-only
Rank countries by change; use indicator_trend for unranked summaries.
metric='percent_change' and 'cagr' use fractions (0.10 = 10%), not percentage units. 'absolute_change' uses percentage points for shares, otherwise the indicator's original units. Compare actual first/latest years and change_units.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| income | No | ||
| metric | No | absolute_change | |
| region | No | ||
| year_end | No | ||
| countries | No | ||
| descending | No | ||
| year_start | No | ||
| country_group | No | ||
| indicator_code | Yes | ||
| min_year_count | No | ||
| min_period_years | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description adds unit and comparison details but doesn't disclose additional behavioral traits like pagination or data completeness. The description adds some value beyond annotations but is not extensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with the purpose front-loaded and unit details in a separate paragraph. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 12 parameters with 0% schema coverage and an output schema, the description is insufficient to guide correct usage. It does not explain filtering parameters, ordering, or the meaning of the change units in relation to the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains metric units but leaves 11 other parameters unexplained. The parameter names like min_year_count and min_period_years are ambiguous without description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verb 'rank' and resource 'countries by change', and explicitly differentiates from indicator_trend for unranked summaries. The purpose is clear and distinguishes from siblings, though 'change' could be more precise about the time range.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names alternative tool for unranked summaries, and provides unit conventions. Gives clear context for when to use this tool vs indicator_trend, and clarifies metric units.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_cacheAIdempotent
Download or re-download the public GHED workbook and rebuild SQLite.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate idempotence and non-destructiveness. The description adds the specific action (download, rebuild) but lacks details on potential side effects like network usage, time cost, or impact on concurrent operations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no redundancy. The description is concise and front-loads the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and the presence of an output schema, the description sufficiently covers the tool's purpose and behavior. No additional context is necessary for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so baseline is 4. The description does not need to add meaning beyond what the schema provides, and it correctly omits unnecessary detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool downloads/re-downloads the GHED workbook and rebuilds SQLite, specifying the exact action and resource. This distinguishes it from sibling tools like cache_status or check_for_updates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not mention prerequisites, recommended context, or situations to avoid, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_use_casesBRead-only
Research patterns seen in GHED-using literature, with recommended variables.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint. Description adds context about content (GHED-using literature) but no behavioral details beyond annotations. Minimal additional value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, very concise. Front-loads purpose. No unnecessary words. Appropriate for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists but not visible. Description mentions patterns and recommended variables but no detail on output structure or format. For a tool with no parameters and existing output schema, slightly more context on what to expect would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters, so schema coverage is 100%. Baseline is 4. Description doesn't mention zero parameters, but it's not necessary. No loss of meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states tool researches patterns in GHED-using literature and gives recommended variables. Verb 'research' is somewhat vague but resource 'GHED-using literature' is specific, and it distinguishes from siblings like 'suggest_variables_for_research_question' which focuses on a specific question.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives. Sibling list includes many research-oriented tools, but no comparison or context is provided. Agent must infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_country_group_membershipARead-only
Resolve a curated group code to its ISO3 member list.
Accepts canonical codes (LAC, SSA, LDC, OECD, …), official WB region codes (LCN, SSF, …), and common spellings ("Latin America and Caribbean", "Sub-Saharan Africa", "Least Developed Countries").
| Name | Required | Description | Default |
|---|---|---|---|
| group | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate the tool is read-only and non-exhaustive. The description adds behavioral context by specifying acceptable input formats (canonical codes, WB region codes, common spellings), which aids correct invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first defines the primary function, the second lists acceptable inputs. Every sentence adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description adequately covers the input parameter and acceptable formats. Despite the presence of an output schema, it lacks mention of error handling for invalid inputs, but overall it is sufficient for a straightforward resolution tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description fully compensates by explaining the 'group' parameter accepts canonical codes, WB region codes, and common spellings, providing concrete examples. This adds significant meaning beyond the schema's bare string type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool resolves curated group codes to their ISO3 member lists. It distinguishes itself from sibling tools like list_country_groups and summarize_country_group by focusing on resolving codes to members, not listing or summarizing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly tells when to use the tool: when you have a group code or common spelling and need the ISO3 member list. It could be improved by explicitly stating when not to use it (e.g., for non-curated groups), but the context makes it clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_indicatorsARead-only
Search headline indicators; use search_variables for detailed SHA series.
This is substring search, not semantic search. Use a short fragment such as 'out-of-pocket', 'GGE' or 'che_gdp', not a full research question. If no match, shorten the query or use topics_index. category_1=None searches all variables.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| query | Yes | ||
| category_1 | No | INDICATORS | |
| category_2 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description discloses the substring matching behavior, the need for short fragments, the no-match fallback strategy, and the special category_1=None behavior. This is substantial behavioral context that helps an agent know what to expect during a search.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and sibling differentiation, followed by concise, high-value behavioral and fallback guidance. Every sentence adds operational value without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with an output schema and read-only annotations, the description covers the main usage pitfalls: substring vs semantic search, query format, fallback routes, and the category_1=None behavior. It is slightly incomplete on the semantics of top and category_2, but the schema defaults and titles make those parameters reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for query (substring fragments with examples) and category_1 (None searches all variables), but it does not explain top or category_2. Partial compensation only, so the score is mid-range.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Search headline indicators') and immediately differentiates itself from search_variables for detailed SHA series. It also clarifies that this is substring, not semantic, search, which removes ambiguity about what the tool actually does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use this tool vs search_variables, and when to fall back to topics_index if no match. It also gives concrete query-shaping guidance (short fragments, not full research questions), making the invocation context unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_variablesARead-only
Search all Codebook variables; use search_indicators for headline measures.
This is substring search. Use a short code/name fragment, not a full question. An empty result can mean the wording did not match; shorten the query or use topics_index before concluding that a variable is unavailable.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| query | Yes | ||
| category_1 | No | ||
| category_2 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds behavioral detail: it's substring search, advises short fragments, and explains that empty results may indicate a wording mismatch rather than unavailability. This goes beyond the annotations and helps the agent interpret outcomes correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the primary purpose and immediate alternative. The guidance is efficient without excess wording. It could be slightly more structured (e.g., bullet points), but it's still very readable and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers usage alternatives and query behavior, and the presence of an output schema covers return value documentation. However, it omits the meaning and usage of category_1/category_2, and 'top' pagination is not explained. For a tool with four parameters and no schema descriptions, more detail is needed for full context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It partially explains the 'query' parameter (short fragment, not full question) but says nothing about 'top', 'category_1', or 'category_2'. These are left undocumented, and the description does not bridge that gap for the majority of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Search all Codebook variables' – a clear verb and resource. It also distinguishes itself from search_indicators by directing to that tool for headline measures, so the agent can tell them apart. It doesn't explicitly differentiate from list_variables or other search-like tools, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'use search_indicators for headline measures' – providing a clear alternative and when to use it. It also advises using a short fragment and suggests topics_index as a fallback when results are empty. This gives concrete guidance on when to use this tool versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_variables_for_research_questionARead-only
Map a natural-language research question to likely GHED variables and cautions.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and openWorldHint=true. The description adds value by mentioning 'cautions', suggesting the tool also provides warnings or limitations, which extends behavioral context beyond structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that efficiently communicates action, input, and output. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Has an output schema, which reduces the burden on description for return values. Description covers the core purpose and adds 'cautions'. Could elaborate on what cautions entail, but sufficient given output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single parameter 'question', so the description must compensate. It clarifies the parameter is a 'natural-language research question', but provides no format guidance or examples. Adequate but not enhanced.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Map' and specifies the resource 'natural-language research question' and output 'likely GHED variables and cautions'. It clearly distinguishes from siblings like search_variables that perform string matching rather than conceptual mapping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for research questions but does not explicitly state when to use this tool versus alternatives like search_variables or list_variables. No when-not or exclusion criteria provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_country_groupBRead-only
Summarize an indicator within a country group (curated, regional, or income).
Pass at least one of country_group ("LAC", "OECD", "LDC", …),
region, or income.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | ||
| top_n | No | ||
| income | No | ||
| region | No | ||
| latest_only | No | ||
| country_group | No | ||
| indicator_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description adds no extra behavioral context like side effects or rate limits. Acceptable given annotation coverage, but no added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-loading purpose and usage tip. No fluff, but could be more structured with parameter list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters and 0% schema coverage, the description only covers grouping selection. Missing explanation of other parameters and summary behavior, though output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%; the description only hints at three grouping parameters (country_group, region, income) with examples. Other parameters (indicator_code, year, latest_only, top_n) are not explained, failing to compensate for low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool summarizes an indicator within a country group, specifying types (curated, regional, income). It distinguishes from siblings like compare_country_group, but could be more explicit about the summary's nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a usage tip to pass at least one grouping parameter with examples, but lacks when-to-use guidance or alternatives. Minimal but not comprehensive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
topics_indexARead-only
Curated GHED topic index mapping common user requests to variable codes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, covering safety. The description adds that it is 'curated' and involves 'mapping', which provides context but does not elaborate on behavioral traits like static nature or update frequency. It is adequate given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise, front-loaded with the key term 'Curated', and contains no unnecessary words. It efficiently conveys the tool's core function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, annotated, output schema exists), the description is mostly complete. It explains the purpose but could briefly mention what the output contains or how the mapping is structured, though the output schema likely covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the schema has 100% coverage. The baseline for zero-parameter tools is 4, and the description does not need to add parameter details. It correctly omits irrelevant information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it's a curated index mapping user requests to variable codes. It specifies the resource (GHED topic index) and the action (mapping). However, it does not explicitly differentiate from sibling tools like search_indicators or list_variables, but the unique mapping purpose is implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives is provided. With many sibling tools for searching and listing, the description should indicate that this tool is for translating user requests directly to codes, rather than general search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
versionARead-only
Return workbook version lines and cache provenance.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint. Description adds 'cache provenance', giving more detail but no behavioral traits beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single clear sentence with no wasted words. Perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has zero parameters and output schema exists, so description adequately covers what it returns. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters, so baseline 4. Description adds no parameter info but not needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it returns workbook version lines and cache provenance. Distinct from sibling tools like cache_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use versus alternatives like cache_status or refresh_cache.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.6.1- Changed
compare_trends2 fields changed- added
Input schema / properties / min_period_yearsAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Min Period Years" +} - added
Input schema / properties / min_year_countAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Min Year Count" +}
- Changed
search_indicators2 fields changed- added
Input schema / properties / category_1 / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - removed
Input schema / properties / category_1 / typeRemoved value: -"string"
35 tool updates
v0.5.1- First observed
additive_hierarchy - First observed
assess_data_quality - First observed
build_additive_breakdown - First observed
build_research_package - First observed
build_research_panel - First observed
cache_status - First observed
check_for_updates - First observed
compare_countries - First observed
compare_country_group - First observed
compare_trends - First observed
country_profile - First observed
data_availability - First observed
explain_indicator_relationship - First observed
find_country_code - First observed
get_country_metadata - First observed
get_indicator_data - First observed
get_indicator_metadata - First observed
indicator_trend - First observed
list_countries - First observed
list_country_groups - First observed
list_curated_country_groups - First observed
list_indicators - First observed
list_variable_categories - First observed
list_variables - First observed
methodology_guide - First observed
rank_country_changes - First observed
refresh_cache - First observed
research_use_cases - First observed
resolve_country_group_membership - First observed
search_indicators - First observed
search_variables - First observed
suggest_variables_for_research_question - First observed
summarize_country_group - First observed
topics_index - First observed
version
TDQS
Scored across 35 tools
Multiple tools have heavily overlapping boundaries: compare_countries, compare_country_group, summarize_country_group, and get_indicator_data all retrieve one indicator across country sets, and compare_countries also accepts country_group, making the distinction unclear. Search and list tools similarly overlap across indicators vs. variables, and data quality/metadata tools are hard to separate without deep inspection.
The set mostly uses verb_noun names like list_variables, get_indicator_data, and build_research_panel, but a substantial minority are noun-only commands such as country_profile, version, methodology_guide, topics_index, indicator_trend, and cache_status. The conventions are mixed but still readable and not chaotic.
At 35 tools, the server is well beyond the 25+ threshold and feels over-granular for its domain. Several tools could be consolidated, especially the country-comparison trio and the search/list variants, without losing real capability.
The tool surface is unusually complete for the GHED domain: it covers discovery, metadata, data availability, country and group resolution, single-indicator retrieval, cross-country comparison, trend analysis, research panel construction, export packaging, and cache/version management. There are no obvious dead ends or missing operations for a read-only database server.
Maintenance
Related MCP Connectors
Hosted MCP server exposing US hospital procedure cost data to AI assistants
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
Hosted MCP server for Cliniko — patients, appointments, availability, and invoices for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol (MCP) server that gives Claude access to your WHOOP biometric data — recovery, sleep, strain, and workouts.33 npmMIT
- FlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that connects Claude to your Google Sheets for personal finance tracking, analysis, and reporting — all through natural language.2-
- AlicenseNot gradedqualityAmaintenanceA Model Context Protocol (MCP) server that brings your Withings health data into Claude, allowing natural conversation access to sleep patterns, body measurements, workouts, heart data, and more.42MIT
- AlicenseAqualityAmaintenanceA Model Context Protocol (MCP) server that gives AI assistants direct access to the World Health Organization's Global Health Observatory (GHO) for comparative health systems research.15MIT