Harmonize Free-Text Terms to Standard Codes
harmonize_termsMap lists of free-text clinical terms to standard ICD-11, RxNorm/ATC, or LOINC codes, returning ranked candidates and confidence labels for review.
Instructions
Map a LIST of free-text clinical terms to standard codes in one call, with ranked candidates and a confidence label for each — the building block of a reviewable crosswalk.
Use this tool to:
Harmonize a column of diagnoses, drugs or lab names from a dataset to ICD-11 / RxNorm (+ ATC) / LOINC
Triage which terms map cleanly (exact / strong) and which need a person (needs_review)
Build a crosswalk you can audit: every row keeps its candidates, scores and sources
Give each term its domain: diagnosis → ICD-11; drug → RxNorm concepts (ingredients first) plus the ATC classes of the term; lab → LOINC. Up to 50 terms per call — a longer list is refused with a validation error: split it into batches of 50. Repeated term+domain pairs are looked up once. max_candidates keeps 1-5 per term (default 3).
Every candidate carries match_score (lexical, 0-1, the find_equivalent formula) and match_type: exact = same words after normalization; strong = every term word is in the title (or the matched synonym) and score ≥ 0.85; needs_review = anything else. A one-word term is exact or needs_review, never strong ("Tylenol" vs "Tylenol PM" is a different product). Synonyms, abbreviations ("MI", "HbA1c") and misspellings land in needs_review or no_candidates — the label errs toward asking a person. Candidates sharing no word with the term are dropped, as are LOINC codes named "Deprecated". For diagnoses, a candidate is also scored against the synonyms WHO matched (e.g. "hypertension NOS" for Essential hypertension), reported in matched_label; postcoordinated clusters (codes with "/" or "&") are left out — build those with icd11_postcoordination. Lab names are ambiguous without specimen and property: "glucose" matches over a thousand LOINC codes, so write "glucose serum" or expect needs_review. One failed lookup does not fail the batch: that row comes back with status "error".
Terms are searched in English and sent to the WHO and NLM APIs — de-identify the list first. For Brazilian Portuguese diagnoses use cid10_search; to check codes you already have, use validate_codes; for one term across every terminology, use find_equivalent. Record the vocabulary versions with the provenance blocks (one per source) and terminology_versions.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| terms | Yes | Terms to harmonize, each with its domain. Up to 50 per call; repeated term+domain pairs are looked up once. | |
| max_candidates | No | Candidates kept per term, best first (1-5, default 3). |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| total | Yes | Terms submitted. | |
| counts | Yes | ||
| ranking | Yes | ||
| results | Yes | One entry per submitted term, in request order. | |
| provenance | Yes | One provenance block per upstream source that contributed to this response (contract v1.1; licenses are never merged; each block carries the origin diagnostics of ITS source) | |
| attribution | Yes | Canonical source URLs of this response (attribution list) | |
| unique_lookups | Yes | Distinct term+domain pairs actually looked up. |