Skip to main content
Glama

Harmonize Free-Text Terms to Standard Codes

harmonize_terms
Read-onlyIdempotent

Map lists of free-text clinical terms to standard ICD-11, RxNorm/ATC, or LOINC codes, returning ranked candidates and confidence labels for review.

Instructions

Map a LIST of free-text clinical terms to standard codes in one call, with ranked candidates and a confidence label for each — the building block of a reviewable crosswalk.

Use this tool to:

  • Harmonize a column of diagnoses, drugs or lab names from a dataset to ICD-11 / RxNorm (+ ATC) / LOINC

  • Triage which terms map cleanly (exact / strong) and which need a person (needs_review)

  • Build a crosswalk you can audit: every row keeps its candidates, scores and sources

Give each term its domain: diagnosis → ICD-11; drug → RxNorm concepts (ingredients first) plus the ATC classes of the term; lab → LOINC. Up to 50 terms per call — a longer list is refused with a validation error: split it into batches of 50. Repeated term+domain pairs are looked up once. max_candidates keeps 1-5 per term (default 3).

Every candidate carries match_score (lexical, 0-1, the find_equivalent formula) and match_type: exact = same words after normalization; strong = every term word is in the title (or the matched synonym) and score ≥ 0.85; needs_review = anything else. A one-word term is exact or needs_review, never strong ("Tylenol" vs "Tylenol PM" is a different product). Synonyms, abbreviations ("MI", "HbA1c") and misspellings land in needs_review or no_candidates — the label errs toward asking a person. Candidates sharing no word with the term are dropped, as are LOINC codes named "Deprecated". For diagnoses, a candidate is also scored against the synonyms WHO matched (e.g. "hypertension NOS" for Essential hypertension), reported in matched_label; postcoordinated clusters (codes with "/" or "&") are left out — build those with icd11_postcoordination. Lab names are ambiguous without specimen and property: "glucose" matches over a thousand LOINC codes, so write "glucose serum" or expect needs_review. One failed lookup does not fail the batch: that row comes back with status "error".

Terms are searched in English and sent to the WHO and NLM APIs — de-identify the list first. For Brazilian Portuguese diagnoses use cid10_search; to check codes you already have, use validate_codes; for one term across every terminology, use find_equivalent. Record the vocabulary versions with the provenance blocks (one per source) and terminology_versions.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
termsYesTerms to harmonize, each with its domain. Up to 50 per call; repeated term+domain pairs are looked up once.
max_candidatesNoCandidates kept per term, best first (1-5, default 3).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
totalYesTerms submitted.
countsYes
rankingYes
resultsYesOne entry per submitted term, in request order.
provenanceYesOne provenance block per upstream source that contributed to this response (contract v1.1; licenses are never merged; each block carries the origin diagnostics of ITS source)
attributionYesCanonical source URLs of this response (attribution list)
unique_lookupsYesDistinct term+domain pairs actually looked up.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv2.0.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, open-world and non-destructive, but the description adds substantial behavior beyond them: 50-term hard limit with a validation error, dedup of repeated term+domain pairs, per-row error isolation ('one failed lookup does not fail the batch'), needs_review bias, dropped deprecated LOINC codes, and a de-identification requirement before hitting WHO/NLM APIs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line summary, then structured bullets for usage, then behavioral rules. Dense but nearly every sentence carries a rule; minor redundancy in restating the 50-term cap twice and the domain mapping already in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity, nested term objects, and presence of an output schema, the description covers everything an agent needs: batch limits, error semantics, match_type scoring rules, why 'glucose' alone fails, and pointer to provenance tools (terminology_versions, per-source provenance blocks). Return values are correctly left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the description mostly restates the domain enum mapping that the schema already documents. It does add value on the terms parameter (batching at 50, dedup behavior, English-only search) and on max_candidates (default 3, best-first ordering), but this is elaboration rather than information the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (harmonize/map) plus resource (list of free-text clinical terms) plus output shape (ranked candidates with confidence labels). It explicitly positions itself as the batch/building-block counterpart to find_equivalent, so an agent can distinguish it from siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Has an explicit 'Use this tool to' list covering the three concrete scenarios (column harmonization, triage, auditable crosswalk), and names when to use alternatives instead: cid10_search for pt-BR diagnoses, validate_codes for existing codes, find_equivalent for one term across all terminologies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.