<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Digit. Health</journal-id>
<journal-title>Frontiers in Digital Health</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Digit. Health</abbrev-journal-title>
<issn pub-type="epub">2673-253X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdgth.2025.1495040</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Digital Health</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A simplified retriever to improve accuracy of phenotype normalizations by large language models</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes"><name><surname>Hier</surname><given-names>Daniel B.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref><uri xlink:href="https://loop.frontiersin.org/people/1230804/overview"/><role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/><role content-type="https://credit.niso.org/contributor-roles/data-curation/"/><role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/><role content-type="https://credit.niso.org/contributor-roles/methodology/"/><role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/><role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/></contrib>
<contrib contrib-type="author"><name><surname>Do</surname><given-names>Thanh Son</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/2840956/overview" /><role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/><role content-type="https://credit.niso.org/contributor-roles/validation/"/><role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/><role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/></contrib>
<contrib contrib-type="author"><name><surname>Obafemi-Ajayi</surname><given-names>Tayo</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/1289170/overview" /><role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/><role content-type="https://credit.niso.org/contributor-roles/methodology/"/><role content-type="https://credit.niso.org/contributor-roles/validation/"/><role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/><role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/></contrib>
</contrib-group>
<aff id="aff1"><label><sup>1</sup></label><institution>Department of Neurology and Rehabilitation, University of Illinois at Chicago</institution>, <addr-line>Chicago, IL</addr-line>, <country>United States</country></aff>
<aff id="aff2"><label><sup>2</sup></label><institution>Department of Computer Science, Missouri State University</institution>, <addr-line>Springfield, MO</addr-line>, <country>United States</country></aff>
<aff id="aff3"><label><sup>3</sup></label><institution>Engineering Program, Missouri State University</institution>, <addr-line>Springfield, MO</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p><bold>Edited by:</bold> Phoey Lee Teh, Glynd&#x0175;r University, United Kingdom</p></fn>
<fn fn-type="edited-by"><p><bold>Reviewed by:</bold> Elena Cardillo, National Research Council (CNR), Italy</p>
<p>Daisy Monika Lal, Lancaster University, United Kingdom</p></fn>
<corresp id="cor1"><label>&#x002A;</label><bold>Correspondence:</bold> Daniel B. Hier <email>dhier@uic.edu</email></corresp>
</author-notes>
<pub-date pub-type="epub"><day>04</day><month>03</month><year>2025</year></pub-date>
<pub-date pub-type="collection"><year>2025</year></pub-date>
<volume>7</volume><elocation-id>1495040</elocation-id>
<history>
<date date-type="received"><day>11</day><month>09</month><year>2024</year></date>
<date date-type="accepted"><day>12</day><month>02</month><year>2025</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 Hier, Do and Obafemi-Ajayi.</copyright-statement>
<copyright-year>2025</copyright-year><copyright-holder>Hier, Do and Obafemi-Ajayi</copyright-holder><license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Large language models have shown improved accuracy in phenotype term normalization tasks when augmented with retrievers that suggest candidate normalizations based on term definitions. In this work, we introduce a simplified retriever that enhances large language model accuracy by searching the Human Phenotype Ontology (HPO) for candidate matches using contextual word embeddings from BioBERT without the need for explicit term definitions. Testing this method on terms derived from the clinical synopses of Online Mendelian Inheritance in Man (OMIM<sup>&#x00AE;</sup>), we demonstrate that the normalization accuracy of GPT-4o increases from a baseline of 62&#x0025; without augmentation to 85&#x0025; with retriever augmentation. This approach is potentially generalizable to other biomedical term normalization tasks and offers an efficient alternative to more complex retrieval methods.</p>
</abstract>
<kwd-group>
<kwd>phenotype normalization</kwd>
<kwd>large language model</kwd>
<kwd>small language model</kwd>
<kwd>cosine similarity</kwd>
<kwd>HPO</kwd>
<kwd>OMIM</kwd>
<kwd>retrievalaugmented generation</kwd>
</kwd-group><counts>
<fig-count count="1"/>
<table-count count="2"/><equation-count count="4"/><ref-count count="38"/><page-count count="7"/><word-count count="0"/></counts><custom-meta-wrap><custom-meta><meta-name>section-at-acceptance</meta-name><meta-value>Health Informatics</meta-value></custom-meta></custom-meta-wrap>
</article-meta>
</front>
<body><sec id="s1" sec-type="intro"><title>Introduction</title>
<p>Large pre-trained language models are increasingly used in healthcare care, showing promise in performing a variety of complex natural language processing (NLP) tasks, such as text summarization, concept recognition, and answer questions (<xref ref-type="bibr" rid="B1">1</xref>&#x2013;<xref ref-type="bibr" rid="B3">3</xref>). Large language models can identify medical concepts in the text and normalize them to an ontology (<xref ref-type="bibr" rid="B4">4</xref>&#x2013;<xref ref-type="bibr" rid="B7">7</xref>). However, when large language models normalize medical terms to a standard ontology such as the human phenotype ontology (HPO), the retrieved code is not always accurate.</p>
<table-wrap id="T2" position="float"><label>Table 2</label>
<caption><p>GPT-3.5-Turbo may select an HPO term as semantically equivalent that did not have the highest cosine similarity by BioBERT word embeddings.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="center"/>
<col align="left"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Term to normalize</th>
<th valign="top" align="center">BioBERT match by CS</th>
<th valign="top" align="center">CS</th>
<th valign="top" align="center">Large language model + retriever match</th>
<th valign="top" align="center">CS</th>
<th valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM1"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Absent ankle jerks</td>
<td valign="top" align="left">Absent knee jerk reflex</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="left">Absent ankle reflexes</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.01</td>
</tr>
<tr>
<td valign="top" align="left">Pale fundi</td>
<td valign="top" align="left">Pale eyelashes</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="left">Depigmented fundus</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.01</td>
</tr>
<tr>
<td valign="top" align="left">Lack of speech</td>
<td valign="top" align="left">Poor speech discrimination</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="left">Absent speech development</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.01</td>
</tr>
<tr>
<td valign="top" align="left">Disinhibition</td>
<td valign="top" align="left">Inactivity</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="left">Social disinhibition</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.01</td>
</tr>
<tr>
<td valign="top" align="left">Hand weakness</td>
<td valign="top" align="left">Shoulder weakness</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="left">Hand muscle weakness</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.01</td>
</tr>
<tr>
<td valign="top" align="left">Bilateral foot drop</td>
<td valign="top" align="left">Bilateral clubfoot</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="left">Foot drop</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.01</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn1"><p><italic>Note:</italic> The table shows the cosine similarity (CS) between the &#x201C;term to normalize&#x201D; and the term choice by the BioBERT method and the term choice by the large language model + Retriever. Using the BioBERT method, the term in the HPO with maximal CS to the &#x201C;term to normalize&#x201D; is chosen. The large language model + Retriever method selects from 20 candidate terms the the term with best &#x201C;semantic equivalence&#x201D; while ignoring cosine similarity. As shown in this Table, in some cases large language model + Retriever can outperform the BioBERT method due to its ability to pick terms that are better matches to the &#x201C;term to normalize&#x201D; that do not have the highest cosine similarities. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM2"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> is the difference between the cosine similarities for the term selected for each method.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Shlyk et al. (<xref ref-type="bibr" rid="B8">8</xref>) demonstrated that the accuracy of large language models in term normalization tasks can be improved with Retrieval-Augmented Entity Linking (REAL). This method generates definitions of HPO terms and target terms needing normalization. The definitions are converted into word embeddings, and cosine similarity is used to identify the three closest candidate terms. The large language model then selects the best normalization from these candidates.</p>
<p>In this work, we introduce a simplified but effective retriever that bypasses the need for definition generation. Instead, it matches HPO terms to target terms using BioBERT contextual word embeddings, identifying the closest matches by semantic similarity. By prompting GPT-4o with the 20 closest candidate terms, we enable the model to take advantage of its implicit knowledge of HPO terms and select a semantically equivalent normalization. This approach achieves accuracy comparable to more complex methods without using explicit definitions.</p>
<p>Phenotyping, which involves recognizing the signs and symptoms of disease in patients and mapping them to an appropriate ontology such as HPO, is critical for precision medicine (<xref ref-type="bibr" rid="B9">9</xref>&#x2013;<xref ref-type="bibr" rid="B13">13</xref>). Manual phenotyping is labor intensive, which makes high-throughput automated methods essential (<xref ref-type="bibr" rid="B11">11</xref>, <xref ref-type="bibr" rid="B14">14</xref>&#x2013;<xref ref-type="bibr" rid="B18">18</xref>). Phenotyping can be seen as part of the broader task of term normalization in different vocabularies, such as drugs to RXNORM (<xref ref-type="bibr" rid="B19">19</xref>), diseases to ICD-11 (<xref ref-type="bibr" rid="B20">20</xref>), or laboratory tests to LOINC (<xref ref-type="bibr" rid="B21">21</xref>).</p>
<p>A distinction can be made between surface phenotyping, which assigns a diagnosis from a disease ontology (<xref ref-type="bibr" rid="B22">22</xref>), and deep phenotyping, which assigns an HPO concept and ID to each symptom (<xref ref-type="bibr" rid="B11">11</xref>). Our work focuses on deep phenotyping (<xref ref-type="bibr" rid="B17">17</xref>).</p>
<p>Although often performed together, concept extraction (identification) and concept normalization are distinct NLP tasks. In concept extraction, the goal is to find relevant medical concepts within free text. Concept normalization involves matching irregular medical terms to standardized terms in an ontology and their corresponding machine-codes. If the irregular term to normalize has an exact match in the ontology, the task is straightforward such a matching &#x201C;hyporeflexia&#x201D; to its standard form in the HPO which is &#x201C;Hyporeflexia.&#x201D; If the irregular term has no exact match (e.g., &#x201C;diminished reflexes&#x201D;), the model must select the most semantically similar concept (e.g., &#x201C;decreased reflexes&#x201D;). Additionally, concept normalization involves mapping irregular terms to their standard forms in an ontology and their corresponding machine codes (&#x201C;hyporeflexia,&#x201D; HP:0001265).</p>
<p>Advances in deep learning, transformer architectures, and dynamic word embeddings have facilitated the development of tools that perform concept identification and normalization such as Doc2Hpo, ClinPheno, BERN2, PhenoBERT, and FastHPOCR (<xref ref-type="bibr" rid="B23">23</xref>&#x2013;<xref ref-type="bibr" rid="B30">30</xref>). Evaluations of concept recognition tools show F1 values ranging from 0.50 to 0.76, with FastHPOCR performing the best on manually annotated corpora (<xref ref-type="bibr" rid="B29">29</xref>, <xref ref-type="bibr" rid="B31">31</xref>).</p>
<p>Although deep learning and transformer-based methods have shown promise, they require extensive training data, which can be time consuming to acquire. Pre-trained large language models offer an alternative approach by eliminating the need for new training data. Despite their impressive performance, large language models may still make errors in retrieving the correct HPO ID (<xref ref-type="bibr" rid="B32">32</xref>).</p>
<p>Retrieval-augmented generation (RAG) (<xref ref-type="bibr" rid="B33">33</xref>) addresses this problem by using a retriever to provide relevant information, improving the likelihood of generating accurate output (<xref ref-type="bibr" rid="B34">34</xref>). In this paper, we extend the work of Shlyk et al. (<xref ref-type="bibr" rid="B8">8</xref>) by demonstrating that a simplified retriever, which relies on embedded terms, can improve normalization accuracy without the need to generate embedded term definitions. By prompting a large language model with candidate terms that have similarity to the target term, we achieve high accuracy with a more efficient approach.</p>
</sec>
<sec id="s2" sec-type="methods"><title>Methods</title>
<sec id="s2a"><title>Experimental plan</title>
<p>We selected 1,820 phenotypic terms from the Clinical Features sections of OMIM summaries as a test set for term normalization. In the first experimental condition, the NLP models (spaCy and BioBERT) normalized the terms by selecting the best-matching HPO term and HPO ID based on the cosine similarity of the embedded word vectors. In the second experimental condition, language models were prompted to normalize a term via the OpenAI API to the best matching HPO term and HPO ID. In the third experimental condition, the prompts for the language models were augmented with up to 50 candidate terms and HPO IDs generated based on the cosine similarity between the BioBERT word embeddings and the term to be normalized. For each experimental condition, we calculated accuracy, F1, recall, and precision of term normalization.</p>
</sec>
<sec id="s2b"><title>Data</title>
<p>Terms to normalize were the signs and symptoms of neurogenetic diseases derived from <italic>Clinical Feature</italic> summaries in the OMIM database (<xref ref-type="bibr" rid="B35">35</xref>). We downloaded clinical feature summaries for 236 neurogenetic diseases, consisting of 175,724 tokens via the OMIM API (<ext-link ext-link-type="uri" xlink:href="https://api.omim.org">https://api.omim.org</ext-link>). GPT-3.5-Turbo was used to identify 2,023 terms for normalization (mean 16.5 signs per disease) (<xref ref-type="bibr" rid="B18">18</xref>, <xref ref-type="bibr" rid="B36">36</xref>). The text for extracting signs and symptoms was passed to the GPT-3.5 Turbo API with the following prompt:
</p>
<p><monospace>prompt =</monospace></p>
<p><monospace>(You are a neurologist analyzing a case summary.</monospace></p>
<p><monospace>The input is a JSON object containing:</monospace></p>
<p><monospace>&#x2018;&#x2018;clinical Features&#x2019;&#x2019;</monospace></p>
<p><monospace>Your task is to extract all relevant</monospace></p>
<p><monospace>neurological symptoms (patient complaints) and signs (findings on examination).</monospace></p>
<p><monospace>Exclude any signs and symptoms related to family members.</monospace></p>
<p><monospace>Please respond with the findings organized into</monospace></p>
<p><monospace>a dictionary under the key &#x2018;&#x2018;Signs.&#x2019;&#x2019;</monospace></p>
<p><monospace>Each sign should be distinctly listed.</monospace></p>
<p><monospace>Here is the format for your response:</monospace></p>
<p><monospace>{</monospace></p>
<p><monospace>&#x2018;&#x2018;Signs&#x2019;&#x2019;: [&#x2018;&#x2018;sign a,&#x2019;&#x2019; &#x2018;&#x2018;sign b,&#x2019;&#x2019; &#x2018;&#x2018;sign c&#x2019;&#x2019;]</monospace></p>
<p><monospace>}</monospace></p>
<p><monospace>Report only signs and symptoms observable by</monospace></p>
<p><monospace>the physician at the bedside.</monospace></p>
<p><monospace>Ignore all laboratory, pathological,</monospace></p>
<p><monospace>and radiological signs)</monospace>
</p>
<p>A domain expert excluded 203 malformed terms (e.g., vague, contradictory, verbose, or ambiguous phrases). These terms were excluded because they were judged to be difficult to normalize. Examples of malformed terms that would be difficult to normalize included:</p>
<p><monospace>impaired visual pathways</monospace></p>
<p><monospace>initial good response to dopaminergic therapy</monospace></p>
<p><monospace>intermittent microsaccadic pursuits</monospace></p>
<p><monospace>intermittent mobility</monospace></p>
<p><monospace>intermittent tetanic contraction</monospace></p>
<p><monospace>intrusive square wave jerks</monospace></p>
<p><monospace>jerky voice</monospace></p>
<p><monospace>kineto rigid syndrome</monospace></p>
<p><monospace>legs and arms</monospace>
</p>
<p>The final test dataset consisted of 1,820 terms to normalize.</p>
<p>The Human Phenotype Ontology (HPO) was downloaded as a comma-separated value (CSV) file from NCBO BioPortal (<xref ref-type="bibr" rid="B37">37</xref>). A list of 17,957 HPO entry terms was expanded to 30,234 by adding all available synonyms. Each HPO entry term was associated with a corresponding HPO ID, formatted as <monospace>HP:nnnnnnn</monospace>, where <monospace>n</monospace> is a digit between 0 and 9.</p>
</sec>
<sec id="s2c"><title>Term normalization using NLP-based methods</title>
<p><bold>spaCy</bold>: SpaCy was combined with <monospace>en&#x005F;core&#x005F;web&#x005F;lg</monospace> word embeddings. Vectors were generated for each HPO entry term and stored as a Python dictionary.</p>
<p><bold>BioBERT</bold>: We utilized the BioBERT v1.1 model (<monospace>dmis-lab/biobert-base-cased-v1.1</monospace>) to compute embeddings for target terms and HPO entry terms (<xref ref-type="bibr" rid="B38">38</xref>). Each term (target or HPO term) was tokenized using the BioBERT tokenizer. The resulting embeddings were computed using the BioBERT transformer model, and the mean of the token embeddings across all tokens was used as the global embedding vector:
</p>
<p><monospace>inputs = tokenizer(term, return_tensors=&#x2018;&#x2018;pt,&#x2019;&#x2019; truncation=True, padding=True, max_length=128)</monospace></p>
<p><monospace>with torch.no_grad():</monospace></p>
<p><monospace>outputs = model(**inputs)</monospace></p>
<p><monospace>embedding = outputs.last_hidden_state.mean(dim=1).squeeze().numpy()</monospace>
</p>
<p>HPO entry terms and their corresponding IDs were preprocessed, and their embeddings were calculated and stored in a CSV file for efficiency. For each target term, cosine similarity between its BioBERT embedding and the precomputed HPO embeddings was calculated. The HPO term with the highest similarity score was selected as the &#x201C;best match&#x201D;:
</p>
<p><monospace>similarities = cosine_similarity(term_vector, hpo_embeddings).flatten()</monospace></p>
<p><monospace>best_match_idx = np.argmax(similarities)</monospace>
</p>
<p>The same method was used to retrieve the <bold>k</bold> best matches for inputting candidate terms to the the large language model with retriever methods.</p>
<p><bold>Doc2Hpo</bold>: Terms were normalized using the Doc2Hpo API at <ext-link ext-link-type="uri" xlink:href="https://doc2hpo.wglab.org/parse/acdat">https://doc2hpo.wglab.org/parse/acdat</ext-link> using the string-based matching engine (<xref ref-type="bibr" rid="B25">25</xref>).</p>
</sec>
<sec id="s2d"><title>Term normalization using large language models</title>
<p>Three large language models were evaluated for term normalization: GPT-4o, GPT-3.5-Turbo, and GPT-4o-mini from OpenAI (San Francisco, CA) via its API (<ext-link ext-link-type="uri" xlink:href="https://api.openai.com">https://api.openai.com</ext-link>). Each of the terms to normalize was passed to the API with the following prompt:
</p>
<p><monospace>prompt = (</monospace></p>
<p><monospace>You are given a term to normalize to a concept</monospace></p>
<p><monospace>from the Human Phenotype Ontology and return</monospace></p>
<p><monospace>the best match and its HPO ID.</monospace></p>
<p><monospace>&#x2018;&#x2018;Term: {term}&#x2019;&#x2019;</monospace></p>
<p><monospace>Pick the best one and return it in JSON format:</monospace></p>
<p><monospace>{&#x2018;&#x2018;best_match&#x2019;&#x2019;: &#x2018;&#x2018;term,&#x2019;&#x2019; &#x2018;&#x2018;HPO ID&#x2019;&#x2019;: &#x2018;&#x2018;HP:nnnnnnn&#x2019;&#x2019;}</monospace></p>
</sec>
<sec id="s2e"><title>Term normalization using large language models enhanced with retrieval augmentation</title>
<p>The performance of large language models for term normalization was enhanced by augmenting the prompt with up to 50 candidate terms. The top 20 candidates from the BioBERT embeddings were used in the final analysis.
</p>
<p><monospace>prompt = (</monospace></p>
<p><monospace>You are given a term to normalize to a concept from the Human Phenotype Ontology and its HPO_ID:</monospace></p>
<p><monospace>Term: {term}</monospace></p>
<p><monospace>Possible matches: [match_1&#x2026;match_20]</monospace></p>
<p><monospace>Pick the best one from the above matches and return it in JSON format:</monospace></p>
<p><monospace>{&#x2018;&#x2018;best_match&#x2019;&#x2019;: &#x2018;&#x2018;term,&#x2019;&#x2019; &#x2018;&#x2018;hpo_id&#x2019;&#x2019;: &#x2018;&#x2018;HP:xxxxxxx&#x2019;&#x2019;}</monospace></p>
</sec>
<sec id="s2f"><title>Assessment of semantic equivalence</title>
<p>We evaluated the semantic equivalence of the normalized terms by comparing the &#x201C;best matches&#x201D; to the original <italic>terms to normalize</italic>. Among the 1,820 terms, 438 had exact matches in the list of HPO entry terms. To assess semantic equivalence, we employed three complementary approaches:
<list list-type="simple">
<list-item><label>1.</label>
<p><bold>Cosine Similarity:</bold> We calculated the cosine similarity between the embeddings of the term to normalize and the candidate HPO terms using BioBERT embeddings.</p></list-item>
<list-item><label>2.</label>
<p><bold>GPT-3.5-Turbo Judgment:</bold> GPT-3.5-Turbo was prompted to assess semantic equivalence by returning a binary judgment (equivalent or not) for each input term.</p></list-item>
<list-item><label>3.</label>
<p><bold>Expert Review:</bold> Domain experts in clinical terminology provided the final judgment on semantic equivalence, taking into account the cosine similarity and the GPT-3.5-Turbo binary judgment.</p></list-item>
</list>A term was deemed an accurate match (True Positive, TP) if it was semantically equivalent to the original term and was mapped to the correct HPO ID. A False Positive (FP) occurred when a term was semantically incorrect or the HPO ID was inaccurate. Malformed terms that were previously excluded from normalization attempts were not counted as True Negatives (TN). If a model failed to return any normalization for a given term, it was considered a False Negative (FN). Accuracy, F1, recall, and precision were calculated using standard formulas:
<list list-type="simple">
<list-item>
<p>Accuracy = (TP + TN) / (TP + TN + FP + FN)</p></list-item>
<list-item>
<p>Precision = TP / (TP + FP)</p></list-item>
<list-item>
<p>Recall = TP / (TP + FN)</p></list-item>
<list-item>
<p>F1 = 2 <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM3"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> (Precision <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM4"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> Recall) / (Precision + Recall)</p></list-item>
</list>Metrics were reported to two decimal places, reflecting the precision and reproducibility of the calculations.</p>
</sec>
</sec>
<sec id="s3" sec-type="results"><title>Results</title>
<p><xref ref-type="table" rid="T1">Table&#x00A0;1</xref> shows the model accuracies for the phenotype normalization of the 1,820 <italic>terms to normalize</italic>. To be rated as &#x201C;accurate,&#x201D; the <italic>term to normalize</italic> had to be semantically equivalent to the HPO entry term and and the model had to retrieve the term&#x2019;s correct HPO ID.</p>
<table-wrap id="T1" position="float"><label>Table 1</label>
<caption><p>Model metrics for term normalization to HPO.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Method</th>
<th valign="top" align="center">Accuracy</th>
<th valign="top" align="center">F1</th>
<th valign="top" align="center">Recall</th>
<th valign="top" align="center">Precision</th>
<th valign="top" align="center">N</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">spaCy embeddings cosine similarity</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">BioBERT embeddings by cosine similarity</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">GPT-4o mini</td>
<td valign="top" align="center">0.12</td>
<td valign="top" align="center">0.21</td>
<td valign="top" align="center">0.40</td>
<td valign="top" align="center">0.40</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">GPT-3.5-Turbo</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">GPT-4o</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">Doc2Hpo API</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">GPT-3.5-Turbo with retrieval augmentation</td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">1,820</td>
</tr>
<tr>
<td valign="top" align="left">GPT-4o with retrieval augmentation</td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">1,820</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn2"><p><italic>Note</italic>: The sample size (N) for all methods is 1820. For augmented methods, the language models were presented with a set of 20 candidate terms generated by the retriever. False negatives (FN) were excluded as &#x201C;malformed terms&#x201D; (see Methods).</p></fn>
<fn id="table-fn1a"><p>Bold values show models with highest accuracy on term normalization task.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>The spaCy and BioBERT methods used word embeddings and NLP algorithms to find the best match in a complete table of HPO terms. The spaCy embeddings were general-purpose word embeddings, whereas the BioBERT embeddings were optimized for biomedical terminologies. spaCy averaged the vectors of the component tokens to get a global term vector, whereas BioBERT utilized a transformer architecture and hidden states to generate a global term vector. Both methods used cosine similarities to find the best match for each <italic>term to normalize</italic> and candidate terms in the HPO. The BioBERT method, at 69&#x0025; accuracy, outperformed the spaCy method at 46&#x0025;.</p>
<p>The GPT-4o-mini, GPT-3.5-Turbo, and GPT-4o models without a retriever had no access to external data sources and relied entirely on pre-training to find HPO IDs. Errors made by these language models typically involved the retrieval of incorrect HPO IDs rather than errors in the HPO entry terms. In most cases, when the models made an error, the HPO entry term was correct or nearly correct, but the HPO ID was inaccurate and matched an incorrect concept in the HPO. Among large language models without a retriever, GPT-4o, the largest and most advanced model, performed best with a accuracy of 62&#x0025;, while GPT-4o-mini, the smallest model, performed worse with an accuracy of 12&#x0025;.</p>
<p>The best-performing methods combined a language model with a retriever. Each model had 20 candidate normalizations for a <italic>term to normalize</italic> based on the closest embedding similarities. Each language model was prompted to choose the &#x201C;best match&#x201D; from the twenty closest candidates. In this scenario, the GPT-4o and GPT-3.5-Turbo models outperformed the BioBERT method with 85&#x0025; to 88&#x0025; accuracies (<xref ref-type="table" rid="T1">Table&#x00A0;1</xref>).</p>
<p>We investigated the optimal number of potential matches to submit to GPT-4o or GPT-3.5-Turbo (<xref ref-type="fig" rid="F1">Figure&#x00A0;1</xref>). Accuracy improved as the number of candidate normalization increased to 20 and then reached a plateau with further increases in candidates not improving accuracy of normalization.</p>
<fig id="F1" position="float"><label>Figure 1</label>
<caption><p>RAG candidate choices vs. accuracy. Accuracy was evaluated for normalization by GPT-3.5-Turbo with candidate prompts in the range of 1 to 50. Accuracy improved until reaching a plateau at 20 candidate terms in the prompt.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1495040-g001.tif"/>
</fig>
</sec>
<sec id="s4" sec-type="discussion"><title>Discussion</title>
<p>The results indicate that the most accurate method for phenotype term normalization combines a language model with a retriever. GPT-4o and GPT-3.5-Turbo, when paired with a retriever, achieved the highest accuracies of 85&#x0025; to 88&#x0025;, demonstrating the benefits of augmenting language models with a retrieval mechanism.</p>
<p>The standalone BioBERT, specifically optimized for biomedical text, performed better than spaCy or GPT-4o without retrieval, with an accuracy of 69&#x0025;. This highlights the limitations of a standalone large language model relying solely on pre-training when no external retrieval is available. BioBERT&#x2019;s ability to generate specialized biomedical embeddings allowed it to outperform GPT-3.5-Turbo and GPT-4o without a retriever.</p>
<p>GPT-4o-mini, smaller and less pre-trained than GPT-4o and GPT-3.5-Turbo, showed weak performance on term normalization with 12&#x0025; accuracy, highlighting its lack of exposure to HPO terms and their HPO IDs during training.</p>
<p>Examining the cases where large language model combined with a retriever outperformed BioBERT reveals that the &#x201C;best match&#x201D; selected by large language model can deviate from the candidate term with the highest cosine similarity (<xref ref-type="table" rid="T2">Table 2</xref>). This demonstrates the strength of large language models in interpreting semantic equivalence beyond simple cosine metrics. For example, &#x201C;foot drop&#x201D; was selected as a better semantic match for &#x201C;bilateral foot drop&#x201D; than &#x201C;bilateral clubfoot,&#x201D; despite having a lower cosine similarity score. Similarly, &#x201C;depigmented fundus&#x201D; was preferred for &#x201C;pale fundi&#x201D; over &#x201C;pale eyelashes.&#x201D; The ability of a large language model to override cosine similarity based on contextual understanding explains the superior accuracy of the retriever-augmented method. Part of this superiority likely resides on focusing attention on the semantically most important words in compound terms such as &#x201C;depigmented fundu&#x201D; where the focus is appropriately directed to &#x201C;fundus&#x201D; in favor of the less important word &#x201C;depigmented.&#x201D;</p>
<p>Our results suggest that, while large language models possess significant capabilities, their performance on phenotype normalization tasks can be enhanced by retrieval augmentation. The accuracy of normalization improves when the large language model is presented with candidate terms selected via cosine similarity. The plateau observed for 20 candidate terms in the prompt suggests that presenting more candidates neither degrades nor improves accuracy (<xref ref-type="fig" rid="F1">Figure&#x00A0;1</xref>).</p>
<p>Shlyk et al. (<xref ref-type="bibr" rid="B8">8</xref>) demonstrated the efficacy of retrieval-augmented entity linking using embedded term definitions. However, our approach, which bypasses the need for explicit definitions, achieves comparable results. This suggests that advanced models like GPT-4o and GPT-3.5-Turbo can draw on pre-trained knowledge to assess semantic equivalence without relying on external definitions, streamlining the normalization process.</p>
<p>Regarding limitations, our study focused solely on term normalization and did not evaluate term identification. Additionally, our definition of &#x201C;semantic equivalence&#x201D; remains qualitative rather than exact. Determining whether &#x201C;impaired sensation&#x201D; is semantically equivalent to &#x201C;decreased sensation&#x201D; can be approached in three ways: (1) by measuring cosine similarity between appropriate word embeddings, (2) by the judgment of large language models trained for semantic reasoning, or (3) by gold-standard review from human domain experts. However, as datasets grow beyond 2,000 terms, human review becomes increasingly impractical, particularly for large ontologies such as SNOMED CT with over 400,000 terms. Another limitation is the relatively small and specialized list of terns derived from the OMIM summaries of neurogenetic diseases, which may not capture the full diversity of phenotype terms that require normalization. Expanding the dataset to cover a broader range of phenotypic terms from other domains could provide further insight.</p>
<p>Looking ahead, our retriever-based method is potentially generalizable to other terminologies, such as Gene Ontology (GO), UniProt, and SNOMED CT. By avoiding term definitions and relying solely on word embeddings, our approach could simplify normalization tasks in domains where definitions are ambiguous or difficult to generate. Further research into alternative retrieval strategies could improve the accuracy of the model and expand its applicability.</p>
<p>In conclusion, retrieval-augmented prompts based on BioBERT word embeddings improve the accuracy of phenotype normalization tasks. This simplified retriever performs as well as methods based on term definitions, offering an alternative solution for biomedical term normalization.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability"><title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. The Human Phenotype Ontology (HPO), which includes 22,929 classes, is available for download as a CSV file from the National Center for Biomedical Ontology at <ext-link ext-link-type="uri" xlink:href="https://bioportal.bioontology.org/ontologies/HP">https://bioportal.bioontology.org/ontologies/HP</ext-link>. A dataset containing 271,702 HPO annotations for 8,259 diseases in OMIM and 4,283 diseases in Orphanet can be accessed at <ext-link ext-link-type="uri" xlink:href="https://hpo.jax.org/data/annotations">https://hpo.jax.org/data/annotations</ext-link>. The Python code used in this project is publicly available on the project's GitHub repository at <ext-link ext-link-type="uri" xlink:href="https://github.com/clslabMSU/simplified-retriever">https://github.com/clslabMSU/simplified-retriever</ext-link>.</p>
</sec>
<sec id="s6" sec-type="ethics-statement"><title>Ethics statement</title>
<p>Ethical approval was not required for the study involving humans in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was not required from the participants or the participants&#x2019; legal guardians/next of kin in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec id="s7" sec-type="author-contributions"><title>Author contributions</title>
<p>DBH: Conceptualization, data curation, formal analysis, methodology, writing &#x2013; original draft, writing &#x2013; review &#x0026; editing. TSD: Formal analysis, validation, writing &#x2013; original draft, writing &#x2013; review &#x0026; editing. TO-A: Formal analysis, methodology, validation, writing &#x2013; original draft, writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec id="s8" sec-type="funding-information"><title>Funding</title>
<p>The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.</p>
</sec>
<sec id="s9" sec-type="COI-statement"><title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer"><title>Publisher&#x0027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list><title>References</title>
<ref id="B1"><label>1.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname><given-names>S</given-names></name><name><surname>Jin</surname><given-names>Q</given-names></name><name><surname>Yeganova</surname><given-names>L</given-names></name><name><surname>Lai</surname><given-names>PT</given-names></name><name><surname>Zhu</surname><given-names>Q</given-names></name><name><surname>Chen</surname><given-names>X</given-names></name></person-group>, et al. <article-title>Opportunities and challenges for ChatGPT and large language models in biomedicine and health</article-title>. <source>Brief Bioinform</source>. (<year>2023</year>) <volume>25</volume>(<issue>1</issue>):<fpage>bbad493</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbad493</pub-id><pub-id pub-id-type="pmid">38168838</pub-id></citation></ref>
<ref id="B2"><label>2.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romano</surname><given-names>MF</given-names></name><name><surname>Shih</surname><given-names>LC</given-names></name><name><surname>Paschalidis</surname><given-names>IC</given-names></name><name><surname>Au</surname><given-names>R</given-names></name><name><surname>Kolachalama</surname><given-names>VB</given-names></name></person-group>. <article-title>Large language models in neurology research and future practice</article-title>. <source>Neurology</source>. (<year>2023</year>) <volume>101</volume>(<issue>23</issue>):<fpage>1058</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1212/WNL.0000000000207967</pub-id><pub-id pub-id-type="pmid">37816646</pub-id></citation></ref>
<ref id="B3"><label>3.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clusmann</surname><given-names>J</given-names></name><name><surname>Kolbinger</surname><given-names>FR</given-names></name><name><surname>Muti</surname><given-names>HS</given-names></name><name><surname>Carrero</surname><given-names>ZI</given-names></name><name><surname>Eckardt</surname><given-names>JN</given-names></name><name><surname>Laleh</surname><given-names>NG</given-names></name></person-group>, et al. <article-title>The future landscape of large language models in medicine</article-title>. <source>Commun Med</source>. (<year>2023</year>) <volume>3</volume>(<issue>1</issue>):<fpage>141</fpage>. <pub-id pub-id-type="doi">10.1038/s43856-023-00370-1</pub-id><pub-id pub-id-type="pmid">37816837</pub-id></citation></ref>
<ref id="B4"><label>4.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname><given-names>Y</given-names></name><name><surname>Chen</surname><given-names>Q</given-names></name><name><surname>Du</surname><given-names>J</given-names></name><name><surname>Peng</surname><given-names>X</given-names></name><name><surname>Keloth</surname><given-names>VK</given-names></name><name><surname>Zuo</surname><given-names>X</given-names></name></person-group>, et al. <article-title>Improving large language models for clinical named entity recognition via prompt engineering</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2024</year>) <volume>31</volume>:<fpage>ocad259</fpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocad259</pub-id></citation></ref>
<ref id="B5"><label>5.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ge</surname><given-names>Y</given-names></name><name><surname>Guo</surname><given-names>Y</given-names></name><name><surname>Das</surname><given-names>S</given-names></name><name><surname>Al-Garadi</surname><given-names>MA</given-names></name><name><surname>Sarker</surname><given-names>A</given-names></name></person-group>. <article-title>Few-shot learning for medical text: a review of advances, trends, and opportunities</article-title>. <source>J Biomed Inform</source>. (<year>2023</year>) <volume>144</volume>:<fpage>104458</fpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2023.104458</pub-id><pub-id pub-id-type="pmid">37488023</pub-id></citation></ref>
<ref id="B6"><label>6.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Averly</surname><given-names>R</given-names></name><name><surname>Ning</surname><given-names>X</given-names></name></person-group>. <article-title>Entity decomposition with filtering: a zero-shot clinical named entity recognition framework.</article-title> <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:240704629</italic> (2024)</comment>.</citation></ref>
<ref id="B7"><label>7.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahoo</surname><given-names>SS</given-names></name><name><surname>Plasek</surname><given-names>JM</given-names></name><name><surname>Xu</surname><given-names>H</given-names></name><name><surname>Uzuner</surname><given-names>&#x00D6;</given-names></name><name><surname>Cohen</surname><given-names>T</given-names></name><name><surname>Yetisgen</surname><given-names>M</given-names></name></person-group>, et al. <article-title>Large language models for biomedicine: foundations, opportunities, challenges, and best practices</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2024</year>) <volume>31</volume>:<fpage>ocae074</fpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocae074</pub-id></citation></ref>
<ref id="B8"><label>8.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Shlyk</surname><given-names>D</given-names></name><name><surname>Groza</surname><given-names>T</given-names></name><name><surname>Mesiti</surname><given-names>M</given-names></name><name><surname>Montanelli</surname><given-names>S</given-names></name><name><surname>Cavalleri</surname><given-names>E</given-names></name></person-group>. <article-title>REAL: a retrieval-augmented entity linking approach for biomedical concept recognition.</article-title> <comment>In: <italic>Proceedings of the 23rd Workshop on Biomedical Natural Language Processing</italic> (2024). p. 380&#x2013;9</comment>.</citation></ref>
<ref id="B9"><label>9.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashley</surname><given-names>EA</given-names></name></person-group>. <article-title>Towards precision medicine</article-title>. <source>Nat Rev Genet</source>. (<year>2016</year>) <volume>17</volume>(<issue>9</issue>):<fpage>507</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1038/nrg.2016.86</pub-id><pub-id pub-id-type="pmid">27528417</pub-id></citation></ref>
<ref id="B10"><label>10.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Collins</surname><given-names>FS</given-names></name><name><surname>Varmus</surname><given-names>H</given-names></name></person-group>. <article-title>A new initiative on precision medicine</article-title>. <source>N Engl J Med</source>. (<year>2015</year>) <volume>372</volume>(<issue>9</issue>):<fpage>793</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1056/NEJMp1500523</pub-id><pub-id pub-id-type="pmid">25635347</pub-id></citation></ref>
<ref id="B11"><label>11.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robinson</surname><given-names>PN</given-names></name></person-group>. <article-title>Deep phenotyping for precision medicine</article-title>. <source>Hum Mutat</source>. (<year>2012</year>) <volume>33</volume>(<issue>5</issue>):<fpage>777</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1002/humu.22080</pub-id><pub-id pub-id-type="pmid">22504886</pub-id></citation></ref>
<ref id="B12"><label>12.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simmons</surname><given-names>M</given-names></name><name><surname>Singhal</surname><given-names>A</given-names></name><name><surname>Lu</surname><given-names>Z</given-names></name></person-group>. <article-title>Text mining for precision medicine: bringing structure to EHRs and biomedical literature to understand genes and health</article-title>. <source>Transl Biomed Inform</source>. (<year>2016</year>) <volume>939</volume>:<fpage>139</fpage>&#x2013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-10-1503-8_7</pub-id></citation></ref>
<ref id="B13"><label>13.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sitapati</surname><given-names>A</given-names></name><name><surname>Kim</surname><given-names>H</given-names></name><name><surname>Berkovich</surname><given-names>B</given-names></name><name><surname>Marmor</surname><given-names>R</given-names></name><name><surname>Singh</surname><given-names>S</given-names></name><name><surname>El-Kareh</surname><given-names>R</given-names></name></person-group>, et al. <article-title>Integrated precision medicine: the role of electronic health records in delivering personalized treatment</article-title>. <source>Wiley Interdiscip Rev Syst Biol Med</source>. (<year>2017</year>) <volume>9</volume>(<issue>3</issue>):<fpage>e1378</fpage>. <pub-id pub-id-type="doi">10.1002/wsbm.1378</pub-id></citation></ref>
<ref id="B14"><label>14.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahu</surname><given-names>M</given-names></name><name><surname>Gupta</surname><given-names>R</given-names></name><name><surname>Ambasta</surname><given-names>RK</given-names></name><name><surname>Kumar</surname><given-names>P</given-names></name></person-group>. <article-title>Artificial intelligence and machine learning in precision medicine: a paradigm shift in big data analysis</article-title>. <source>Prog Mol Biol Transl Sci</source>. (<year>2022</year>) <volume>190</volume>(<issue>1</issue>):<fpage>57</fpage>&#x2013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1016/bs.pmbts.2022.03.002</pub-id><pub-id pub-id-type="pmid">36008002</pub-id></citation></ref>
<ref id="B15"><label>15.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pathak</surname><given-names>J</given-names></name><name><surname>Kho</surname><given-names>AN</given-names></name><name><surname>Denny</surname><given-names>JC</given-names></name></person-group>. <article-title>Electronic health records-driven phenotyping: challenges, recent advances, and perspectives</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2013</year>) <volume>20</volume>(<issue>e2</issue>):<fpage>e206</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1136/amiajnl-2013-002428</pub-id><pub-id pub-id-type="pmid">24302669</pub-id></citation></ref>
<ref id="B16"><label>16.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Yelugam</surname><given-names>R</given-names></name><name><surname>Azizi</surname><given-names>S</given-names></name><name><surname>Carrithers</surname><given-names>MD</given-names></name><name><surname>Wunsch II</surname><given-names>DC</given-names></name></person-group>. <article-title>High throughput neurological phenotyping with MetaMap</article-title>. <source>Eur Sci J</source>. (<year>2022</year>) <volume>18</volume>:<fpage>37</fpage>&#x2013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.19044/esj.2022.v18n4p37</pub-id></citation></ref>
<ref id="B17"><label>17.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Yelugam</surname><given-names>R</given-names></name><name><surname>Azizi</surname><given-names>S</given-names></name><name><surname>Wunsch II</surname><given-names>DC</given-names></name></person-group>. <article-title>A focused review of deep phenotyping with examples from Neurology</article-title>. <source>Eur Sci J</source>. (<year>2022</year>) <volume>18</volume>:<fpage>4</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.19044/esj.2022.v18n4p4</pub-id></citation></ref>
<ref id="B18"><label>18.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Munzir</surname><given-names>SI</given-names></name><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Carrithers</surname><given-names>MD</given-names></name></person-group>. <article-title>High throughput phenotyping of physician notes with large language and hybrid NLP models.</article-title> <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:240305920</italic> (2024); Accepted in International Conference of the IEEE Engineering in Medicine &#x0026; Biology Society (EMBC 2024). Available online at: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.2403.05920">https://doi.org/10.48550/arXiv.2403.05920</ext-link></comment>.</citation></ref>
<ref id="B19"><label>19.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanna</surname><given-names>J</given-names></name><name><surname>Joseph</surname><given-names>E</given-names></name><name><surname>Brochhausen</surname><given-names>M</given-names></name><name><surname>Hogan</surname><given-names>WR</given-names></name></person-group>. <article-title>Building a drug ontology based on RxNorm and other sources</article-title>. <source>J Biomed Semant</source>. (<year>2013</year>) <volume>4</volume>:<fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1186/2041-1480-4-44</pub-id></citation></ref>
<ref id="B20"><label>20.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fung</surname><given-names>KW</given-names></name><name><surname>Xu</surname><given-names>J</given-names></name><name><surname>Bodenreider</surname><given-names>O</given-names></name></person-group>. <article-title>The new international classification of diseases 11th edition: a comparative analysis with ICD-10 and ICD-10-CM</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2020</year>) <volume>27</volume>(<issue>5</issue>):<fpage>738</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocaa030</pub-id><pub-id pub-id-type="pmid">32364236</pub-id></citation></ref>
<ref id="B21"><label>21.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McDonald</surname><given-names>CJ</given-names></name><name><surname>Huff</surname><given-names>SM</given-names></name><name><surname>Suico</surname><given-names>JG</given-names></name><name><surname>Hill</surname><given-names>G</given-names></name><name><surname>Leavelle</surname><given-names>D</given-names></name><name><surname>Aller</surname><given-names>R</given-names></name></person-group>, et al. <article-title>LOINC, a universal standard for identifying laboratory observations: a 5-year update</article-title>. <source>Clin Chem</source>. (<year>2003</year>) <volume>49</volume>(<issue>4</issue>):<fpage>624</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1373/49.4.624</pub-id><pub-id pub-id-type="pmid">12651816</pub-id></citation></ref>
<ref id="B22"><label>22.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shivade</surname><given-names>C</given-names></name><name><surname>Raghavan</surname><given-names>P</given-names></name><name><surname>Fosler-Lussier</surname><given-names>E</given-names></name><name><surname>Embi</surname><given-names>PJ</given-names></name><name><surname>Elhadad</surname><given-names>N</given-names></name><name><surname>Johnson</surname><given-names>SB</given-names></name></person-group>, et al. <article-title>A review of approaches to identifying patient phenotype cohorts using electronic health records</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2014</year>) <volume>21</volume>(<issue>2</issue>):<fpage>221</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1136/amiajnl-2013-001935</pub-id><pub-id pub-id-type="pmid">24201027</pub-id></citation></ref>
<ref id="B23"><label>23.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>French</surname><given-names>E</given-names></name><name><surname>McInnes</surname><given-names>BT</given-names></name></person-group>. <article-title>An overview of biomedical entity linking throughout the years</article-title>. <source>J Biomed Inform</source>. (<year>2023</year>) <volume>137</volume>:<fpage>104252</fpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2022.104252</pub-id><pub-id pub-id-type="pmid">36464228</pub-id></citation></ref>
<ref id="B24"><label>24.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Jonquet</surname><given-names>C</given-names></name><name><surname>Shah</surname><given-names>NH</given-names></name><name><surname>Youn</surname><given-names>CH</given-names></name><name><surname>Musen</surname><given-names>MA</given-names></name><name><surname>Callendar</surname><given-names>C</given-names></name><name><surname>Storey</surname><given-names>MA</given-names></name></person-group>. <article-title>NCBO annotator: semantic annotation of biomedical data.</article-title> <comment>In: <italic>ISWC 2009-8th International Semantic Web Conference, Poster and Demo Session</italic> (2009). Poster 171</comment>.</citation></ref>
<ref id="B25"><label>25.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>C</given-names></name><name><surname>Peres Kury</surname><given-names>FS</given-names></name><name><surname>Li</surname><given-names>Z</given-names></name><name><surname>Ta</surname><given-names>C</given-names></name><name><surname>Wang</surname><given-names>K</given-names></name><name><surname>Weng</surname><given-names>C</given-names></name></person-group>. <article-title>Doc2Hpo: a web application for efficient and accurate HPO concept curation</article-title>. <source>Nucleic Acids Res</source>. (<year>2019</year>) <volume>47</volume>(<issue>W1</issue>):<fpage>W566</fpage>&#x2013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz386</pub-id><pub-id pub-id-type="pmid">31106327</pub-id></citation></ref>
<ref id="B26"><label>26.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feng</surname><given-names>Y</given-names></name><name><surname>Qi</surname><given-names>L</given-names></name><name><surname>Tian</surname><given-names>W</given-names></name></person-group>. <article-title>PhenoBERT: a combined deep learning method for automated recognition of human phenotype ontology</article-title>. <source>IEEE/ACM Trans Comput Biol Bioinform</source>. (<year>2022</year>) <volume>20</volume>(<issue>2</issue>):<fpage>1269</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2022.3170301</pub-id></citation></ref>
<ref id="B27"><label>27.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sung</surname><given-names>M</given-names></name><name><surname>Jeong</surname><given-names>M</given-names></name><name><surname>Choi</surname><given-names>Y</given-names></name><name><surname>Kim</surname><given-names>D</given-names></name><name><surname>Lee</surname><given-names>J</given-names></name><name><surname>Kang</surname><given-names>J</given-names></name></person-group>. <article-title>BERN2: an advanced neural biomedical named entity recognition and normalization tool</article-title>. <source>Bioinformatics</source>. (<year>2022</year>) <volume>38</volume>(<issue>20</issue>):<fpage>4837</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btac598</pub-id><pub-id pub-id-type="pmid">36053172</pub-id></citation></ref>
<ref id="B28"><label>28.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shefchek</surname><given-names>KA</given-names></name><name><surname>Harris</surname><given-names>NL</given-names></name><name><surname>Gargano</surname><given-names>M</given-names></name><name><surname>Matentzoglu</surname><given-names>N</given-names></name><name><surname>Unni</surname><given-names>D</given-names></name><name><surname>Brush</surname><given-names>M</given-names></name></person-group>, et al. <article-title>The monarch initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across species</article-title>. <source>Nucleic Acids Res</source>. (<year>2020</year>) <volume>48</volume>(<issue>D1</issue>):<fpage>D704</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz997</pub-id><pub-id pub-id-type="pmid">31701156</pub-id></citation></ref>
<ref id="B29"><label>29.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Groza</surname><given-names>T</given-names></name><name><surname>Gration</surname><given-names>D</given-names></name><name><surname>Baynam</surname><given-names>G</given-names></name><name><surname>Robinson</surname><given-names>PN</given-names></name></person-group>. <article-title>FastHPOCR: pragmatic, fast, and accurate concept recognition using the human phenotype ontology</article-title>. <source>Bioinformatics</source>. (<year>2024</year>) <volume>40</volume>(<issue>7</issue>):<fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btae406</pub-id></citation></ref>
<ref id="B30"><label>30.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Deisseroth</surname><given-names>CA</given-names></name><name><surname>Birgmeier</surname><given-names>J</given-names></name><name><surname>Bodle</surname><given-names>EE</given-names></name><name><surname>Bernstein</surname><given-names>JA</given-names></name><name><surname>Bejerano</surname><given-names>G</given-names></name></person-group>. <article-title>ClinPhen extracts and prioritizes patient phenotypes directly from medical records to accelerate genetic disease diagnosis.</article-title> <comment><italic>BioRxiv</italic> [Preprint]. (2018). p. 362111</comment>.</citation></ref>
<ref id="B31"><label>31.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Groza</surname><given-names>T</given-names></name><name><surname>K&#x00F6;hler</surname><given-names>S</given-names></name><name><surname>Doelken</surname><given-names>S</given-names></name><name><surname>Collier</surname><given-names>N</given-names></name><name><surname>Oellrich</surname><given-names>A</given-names></name><name><surname>Smedley</surname><given-names>D</given-names></name></person-group>, et al. <article-title>Automatic concept recognition using the human phenotype ontology reference and test suite corpora</article-title>. <source>Database</source>. (<year>2015</year>) <volume>2015</volume>:<fpage>bav005</fpage>. <pub-id pub-id-type="doi">10.1093/database/bav005</pub-id><pub-id pub-id-type="pmid">25725061</pub-id></citation></ref>
<ref id="B32"><label>32.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Munzir</surname><given-names>SI</given-names></name><name><surname>Stahlfeld</surname><given-names>A</given-names></name><name><surname>Obafemi-Ajayi</surname><given-names>T</given-names></name><name><surname>Carrithers</surname><given-names>MD</given-names></name></person-group>. <article-title>High-throughput phenotyping of clinical text using large language models.</article-title> <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:240801214</italic> (2024); Submitted to IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI), Houston, TX</comment>.</citation></ref>
<ref id="B33"><label>33.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhang</surname><given-names>Y</given-names></name><name><surname>Li</surname><given-names>Y</given-names></name><name><surname>Cui</surname><given-names>L</given-names></name><name><surname>Cai</surname><given-names>D</given-names></name><name><surname>Liu</surname><given-names>L</given-names></name><name><surname>Fu</surname><given-names>T</given-names></name></person-group>, et al. <article-title>Siren&#x2019;s song in the AI ocean: a survey on hallucination in large language models.</article-title> <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:230901219</italic> (2023)</comment>.</citation></ref>
<ref id="B34"><label>34.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Gao</surname><given-names>Y</given-names></name><name><surname>Xiong</surname><given-names>Y</given-names></name><name><surname>Gao</surname><given-names>X</given-names></name><name><surname>Jia</surname><given-names>K</given-names></name><name><surname>Pan</surname><given-names>J</given-names></name><name><surname>Bi</surname><given-names>Y</given-names></name></person-group>, et al. <article-title>Retrieval-augmented generation for large language models: a survey.</article-title> <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:231210997</italic> (2023)</comment>.</citation></ref>
<ref id="B35"><label>35.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amberger</surname><given-names>JS</given-names></name><name><surname>Bocchini</surname><given-names>CA</given-names></name><name><surname>Schiettecatte</surname><given-names>F</given-names></name><name><surname>Scott</surname><given-names>AF</given-names></name><name><surname>Hamosh</surname><given-names>A</given-names></name></person-group>. <article-title>OMIM.org: online mendelian inheritance in man (OMIM&#x00AE;), an online catalog of human genes and genetic disorders</article-title>. <source>Nucleic Acids Res</source>. (<year>2015</year>) <volume>43</volume>(<issue>D1</issue>):<fpage>D789</fpage>&#x2013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku1205</pub-id><pub-id pub-id-type="pmid">25428349</pub-id></citation></ref>
<ref id="B36"><label>36.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Munzir</surname><given-names>SI</given-names></name><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Oommen</surname><given-names>C</given-names></name><name><surname>Carrithers</surname><given-names>MD</given-names></name></person-group>. <article-title>A large language model outperforms other computational approaches to the high-throughput phenotyping of physician notes.</article-title> <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:240614757</italic> (2024); Accepted in AMIA Annual Symposium 2024. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2406.14757">https://arxiv.org/abs/2406.14757</ext-link></comment>.</citation></ref>
<ref id="B37"><label>37.</label><citation citation-type="other"><collab>Human Phenotype Ontology</collab>. <article-title>NCBO BioPortal (2024).</article-title> <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://bioportalbioontologyorg/ontologies/HP">https://bioportalbioontologyorg/ontologies/HP</ext-link> (accessed July 27, 2024)</comment>.</citation></ref>
<ref id="B38"><label>38.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname><given-names>J</given-names></name><name><surname>Yoon</surname><given-names>W</given-names></name><name><surname>Kim</surname><given-names>S</given-names></name><name><surname>Kim</surname><given-names>D</given-names></name><name><surname>Kim</surname><given-names>S</given-names></name><name><surname>So</surname><given-names>CH</given-names></name></person-group>, et al. <article-title>BioBERT: a pre-trained biomedical language representation model for biomedical text mining</article-title>. <source>Bioinformatics</source>. (<year>2020</year>) <volume>36</volume>(<issue>4</issue>):<fpage>1234</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz682</pub-id><pub-id pub-id-type="pmid">31501885</pub-id></citation></ref></ref-list>
</back>
</article>