<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Digit. Health</journal-id>
<journal-title>Frontiers in Digital Health</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Digit. Health</abbrev-journal-title>
<issn pub-type="epub">2673-253X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdgth.2023.1075771</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Digital Health</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Inter-rater agreement for the annotation of neurologic signs and symptoms in electronic health records</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Oommen</surname><given-names>Chelsea</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/2100556/overview"/></contrib>
<contrib contrib-type="author"><name><surname>Howlett-Prieto</surname><given-names>Quentin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/2041607/overview" /></contrib>
<contrib contrib-type="author"><name><surname>Carrithers</surname><given-names>Michael D.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/1682797/overview" /></contrib>
<contrib contrib-type="author" corresp="yes"><name><surname>Hier</surname><given-names>Daniel B.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref><uri xlink:href="https://loop.frontiersin.org/people/1230804/overview" /></contrib>
</contrib-group>
<aff id="aff1"><label><sup><sup>1</sup></sup></label><addr-line>Department of Neurology and Rehabilitation</addr-line>, <institution>University of Illinois at Chicago</institution>, <addr-line>Chicago, IL</addr-line>, <country>United States</country></aff>
<aff id="aff2"><label><sup><sup>2</sup></sup></label><addr-line>Department of Electrical and Computer Engineering</addr-line>, <institution>Missouri University of Science and Technology</institution>, <addr-line>Rolla, MO</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p><bold>Edited by:</bold> Lina F. Soualmia, Universit&#x00E9; de Rouen, France</p></fn>
<fn fn-type="edited-by"><p><bold>Reviewed by:</bold> Xia Jing, Clemson University, United States Alec Chapman, The University of Utah, United States</p></fn>
<corresp id="cor1"><label>&#x002A;</label><bold>Correspondence:</bold> Daniel B. Hier <email>hierd@mst.edu</email></corresp>
</author-notes>
<pub-date pub-type="epub"><day>13</day><month>06</month><year>2023</year></pub-date>
<pub-date pub-type="collection"><year>2023</year></pub-date>
<volume>5</volume><elocation-id>1075771</elocation-id>
<history>
<date date-type="received"><day>20</day><month>10</month><year>2022</year></date>
<date date-type="accepted"><day>26</day><month>05</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Oommen, Howlett-Prieto, Carrithers and Hier.</copyright-statement>
<copyright-year>2023</copyright-year><copyright-holder>Oommen, Howlett-Prieto, Carrithers and Hier</copyright-holder><license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The extraction of patient signs and symptoms recorded as free text in electronic health records is critical for precision medicine. Once extracted, signs and symptoms can be made computable by mapping to signs and symptoms in an ontology. Extracting signs and symptoms from free text is tedious and time-consuming. Prior studies have suggested that inter-rater agreement for clinical concept extraction is low. We have examined inter-rater agreement for annotating neurologic concepts in clinical notes from electronic health records. After training on the annotation process, the annotation tool, and the supporting neuro-ontology, three raters annotated 15 clinical notes in three rounds. Inter-rater agreement between the three annotators was high for text span and category label. A machine annotator based on a convolutional neural network had a high level of agreement with the human annotators but one that was lower than human inter-rater agreement. We conclude that high levels of agreement between human annotators are possible with appropriate training and annotation tools. Furthermore, more training examples combined with improvements in neural networks and natural language processing should make machine annotators capable of high throughput automated clinical concept extraction with high levels of agreement with human annotators.</p>
</abstract>
<kwd-group>
<kwd>natural language processing</kwd>
<kwd>annotation</kwd>
<kwd>electronic health records</kwd>
<kwd>phenotype</kwd>
<kwd>clinical concept extraction</kwd>
<kwd>inter-rater agreement</kwd>
<kwd>neural networks</kwd>
<kwd>signs and symptoms</kwd>
</kwd-group><counts>
<fig-count count="4"/>
<table-count count="0"/><equation-count count="53"/><ref-count count="36"/><page-count count="0"/><word-count count="0"/></counts><custom-meta-wrap><custom-meta><meta-name>section-at-acceptance</meta-name><meta-value>Health Informatics</meta-value></custom-meta></custom-meta-wrap>
</article-meta>
</front>
<body><sec id="s1" sec-type="intro"><title>Introduction</title>
<p>Extracting medical concepts from electronic health records is key to precision medicine (<xref ref-type="bibr" rid="B1">1</xref>). The signs and symptoms of patients (part of the patient phenotype) are generally recorded as free text in progress notes, admission notes, and discharge summaries (<xref ref-type="bibr" rid="B2">2</xref>). Clinical phenotyping of patients involves the mapping of free text to defined terms that are concepts in an ontology (<xref ref-type="bibr" rid="B3">3</xref>,<xref ref-type="bibr" rid="B4">4</xref>). This is a two-step process that involves identifying appropriate text spans in narratives and then converting the text spans to target concepts in an ontology (<xref ref-type="bibr" rid="B5">5</xref>,<xref ref-type="bibr" rid="B6">6</xref>). The process of mapping free text to defined classes in an ontology, illustrated in (1) and (2), has been termed <bold>normalization</bold> (<xref ref-type="bibr" rid="B7">7</xref>,<xref ref-type="bibr" rid="B8">8</xref>).<disp-formula id="disp-formula1"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM1"><mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi mathvariant="normal">patient</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">movements</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">were</mml:mi></mml:mrow></mml:mrow><mml:mspace width="thinmathspace"/><mml:mrow><mml:mi mathvariant="bold">a</mml:mi><mml:mi mathvariant="bold">t</mml:mi><mml:mi mathvariant="bold">a</mml:mi><mml:mi mathvariant="bold">x</mml:mi><mml:mi mathvariant="bold">i</mml:mi><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mrow><mml:mi mathvariant="bold">a</mml:mi><mml:mi mathvariant="bold">t</mml:mi><mml:mi mathvariant="bold">a</mml:mi><mml:mi mathvariant="bold">x</mml:mi><mml:mi mathvariant="bold">i</mml:mi><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">UMLS</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">CUI</mml:mi></mml:mrow></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="bold">C</mml:mi><mml:mn mathvariant="bold">0004134</mml:mn></mml:mrow></mml:math></disp-formula><disp-formula id="disp-formula2"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM2"><mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi mathvariant="normal">freetext</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">clinical</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">concept</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">machine</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">readable</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">code</mml:mi></mml:mrow></mml:mrow></mml:math></disp-formula>In this example 1, an annotator highlights the term ataxic, then it is mapped to the concept ataxia, and the UMLS code CUI C0004134 is retrieved (<xref ref-type="bibr" rid="B9">9</xref>). This is a slow and error-prone process for human annotators. Agreement between human raters for annotation of clinical text is often low. A study on the agreement for SNOMED CT codes between coders from three professional coding companies yielded about 50 percent agreement for exact matches with slightly higher agreement when adjusted for near matches (<xref ref-type="bibr" rid="B10">10</xref>). Another study of SNOMED CT coding of ophthalmology notes yielded low levels of inter-rater agreement ranging from 33 to 64&#x0025; (<xref ref-type="bibr" rid="B11">11</xref>). Identified sources of disagreement between coders included human errors (lack of applicable medical knowledge, lack of recognition of abbreviations for concepts, and general carelessness), annotation guideline flaws (under specified and unclear guidelines), ontology flaws (polysemy of coded concepts), interface term issues (inconsistent categorization of clinical jargon), and language issues (interpretation difficulties due to use of ellipsis, anaphora, paraphrasing, and other linguistic concepts) (<xref ref-type="bibr" rid="B12">12</xref>).</p>
<p>The goal of high throughput phenotyping is to use natural language processing (NLP) to automate the annotation process (<xref ref-type="bibr" rid="B13">13</xref>). Approaches to high throughput clinical concept extraction have included rule-based systems, traditional machine learning algorithms, deep learning algorithms, and hybrid methods that combine algorithms (<xref ref-type="bibr" rid="B6">6</xref>). Tools for concept extraction based on rules, linguistic analysis, and statistical models, such as cTAKES and MetaMap, generally have accuracy and recall between 0.38 and 0.66 (<xref ref-type="bibr" rid="B5">5</xref>,<xref ref-type="bibr" rid="B14">14</xref>,<xref ref-type="bibr" rid="B15">15</xref>). Neural networks are being used for concept recognition with increasing success. Arbabi et al. developed a convolutional neural network that matches input phrases to concepts in the Human Phenotype Ontology with high accuracy (<xref ref-type="bibr" rid="B16">16</xref>). Other deep learning approaches, including neural networks based on bidirectional encoder representations from transformers (BERT), show promise for automated clinical concept extraction (<xref ref-type="bibr" rid="B5">5</xref>,<xref ref-type="bibr" rid="B6">6</xref>,<xref ref-type="bibr" rid="B17">17</xref>,<xref ref-type="bibr" rid="B18">18</xref>).</p>
<p>In this paper, we examine inter-rater agreement for text-span identification of neurological concepts in notes from electronic health records. In addition to the agreement between human annotators, we examine the agreement between human annotators and a machine annotator based on a convolutional neural network.</p>
</sec>
<sec id="s2" sec-type="methods"><title>Methods</title>
<sec id="s2a"><title>Annotation tool</title>
<p>Prodigy (Explosion AI, Berlin, Germany) was used to annotate neurologic concepts in the EHR physician notes. Prodigy runs under python in the terminal mode of macOS, Windows, or Linux. It creates a web interface locally (<xref ref-type="fig" rid="F1">Figures&#x00A0;1A</xref>,<xref ref-type="fig" rid="F1">B</xref>). As input, Prodigy requires free text to be converted to JSON format.<disp-formula id="disp-formula3"><label>(3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM3"><mml:mrow><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="normal">text</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x2033;</mml:mo></mml:msup><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="normal">The</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">patient</mml:mi><mml:mspace width="thinmathspace"/><mml:mi mathvariant="normal">had</mml:mi></mml:mrow></mml:mrow><mml:mspace width="thinmathspace"/><mml:mrow><mml:mi mathvariant="bold">w</mml:mi><mml:mi mathvariant="bold">e</mml:mi><mml:mi mathvariant="bold">a</mml:mi><mml:mi mathvariant="bold">k</mml:mi><mml:mi mathvariant="bold">n</mml:mi><mml:mi mathvariant="bold">e</mml:mi><mml:mi mathvariant="bold">s</mml:mi><mml:mi mathvariant="bold">s</mml:mi></mml:mrow><mml:mspace width="thinmathspace"/><mml:mrow><mml:mrow><mml:mi mathvariant="normal">and</mml:mi></mml:mrow></mml:mrow><mml:mspace width="thickmathspace" /><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="bold">s</mml:mi><mml:mi mathvariant="bold">e</mml:mi><mml:mi mathvariant="bold">n</mml:mi><mml:mi mathvariant="bold">s</mml:mi><mml:mi mathvariant="bold">o</mml:mi><mml:mi mathvariant="bold">r</mml:mi><mml:mi mathvariant="bold">y</mml:mi><mml:mspace width="thickmathspace" /><mml:mi mathvariant="bold">l</mml:mi><mml:mi mathvariant="bold">o</mml:mi><mml:mi mathvariant="bold">s</mml:mi><mml:mi mathvariant="bold">s</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x2033;</mml:mo></mml:msup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:math></disp-formula>Each line of text from a JSON file 3, appears as a separate screen for annotation by Prodigy (<xref ref-type="fig" rid="F1">Figures&#x00A0;1A</xref>,<xref ref-type="fig" rid="F1">B</xref>). Annotations are stored in an SQLite database and are exportable with annotations and text spans as a JSON file. Prodigy is integrated with the <italic>spaCy</italic> natural language processing toolkit (Explosion AI) and can train neural networks for named entity recognition and text classification.</p>
<fig id="F1" position="float"><label>Figure 1</label>
<caption><p>(<bold>A</bold>) Annotator screen for a patient with multiple sclerosis. The patient complains of imbalance, leg weakness, and pain, and these concepts have been annotated. Imbalance and pain are labeled as unigrams; leg weakness is labeled a bigram. Annotators were trained to ignore laterality (e.g., right leg weakness.) Each Prodigy screen reflects one line of text from the JSON input file. This screen has three potential items to contribute to the Kappa statistic: imbalance, leg weakness, and pain. (<bold>B</bold>) Annotator screen for neurological concepts for a patient with multiple sclerosis. The patient denies problems with vision, sensation, bladder, bowel, gait, or falls. The annotators are trained not to annotate negated concepts. The NN had no specific negation rule but learned not to tag negated concepts through training examples. Since there are no signs and symptoms in this screen, if both annotators show no annotations, a score of 1 is assigned to the Kappa statistic for agreement on this screen. If one annotator shows no annotations and another shows annotations on this screen, annotator disagreement is scored.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-05-1075771-g001.tif"/>
</fig>
<p>The Kappa statistic was used to assess agreement between the three annotators and the neural network. The Kappa statistic corrects observed rater agreement for chance rater agreement. It ranges from 0 to 1, where 1 is complete agreement, 0 is a chance agreement. Values of Kappa of 0.6 to 0.79 are considered substantial agreement, values between 0.8 and 0.90 are considered strong agreement, and values over 0.90 are considered near perfect agreement (<xref ref-type="bibr" rid="B19">19</xref>,<xref ref-type="bibr" rid="B20">20</xref>). For each line of text that had one or more annotations (3), the agreement was rated 1 for the annotations if both annotators agreed and rated 0 if the annotators disagreed. A line of text with no annotations (null&#x005F;annotations) by either annotator was scored 1 for agreement. The total number of annotations considered by the Kappa statistic for two raters <bold>A</bold> and <bold>B</bold> was (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM1"><mml:mi>A</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>B</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">null</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi mathvariant="normal">annotations</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>).</p>
</sec>
<sec id="s2b"><title>Rater training and instructions</title>
<p>Three annotators participated in the research. Annotator 1 (A1) was a senior neurologist, Annotator 2 (A2) was a pre-medical student majoring in neuroscience, and Annotator 3 (A3) was a third-year medical student. Raters first reviewed neurologic signs and symptoms in the neuro-ontology of neurological concepts (<xref ref-type="bibr" rid="B21">21</xref>) and then were instructed to find all neurological concepts in the neurology notes. Signs and symptoms (ataxia, fatigue, weakness, memory loss, etc.) were annotated but not disease entities (Alzheimer&#x2019;s disease, multiple sclerosis, etc.) Raters annotated the neurologic concepts and ignored laterality and other modifiers (e.g., <italic>arm pain</italic> for <italic>right arm pain</italic>, <italic>back pain</italic> for <italic>severe back pain</italic>, etc.) In addition, annotators tagged each text span with an category label (see <xref ref-type="fig" rid="F1">Figures&#x00A0;1A</xref>,<xref ref-type="fig" rid="F1">B</xref>). Category labels included <italic>unigrams</italic> (one-word concepts such as ataxia), <italic>bigrams</italic> (two-word concepts such as double vision), <italic>trigrams</italic> (three-word concepts such as low back pain), <italic>tetragrams</italic> (four-word concepts such as relative afferent pupil defect), <italic>extended</italic> (text span annotations longer than four words), <italic>compound</italic> (multiple concepts in one text span such as brisk ankle and knee reflex), and <italic>tabular</italic> (concepts represented in tabular or columnar format, usually showed right and left body sides). Our motivation for tagging signs and symptoms by the length and type of the text span was a hypothesis that neural networks trained to recognize signs and symptoms in medical text would exhibit lower accuracies with longer text spans. This hypothesis was confirmed by a recent study from our group (<xref ref-type="bibr" rid="B18">18</xref>).</p>
</sec>
<sec id="s2c"><title>The machine annotator</title>
<p>The machine annotator (NN) was a neural network that was trained to recognize text spans containing neurology concepts in the electronic health record physician notes. The NN was the default spaCy named entity recognition model based on a four-layer convolutional neural network (CNN) that looked at four words on either side of each token using <italic>tok2vec</italic> with an initial learning rate <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM2"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The default parameters provided by Prodigy were used for training. NN was trained on 11,000 manually annotated sentences derived from neurology textbooks, online neurological disease descriptions, and electronic health record notes. Further details on training the NN are available in (<xref ref-type="bibr" rid="B18">18</xref>).</p>
</sec>
<sec id="s2d"><title>Annotations</title>
<p>Five patient EHR notes were annotated for each of the three rounds. The annotation of EHR clinical notes for research purposes was approved by the Institutional Review Board of the University of Illinois (UIC Neuroimmunology Biobank 2017-0520Z). Informed patient consent for use of clinical notes was obtained from all subjects through the UIC Biobank Project. Three human annotators (A1, A2, and A3) and the machine annotator (NN) annotated each note. After each round, the annotators met and reviewed any annotation disagreements. The annotations of each annotator were stored in an SQLite database and exported as a JSON file for scoring for inter-rater agreement in Python. Text spans were mapped to concepts in the neuro-ontology (<xref ref-type="bibr" rid="B21">21</xref>) utilizing a lookup table with 3,500 target phrases and the similarity method from spaCy (<xref ref-type="bibr" rid="B22">22</xref>) (pp. 152&#x2013;54). Univariate analysis of variance and Cohen&#x2019;s Kappa statistic were calculated with SPSS (IBM, version 28).</p>
</sec>
</sec>
<sec id="s3" sec-type="results"><title>Results</title>
<p>Annotators identified neurological signs and symptoms in physician notes from electronic health records. Each annotator identified the text span associated with each sign and symptom and assigned a category label to each annotation (e.g., unigram, bigram, trigram, etc.) Inter-rater agreement (adjusted and unadjusted) was calculated between the three human annotators and the machine annotator (NN).</p>
<p>Although five EHR notes were annotated for each round, the notes varied in length. Each line in the EHR note was converted to a single line in the JSON file and generated one annotation screen in the Prodigy annotator. Round 1 had 625 annotation screens with 139 signs and symptoms to annotate, Round 2 had 674 annotation screens with 205 signs and symptoms to annotate, and Round 3 had 523 annotation screens with 138 signs and symptoms to annotate. Since the number of signs and symptoms was less than the number of annotation screens, many annotation screens had no signs or symptoms to annotate (null screens). When both annotators agreed that the annotation screen had no signs or symptoms, this was scored as annotator agreement for both the adjusted and unadjusted metrics (Kappa and concordance).</p>
<p>Concordance (unadjusted agreement) on the text span task was <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM3"><mml:mn>88.9</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo>&#x00B1;</mml:mo><mml:mn>3.2</mml:mn></mml:math></inline-formula> (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM4"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SD</mml:mi></mml:mrow></mml:math></inline-formula>) between the human annotators and was <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM5"><mml:mn>83.9</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo>&#x00B1;</mml:mo><mml:mn>4.6</mml:mn></mml:math></inline-formula> (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM6"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SD</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> between the human annotators and the machine annotator (human-human mean was higher, one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM7"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM8"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.016</mml:mn></mml:math></inline-formula>). Concordance (unadjusted agreement) on the category label task was <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM9"><mml:mn>87.7</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo>&#x00B1;</mml:mo><mml:mn>4.4</mml:mn></mml:math></inline-formula> (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM10"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SD</mml:mi></mml:mrow></mml:math></inline-formula>) between human annotators and was <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM11"><mml:mn>84.6</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo>&#x00B1;</mml:mo><mml:mn>5.5</mml:mn></mml:math></inline-formula> (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM12"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SD</mml:mi></mml:mrow></mml:math></inline-formula>) between the human annotators and the machine annotator (means did not differ, one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM13"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM14"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.212</mml:mn></mml:math></inline-formula>).</p>
<p>Cohen&#x2019;s Kappa statistic <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM15"><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03BA;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> was high for both the text span task (0.715 to 0.893) and the category label task (0.72 to 0.89) (<xref ref-type="fig" rid="F2">Figures&#x00A0;2A</xref>,<xref ref-type="fig" rid="F2">B</xref>). On the text span identification task (<xref ref-type="fig" rid="F3">Figure&#x00A0;3A</xref>) <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM16"><mml:mi>&#x03BA;</mml:mi></mml:math></inline-formula> was higher for the human-human pairs (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM17"><mml:mn>0.85</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula> <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM18"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SD</mml:mi></mml:mrow></mml:math></inline-formula>) than the human-machine pairs (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM19"><mml:mn>0.76</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.06</mml:mn></mml:math></inline-formula>). On the category label task, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM20"><mml:mi>&#x03BA;</mml:mi></mml:math></inline-formula> (<xref ref-type="fig" rid="F3">Figure&#x00A0;3B</xref>) was similar between the human-human pairs (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM21"><mml:mn>0.83</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula> <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM22"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SD</mml:mi></mml:mrow></mml:math></inline-formula>) and the human-machine pairs (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM23"><mml:mn>0.82</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.06</mml:mn></mml:math></inline-formula>). <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM24"><mml:mi>&#x03BA;</mml:mi></mml:math></inline-formula> for the text span task and the category label task did not differ by round (for <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM25"><mml:mi>p</mml:mi></mml:math></inline-formula> values and means see <xref ref-type="fig" rid="F4">Figures&#x00A0;4A</xref>,<xref ref-type="fig" rid="F4">B</xref>).</p>
<fig id="F2" position="float"><label>Figure 2</label>
<caption><p>(<bold>A</bold>) Boxplots for the Kappa statistic for inter-rater agreement for text spans for the neurological concepts. Univariate analysis of variance showed that mean inter-rater agreement differed by rating pair (one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM26"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM27"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.021</mml:mn></mml:math></inline-formula>). Post hoc comparisons by the Bonferroni method showed that pair A1-A2 outperformed pair NN-A2. (<bold>B</bold>) Boxplots for the Kappa statistic for inter-rater agreement for category labels for the neurological concepts. Univariate analysis of variance showed that mean Kappa for category label agreement did not differ by rating pair (one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM28"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.165</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM29"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>).</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-05-1075771-g002.tif"/>
</fig>
<fig id="F3" position="float"><label>Figure 3</label>
<caption><p>(<bold>A</bold>) Kappa statistic for agreement between human-human and human-machine raters for text span. Groups differed, one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM30"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM31"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula>. (<bold>B</bold>) Kappa statistic for agreement between human-human and human-machine raters for category label. Groups did not differ, one way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM32"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM33"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.589</mml:mn></mml:math></inline-formula>.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-05-1075771-g003.tif"/>
</fig>
<fig id="F4" position="float"><label>Figure 4</label>
<caption><p>(<bold>A</bold>) Kappa statistic for inter-rater agreement for text span by round. Round 1: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM34"><mml:mn>0.78</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.03</mml:mn></mml:math></inline-formula> (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM35"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SE</mml:mi></mml:mrow></mml:math></inline-formula>), Round 2: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM36"><mml:mn>0.84</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.03</mml:mn></mml:math></inline-formula>, Round 3: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM37"><mml:mn>0.81</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.03</mml:mn></mml:math></inline-formula>, groups do not differ, one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM38"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM39"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.310</mml:mn></mml:math></inline-formula>. (<bold>B</bold>) Kappa statistic for inter-rater agreement for category label by round. Round 1: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM40"><mml:mn>0.80</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.21</mml:mn></mml:math></inline-formula> (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM41"><mml:mrow><mml:mi mathvariant="normal">mean</mml:mi></mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">SE</mml:mi></mml:mrow></mml:math></inline-formula>). Round 2: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM42"><mml:mn>0.85</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.21</mml:mn></mml:math></inline-formula>, Round 3: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM43"><mml:mn>0.83</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.21</mml:mn></mml:math></inline-formula>, groups do not differ, one-way ANOVA, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM44"><mml:mrow><mml:mi mathvariant="normal">df</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM45"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.306</mml:mn></mml:math></inline-formula>.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-05-1075771-g004.tif"/>
</fig>
</sec>
<sec id="s4" sec-type="discussion"><title>Discussion</title>
<p>Signs and symptoms are an important component of a patient&#x2019;s phenotype. Extracting these phenotypic features from electronic health records and converting them to machine-readable codes makes them computable (<xref ref-type="bibr" rid="B23">23</xref>). These computable phenotypes are critical to precision medicine initiatives (<xref ref-type="bibr" rid="B24">24</xref>&#x2013;<xref ref-type="bibr" rid="B26">26</xref>). Agrawal et al. (<xref ref-type="bibr" rid="B5">5</xref>) have conceptualized clinical entity extraction as a two-step process of text span recognition followed by clinical entity normalization. Text span recognition is the identification of signs and symptoms in the free text; entity normalization is the mapping of this text to canonical signs and symptoms in an ontology such as UMLS (<xref ref-type="bibr" rid="B9">9</xref>). We have focused on an inter-rater agreement for text span annotation. For entity normalization, we depended on a look-up table that mapped text spans to concepts in neuro-ontology. We found high inter-rater concordance (unadjusted agreement) among the human annotators (approximately 89&#x0025;) with a lower concordance (unadjusted) agreement between the human annotators and the machine annotator (approximately 84&#x0025;).</p>
<p>The concordance (unadjusted agreement) for category labels was lower than the inter-rater agreement for text spans which may have been due to factors such as the use of hyphens in the free text of the EHR notes and annotator uncertainty about which types of text spans required the tabular label. The Kappa statistic (adjusted agreement) for human-human raters was between 0.77 and 0.91, and the Kappa statistic for the human-machine agreement was between 0.69 and 0.87 (<xref ref-type="fig" rid="F3">Figure&#x00A0;3A</xref>). We consider the inter-rater adjusted agreement between the human raters (0.77 to 0.91) good, especially when contrasted with the inter-rater adjusted agreement between trained neurologists eliciting patient signs and symptoms (<xref ref-type="bibr" rid="B27">27</xref>,<xref ref-type="bibr" rid="B28">28</xref>). For trained neurologists eliciting signs and symptoms such as weakness, sensory loss, ataxia, aphasia, dysarthria, and drowsiness, the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM46"><mml:mi>&#x03BA;</mml:mi></mml:math></inline-formula> statistic ranges from 0.40 to 0.70 (<xref ref-type="bibr" rid="B27">27</xref>,<xref ref-type="bibr" rid="B28">28</xref>).</p>
<p>The higher levels of agreement in this study may reflect that eliciting a sign or symptom from a patient is more difficult than annotating a sign or symptom in an EHR. Nonetheless, the adjusted agreement (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM47"><mml:mi>&#x03BA;</mml:mi></mml:math></inline-formula>) was higher in this study than in prior annotation studies (<xref ref-type="bibr" rid="B10">10</xref>,<xref ref-type="bibr" rid="B11">11</xref>), possibly reflecting the training of the annotators, the use of a neuro-ontology, the decision not to code severity or laterality of the symptoms, and the use of a sophisticated annotation tool.</p>
<p>We did not find a training effect for the human annotators across rounds (<xref ref-type="fig" rid="F4">Figures&#x00A0;4A</xref>,<xref ref-type="fig" rid="F4">B</xref>). Although the annotators met after each round and discussed discrepancies in their annotations, inter-rater adjusted and unadjusted agreement did not improve significantly between rounds. This suggests that there may be a ceiling for inter-rater agreement for text span annotation with a Kappa of 0.80 to 0.90 and that higher levels of agreement may not be possible due to the complexity of the task and random factors that are not addressable with additional training or experience. This ceiling effect for the human inter-rater agreement has implications for the potential for higher rates of inter-rater agreement between humans and machines (<xref ref-type="fig" rid="F3">Figure&#x00A0;3B</xref>). Mean inter-rater adjusted agreement for text span was higher for the human-human pairs (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM48"><mml:mi>&#x03BA;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.85</mml:mn></mml:math></inline-formula>) than the human-machine pairs (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM49"><mml:mi>&#x03BA;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.76</mml:mn></mml:math></inline-formula>). Additional training examples would likely improve the performance of the machine annotator on the text span and category label tasks. Furthermore, other neural networks are likely to outperform the convolutional neural network (CNN), which is the baseline for Prodigy. We have found that a neural network based on bidirectional encoder representations from transformers (BERT) can improve performance on the text span task by 5 to 10&#x0025; (<xref ref-type="bibr" rid="B18">18</xref>). Others have found that deep learning approaches based on BERT outperform approaches based on CNN for concept identification and extraction tasks (<xref ref-type="bibr" rid="B17">17</xref>). A ceiling effect for inter-rater agreement for annotating signs and symptoms, whether human-human or human-machine, near a <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM50"><mml:mi>&#x03BA;</mml:mi></mml:math></inline-formula> of 0.90 is likely.</p>
<p>Given the heavy documentation burden on physicians and physician burn-out attributed to electronic health records, physician documentation of signs and symptoms will likely continue as free text. Structured documentation of signs and symptoms as an alternative to free text is too burdensome in the current environment (<xref ref-type="bibr" rid="B29">29</xref>&#x2013;<xref ref-type="bibr" rid="B34">34</xref>). A medium-sized medical center with a daily inpatient census of 300 and a daily outpatient census of 2,000 generates at least 5,000 clinical notes daily or over 1.5 million notes annually (unpublished estimates based on two academic medical centers). The sheer volume of clinical notes in electronic health records makes the manual annotation of signs and symptoms impractical. Extracting signs and symptoms for precision medicine initiatives will depend on advances in natural language processing and natural language understanding.</p>
<p>Although high throughput phenotyping of electronic health records by manual methods is impractical (<xref ref-type="bibr" rid="B13">13</xref>), the manual annotation of free text in electronic health records can be used to train neural networks for phenotyping. Neural networks can also speed up the manual annotation process. The annotator Prodigy (<xref ref-type="bibr" rid="B35">35</xref>,<xref ref-type="bibr" rid="B36">36</xref>) has an annotation mode called <italic>ner.correct</italic>, which uses a trained neural network to accelerate the manual annotation of signs and symptoms.</p>
<p>With suitable training and guidelines, high levels of inter-rater agreement between human annotators for signs and symptoms are feasible. Restricting the annotation to a limited domain (e.g., neurological signs and symptoms) and restricted ontology (e.g., neuro-ontology) simplifies manual annotation. Although the inter-rater agreement between human and machine annotators was lower than between human annotators, advances in natural language processing should bring inter-rater agreement between machines and humans closer and make high throughput phenotyping of electronic health records feasible.</p>
<p>This work has limitations. The sample of clinical notes was small (five patient notes per annotation round). A larger sample of notes would have been desirable. The annotation process was restricted to neurological signs and symptoms in neurology notes. The target ontology was a limited neuro-ontology with 1600 concepts (<xref ref-type="bibr" rid="B21">21</xref>). We evaluated only one machine annotator based on a convolutional neural network. Other neural networks are likely to perform better. Our results on an inter-rater agreement might not generalize to other medical domains and ontologies. Although we had three raters for this study, we did not designate any of them as the &#x201C;gold standard,&#x201D; and we elected to calculate inter-rater agreement for each pair of raters separately. In our opinion, unadjusted agreement at the 90&#x0025; level between human raters should be considered high. Likewise, machine annotators that can reach 90&#x0025; unadjusted agreement with human annotators should be considered accurate. Because we lacked a gold standard, we chose to measure the performance of the machine annotator as concordance (unadjusted agreement) and Kappa statistic (adjusted agreement) rather than as accuracy, precision, and recall. Although we used ANOVA to assess the significance of differences in the means for adjusted and unadjusted agreement, we cannot be certain that all assumptions underlying ANOVA were met in our samples, including normality, homogeneity of variance, and independence.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability"><title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s6" sec-type="ethics-statement"><title>Ethics statement</title>
<p>The studies involving human participants were reviewed and approved by the Institutional Review Board of the University of Illinois at Chicago. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s7" sec-type="author-contributions"><title>Author contributions</title>
<p>Concept and design by DH. Data collection by DH, CO, and QH-P. Data analysis by CO and DH. Data interpretation by DH, MC, QH-P, and CO. Initial draft by DH and CO. Revisions, re-writing, and final approval by DH, CO, QH-P, and MC. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8" sec-type="funding-information"><title>Funding</title>
<p>MC acknowledges research funding from the Department of Veterans Affairs (BLR&#x0026;D Merit Award BX000467).</p>
</sec>
<sec id="s9" sec-type="COI-statement"><title>Conflict of interest</title>
<p>MC acknowledges prior support from Biogen.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer"><title>Publisher&#x0027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list><title>References</title>
<ref id="B1"><label>1.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hebbring</surname><given-names>SJ</given-names></name><name><surname>Rastegar-Mojarad</surname><given-names>M</given-names></name><name><surname>Ye</surname><given-names>Z</given-names></name><name><surname>Mayer</surname><given-names>J</given-names></name><name><surname>Jacobson</surname><given-names>C</given-names></name><name><surname>Lin</surname><given-names>S</given-names></name></person-group>. <article-title>Application of clinical text data for phenome-wide association studies (PheWASs)</article-title>. <source>Bioinformatics</source>. (<year>2015</year>) <volume>31</volume>:<fpage>1981</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv076</pub-id><pub-id pub-id-type="pmid">25657332</pub-id></citation></ref>
<ref id="B2"><label>2.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kimia</surname><given-names>AA</given-names></name><name><surname>Savova</surname><given-names>G</given-names></name><name><surname>Landschaft</surname><given-names>A</given-names></name><name><surname>Harper</surname><given-names>MB</given-names></name></person-group>. <article-title>An introduction to natural language processing: how you can get more from those electronic notes you are generating</article-title>. <source>Pediatr Emerg Care</source>. (<year>2015</year>) <volume>31</volume>:<fpage>536</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1097/PEC.0000000000000484</pub-id><pub-id pub-id-type="pmid">26148107</pub-id></citation></ref>
<ref id="B3"><label>3.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alzoubi</surname><given-names>H</given-names></name><name><surname>Alzubi</surname><given-names>R</given-names></name><name><surname>Ramzan</surname><given-names>N</given-names></name><name><surname>West</surname><given-names>D</given-names></name><name><surname>Al-Hadhrami</surname><given-names>T</given-names></name><name><surname>Alazab</surname><given-names>M</given-names></name></person-group>. <article-title>A review of automatic phenotyping approaches using electronic health records</article-title>. <source>Electronics</source>. (<year>2019</year>) <volume>8</volume>:<fpage>1235</fpage>. <pub-id pub-id-type="doi">10.3390/electronics8111235</pub-id></citation></ref>
<ref id="B4"><label>4.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shivade</surname><given-names>C</given-names></name><name><surname>Raghavan</surname><given-names>P</given-names></name><name><surname>Fosler-Lussier</surname><given-names>E</given-names></name><name><surname>Embi</surname><given-names>PJ</given-names></name><name><surname>Elhadad</surname><given-names>N</given-names></name><name><surname>Johnson</surname><given-names>SB</given-names></name></person-group>, et al. <article-title>A review of approaches to identifying patient phenotype cohorts using electronic health records</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2014</year>) <volume>21</volume>:<fpage>221</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1136/amiajnl-2013-001935</pub-id><pub-id pub-id-type="pmid">24201027</pub-id></citation></ref>
<ref id="B5"><label>5.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Agrawal</surname><given-names>M</given-names></name><name><surname>O&#x2019;Connell</surname><given-names>C</given-names></name><name><surname>Fatemi</surname><given-names>Y</given-names></name><name><surname>Levy</surname><given-names>A</given-names></name><name><surname>Sontag</surname><given-names>D</given-names></name></person-group>. <comment>Robust benchmarking for machine learning of clinical entity extraction. <italic>Machine Learning for Healthcare Conference</italic>. PMLR (2020). p. 928&#x2013;49</comment>.</citation></ref>
<ref id="B6"><label>6.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname><given-names>S</given-names></name><name><surname>Chen</surname><given-names>D</given-names></name><name><surname>He</surname><given-names>H</given-names></name><name><surname>Liu</surname><given-names>S</given-names></name><name><surname>Moon</surname><given-names>S</given-names></name><name><surname>Peterson</surname><given-names>KJ</given-names></name></person-group>, et al. <article-title>Clinical concept extraction: a methodology review</article-title>. <source>J Biomed Inform</source>. (<year>2020</year>) <volume>109</volume>:<fpage>103526</fpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2020.103526</pub-id><pub-id pub-id-type="pmid">32768446</pub-id></citation></ref>
<ref id="B7"><label>7.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Mamlin</surname><given-names>BW</given-names></name><name><surname>Heinze</surname><given-names>DT</given-names></name><name><surname>McDonald</surname><given-names>CJ</given-names></name></person-group>. <comment>Automated extraction, normalization of findings from cancer-related free-text radiology reports. <italic>AMIA Annual Symposium Proceedings</italic>. Vol. 2003. American Medical Informatics Association (2003). p. 420</comment>.</citation></ref>
<ref id="B8"><label>8.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leaman</surname><given-names>R</given-names></name><name><surname>Islamaj Do&#x011F;an</surname><given-names>R</given-names></name><name><surname>Lu</surname><given-names>Z</given-names></name></person-group>. <article-title>Dnorm: disease name normalization with pairwise learning to rank</article-title>. <source>Bioinformatics</source>. (<year>2013</year>) <volume>29</volume>:<fpage>2909</fpage>&#x2013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt474</pub-id><pub-id pub-id-type="pmid">23969135</pub-id></citation></ref>
<ref id="B9"><label>9.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bodenreider</surname><given-names>O</given-names></name></person-group>. <article-title>The unified medical language system (UMLS): integrating biomedical terminology</article-title>. <source>Nucleic Acids Res</source>. (<year>2004</year>) <volume>32</volume>:<fpage>D267</fpage>&#x2013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh061</pub-id><pub-id pub-id-type="pmid">14681409</pub-id></citation></ref>
<ref id="B10"><label>10.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andrews</surname><given-names>JE</given-names></name><name><surname>Richesson</surname><given-names>RL</given-names></name><name><surname>Krischer</surname><given-names>J</given-names></name></person-group>. <article-title>Variation of SNOMED CT coding of clinical research concepts among coding experts</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2007</year>) <volume>14</volume>:<fpage>497</fpage>&#x2013;<lpage>506</lpage>. <pub-id pub-id-type="doi">10.1197/jamia.M2372</pub-id><pub-id pub-id-type="pmid">17460128</pub-id></citation></ref>
<ref id="B11"><label>11.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hwang</surname><given-names>JC</given-names></name><name><surname>Alexander</surname><given-names>CY</given-names></name><name><surname>Casper</surname><given-names>DS</given-names></name><name><surname>Starren</surname><given-names>J</given-names></name><name><surname>Cimino</surname><given-names>JJ</given-names></name><name><surname>Chiang</surname><given-names>MF</given-names></name></person-group>. <article-title>Representation of ophthalmology concepts by electronic systems: intercoder agreement among physicians using controlled terminologies</article-title>. <source>Ophthalmology</source>. (<year>2006</year>) <volume>113</volume>:<fpage>511</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/j.ophtha.2006.01.017</pub-id><pub-id pub-id-type="pmid">16488013</pub-id></citation></ref>
<ref id="B12"><label>12.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mi&#x00F1;arro-Gim&#x00E9;nez</surname><given-names>JA</given-names></name><name><surname>Mart&#x00ED;nez-Costa</surname><given-names>C</given-names></name><name><surname>Karlsson</surname><given-names>D</given-names></name><name><surname>Schulz</surname><given-names>S</given-names></name><name><surname>G&#x00F8;eg</surname><given-names>KR</given-names></name></person-group>. <article-title>Qualitative analysis of manual annotations of clinical text with SNOMED CT</article-title>. <source>PLoS ONE</source>. (<year>2018</year>) <volume>13</volume>:<fpage>e0209547</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0209547</pub-id></citation></ref>
<ref id="B13"><label>13.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Yelugam</surname><given-names>R</given-names></name><name><surname>Azizi</surname><given-names>S</given-names></name><name><surname>Wunsch II</surname><given-names>DC</given-names></name></person-group>. <article-title>A focused review of deep phenotyping with examples from neurology</article-title>. <source>Eur Sci J</source>. (<year>2022</year>) <volume>18</volume>:<fpage>4</fpage>&#x2013;<lpage>19</lpage> <comment>(Accessed August 12, 2022)</comment>. <pub-id pub-id-type="doi">10.19044/esj.2022.v18n4p4</pub-id></citation></ref>
<ref id="B14"><label>14.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Divita</surname><given-names>G</given-names></name><name><surname>Zeng</surname><given-names>QT</given-names></name><name><surname>Gundlapalli</surname><given-names>AV</given-names></name><name><surname>Duvall</surname><given-names>S</given-names></name><name><surname>Nebeker</surname><given-names>J</given-names></name><name><surname>Samore</surname><given-names>MH</given-names></name></person-group>. <comment>Sophia: a expedient UMLS concept extraction annotator. <italic>AMIA Annual Symposium Proceedings</italic>. Vol. 2014. American Medical Informatics Association (2014). p. 467</comment>.</citation></ref>
<ref id="B15"><label>15.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Yelugam</surname><given-names>R</given-names></name><name><surname>Azizi</surname><given-names>S</given-names></name><name><surname>Carrithers</surname><given-names>MD</given-names></name><name><surname>Wunsch II</surname><given-names>DC</given-names></name></person-group>. <article-title>High throughput neurological phenotyping with metamap</article-title>. <source>Eur Sci J</source>. (<year>2022</year>) <volume>18</volume>:<fpage>37</fpage>&#x2013;<lpage>49</lpage> <comment>(Accessed August 12, 2022)</comment>. <pub-id pub-id-type="doi">10.19044/esj.2022.v18n4p37</pub-id></citation></ref>
<ref id="B16"><label>16.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arbabi</surname><given-names>A</given-names></name><name><surname>Adams</surname><given-names>DR</given-names></name><name><surname>Fidler</surname><given-names>S</given-names></name><name><surname>Brudno</surname><given-names>M</given-names></name></person-group>. <article-title>Identifying clinical terms in medical text using ontology-guided machine learning</article-title>. <source>JMIR Med Inform</source>. (<year>2019</year>) <volume>7</volume>:<fpage>e12596</fpage>. <pub-id pub-id-type="doi">10.2196/12596</pub-id>. <comment>PMID: 31094361; PMCID: PMC6533869.</comment> <pub-id pub-id-type="pmid">31094361</pub-id></citation></ref>
<ref id="B17"><label>17.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname><given-names>X</given-names></name><name><surname>Bian</surname><given-names>J</given-names></name><name><surname>Hogan</surname><given-names>WR</given-names></name><name><surname>Wu</surname><given-names>Y</given-names></name></person-group>. <article-title>Clinical concept extraction using transformers</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2020</year>) <volume>27</volume>:<fpage>1935</fpage>&#x2013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocaa189</pub-id><pub-id pub-id-type="pmid">33120431</pub-id></citation></ref>
<ref id="B18"><label>18.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Azizi</surname><given-names>S</given-names></name><name><surname>Hier</surname><given-names>D</given-names></name><name><surname>Wunsch</surname><given-names>ID</given-names></name></person-group>. <article-title>Enhanced neurologic concept recognition using a named entity recognition model based on transformers</article-title>. <source>Front Digit Health</source>. (<year>2022</year>) <volume>4</volume>:<fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.3389/fdgth.2022.1065581</pub-id></citation></ref>
<ref id="B19"><label>19.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McHugh</surname><given-names>ML</given-names></name></person-group>. <article-title>Interrater reliability: the kappa statistic</article-title>. <source>Biochem Med</source>. (<year>2012</year>) <volume>22</volume>:<fpage>276</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.11613/BM.2012.031</pub-id></citation></ref>
<ref id="B20"><label>20.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname><given-names>J</given-names></name></person-group>. <article-title>A coefficient of agreement for nominal scales</article-title>. <source>Educ Psychol Meas</source>. (<year>1960</year>) <volume>20</volume>:<fpage>37</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1177/001316446002000104</pub-id></citation></ref>
<ref id="B21"><label>21.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Brint</surname><given-names>SU</given-names></name></person-group>. <article-title>A neuro-ontology for the neurological examination</article-title>. <source>BMC Med Inform Decis Mak</source>. (<year>2020</year>) <volume>20</volume>:<fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1186/s12911-020-1066-7</pub-id><pub-id pub-id-type="pmid">31906929</pub-id></citation></ref>
<ref id="B22"><label>22.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Altinok</surname><given-names>D</given-names></name></person-group>. <source>Mastering spaCy</source>. <publisher-loc>Birmingham, UK</publisher-loc>: <publisher-name>Packt Publishing</publisher-name> (<year>2021</year>).</citation></ref>
<ref id="B23"><label>23.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hier</surname><given-names>DB</given-names></name><name><surname>Yelugam</surname><given-names>R</given-names></name><name><surname>Azizi</surname><given-names>S</given-names></name><name><surname>Wunsch</surname><given-names>DC</given-names></name></person-group>. <article-title>A focused review of deep phenotyping with examples from neurology</article-title>. <source>Eur Sci J</source>. (<year>2022</year>) <volume>18</volume>:<fpage>4</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.19044/esj.2022.v18n4p4</pub-id></citation></ref>
<ref id="B24"><label>24.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haendel</surname><given-names>MA</given-names></name><name><surname>Chute</surname><given-names>CG</given-names></name><name><surname>Robinson</surname><given-names>PN</given-names></name></person-group>. <article-title>Classification, ontology, and precision medicine</article-title>. <source>N Engl J Med</source>. (<year>2018</year>) <volume>379</volume>:<fpage>1452</fpage>&#x2013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1056/NEJMra1615014</pub-id><pub-id pub-id-type="pmid">30304648</pub-id></citation></ref>
<ref id="B25"><label>25.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robinson</surname><given-names>PN</given-names></name></person-group>. <article-title>Deep phenotyping for precision medicine</article-title>. <source>Hum Mutat</source>. (<year>2012</year>) <volume>33</volume>:<fpage>777</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1002/humu.22080</pub-id><pub-id pub-id-type="pmid">22504886</pub-id></citation></ref>
<ref id="B26"><label>26.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Collins</surname><given-names>FS</given-names></name><name><surname>Varmus</surname><given-names>H</given-names></name></person-group>. <article-title>A new initiative on precision medicine</article-title>. <source>N Engl J Med</source>. (<year>2015</year>) <volume>372</volume>:<fpage>793</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1056/NEJMp1500523</pub-id><pub-id pub-id-type="pmid">25635347</pub-id></citation></ref>
<ref id="B27"><label>27.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shinar</surname><given-names>D</given-names></name><name><surname>Gross</surname><given-names>CR</given-names></name><name><surname>Mohr</surname><given-names>JP</given-names></name><name><surname>Caplan</surname><given-names>LR</given-names></name><name><surname>Price</surname><given-names>TR</given-names></name><name><surname>Wolf</surname><given-names>PA</given-names></name></person-group>, et al. <article-title>Interobserver variability in the assessment of neurologic history and examination in the Stroke Data Bank</article-title>. <source>Arch Neurol</source>. (<year>1985</year>) <volume>42</volume>:<fpage>557</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1001/archneur.1985.04060060059010</pub-id><pub-id pub-id-type="pmid">4004598</pub-id></citation></ref>
<ref id="B28"><label>28.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldstein</surname><given-names>LB</given-names></name><name><surname>Bertels</surname><given-names>C</given-names></name><name><surname>Davis</surname><given-names>JN</given-names></name></person-group>. <article-title>Interrater reliability of the NIH stroke scale</article-title>. <source>Arch Neurol</source>. (<year>1989</year>) <volume>46</volume>:<fpage>660</fpage>&#x2013;<lpage>2</lpage>. <pub-id pub-id-type="doi">10.1001/archneur.1989.00520420080026</pub-id><pub-id pub-id-type="pmid">2730378</pub-id></citation></ref>
<ref id="B29"><label>29.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vuokko</surname><given-names>R</given-names></name><name><surname>M&#x00E4;kel&#x00E4;-Bengs</surname><given-names>P</given-names></name><name><surname>Hypp&#x00F6;nen</surname><given-names>H</given-names></name><name><surname>Lindqvist</surname><given-names>M</given-names></name><name><surname>Doupi</surname><given-names>P</given-names></name></person-group>. <article-title>Impacts of structuring the electronic health record: results of a systematic literature review from the perspective of secondary use of patient data</article-title>. <source>Int J Med Inform</source>. (<year>2017</year>) <volume>97</volume>:<fpage>293</fpage>&#x2013;<lpage>303</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijmedinf.2016.10.004</pub-id><pub-id pub-id-type="pmid">27919387</pub-id></citation></ref>
<ref id="B30"><label>30.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname><given-names>GR</given-names></name><name><surname>Friedman</surname><given-names>CP</given-names></name><name><surname>Ryan</surname><given-names>AM</given-names></name><name><surname>Richardson</surname><given-names>CR</given-names></name><name><surname>Adler-Milstein</surname><given-names>J</given-names></name></person-group>. <article-title>Variation in physicians&#x2019; electronic health record documentation and potential patient harm from that variation</article-title>. <source>J Gen Intern Med</source>. (<year>2019</year>) <volume>34</volume>:<fpage>2355</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1007/s11606-019-05025-3</pub-id><pub-id pub-id-type="pmid">31183688</pub-id></citation></ref>
<ref id="B31"><label>31.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joukes</surname><given-names>E</given-names></name><name><surname>Abu-Hanna</surname><given-names>A</given-names></name><name><surname>Cornet</surname><given-names>R</given-names></name><name><surname>de Keizer</surname><given-names>NF</given-names></name></person-group>. <article-title>Time spent on dedicated patient care and documentation tasks before and after the introduction of a structured and standardized electronic health record</article-title>. <source>Appl Clin Inform</source>. (<year>2018</year>) <volume>9</volume>:<fpage>046</fpage>&#x2013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1055/s-0037-1615747</pub-id></citation></ref>
<ref id="B32"><label>32.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosenbloom</surname><given-names>ST</given-names></name><name><surname>Denny</surname><given-names>JC</given-names></name><name><surname>Xu</surname><given-names>H</given-names></name><name><surname>Lorenzi</surname><given-names>N</given-names></name><name><surname>Stead</surname><given-names>WW</given-names></name><name><surname>Johnson</surname><given-names>KB</given-names></name></person-group>. <article-title>Data from clinical notes: a perspective on the tension between structure and flexible documentation</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2011</year>) <volume>18</volume>:<fpage>181</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1136/jamia.2010.007237</pub-id><pub-id pub-id-type="pmid">21233086</pub-id></citation></ref>
<ref id="B33"><label>33.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moy</surname><given-names>AJ</given-names></name><name><surname>Schwartz</surname><given-names>JM</given-names></name><name><surname>Chen</surname><given-names>RJ</given-names></name><name><surname>Sadri</surname><given-names>S</given-names></name><name><surname>Lucas</surname><given-names>E</given-names></name><name><surname>Cato</surname><given-names>KD</given-names></name></person-group>, et al. <article-title>Measurement of clinical documentation burden among physicians and nurses using electronic health records: a scoping review</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2021</year>) <volume>28</volume>:<fpage>998</fpage>&#x2013;<lpage>1008</lpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocaa325</pub-id><pub-id pub-id-type="pmid">33434273</pub-id></citation></ref>
<ref id="B34"><label>34.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Downing</surname><given-names>NL</given-names></name><name><surname>Bates</surname><given-names>DW</given-names></name><name><surname>Longhurst</surname><given-names>CA</given-names></name></person-group>. <article-title>Physician burnout in the electronic health record era: are we ignoring the real cause?</article-title> <source>Ann Intern Med</source>. (<year>2018</year>) <volume>169</volume>:<fpage>50</fpage>&#x2013;<lpage>1</lpage>. <pub-id pub-id-type="doi">10.7326/M18-0139</pub-id><pub-id pub-id-type="pmid">29801050</pub-id></citation></ref>
<ref id="B35"><label>35.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Musabeyezu</surname><given-names>F</given-names></name></person-group>. <comment><italic>Comparative study of annotation tools and techniques</italic> [master&#x2019;s thesis]. African University of Science and Technology (2019)</comment>.</citation></ref>
<ref id="B36"><label>36.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neves</surname><given-names>M</given-names></name><name><surname>&#x0160;eva</surname><given-names>J</given-names></name></person-group>. <article-title>An extensive review of tools for manual annotation of documents</article-title>. <source>Brief Bioinform</source>. (<year>2021</year>) <volume>22</volume>:<fpage>146</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbz130</pub-id><pub-id pub-id-type="pmid">31838514</pub-id></citation></ref></ref-list>
</back>
</article>