<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Big Data</journal-id>
<journal-title>Frontiers in Big Data</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Big Data</abbrev-journal-title>
<issn pub-type="epub">2624-909X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdata.2023.1241899</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Big Data</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Anemia detection through non-invasive analysis of lip mucosa images</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Mahmud</surname> <given-names>Shekhar</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Donmez</surname> <given-names>Turker Berk</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2351545/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Mansour</surname> <given-names>Mohammed</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2240742/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kutlu</surname> <given-names>Mustafa</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1921157/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Freeman</surname> <given-names>Chris</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2518027/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Systems Engineering, Military Technological College</institution>, <addr-line>Muscat</addr-line>, <country>Oman</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Biomedical Engineering, Sakarya University of Applied Sciences, Serdivan</institution>, <addr-line>Sakarya</addr-line>, <country>T&#x000FC;rkiye</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Mechatronics Engineering, Sakarya University of Applied Sciences, Serdivan</institution>, <addr-line>Sakarya</addr-line>, <country>T&#x000FC;rkiye</country></aff>
<aff id="aff4"><sup>4</sup><institution>Electronics and Computer Sciences, University of Southampton</institution>, <addr-line>Southampton</addr-line>, <country>United Kingdom</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Kathiravan Srinivasan, Vellore Institute of Technology, India</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Mala Rajendran, Mepco Schlenk Engineering College, India; Rakesh Chandra Joshi, Dr. A.P.J. Abdul Kalam Technical University, India</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Mohammed Mansour <email>mmansour755&#x00040;gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>10</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>6</volume>
<elocation-id>1241899</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>07</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Mahmud, Donmez, Mansour, Kutlu and Freeman.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Mahmud, Donmez, Mansour, Kutlu and Freeman</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>This paper aims to detect anemia using images of the lip mucosa, where the skin tissue is thin, and to confirm the feasibility of detecting anemia noninvasively and in the home environment using machine learning (ML). Data were collected from 138 patients, including 100 women and 38 men. Six ML algorithms: artificial neural network (ANN), decision tree (DT), k-nearest neighbors (KNN), logistic regression (LR), naive bayes (NB), and support vector machine (SVM) which are widely used in medical applications, were used to classify the collected data. Two different data types were obtained from participants&#x00027; images (RGB red color values and HSV saturation values) as features, with age, sex, and hemoglobin levels utilized to perform classification. The ML algorithm was used to analyze and classify images of the lip mucosa quickly and accurately, potentially increasing the efficiency of anemia screening programs. The accuracy, precision, recall, and F-measure were evaluated to assess how well ML models performed in predicting anemia. The results showed that NB reported the highest accuracy (96%) among the other ML models used. DT, KNN and ANN reported an accuracies of (93%), while LR and SVM had an accuracy of (79%) and (75%) receptively. This research suggests that employing ML approaches to identify anemia will help classify the diagnosis, which will then help to create efficient preventive measures. Compared to blood tests, this noninvasive procedure is more practical and accessible to patients. Furthermore, ML algorithms may be created and trained to assess lip mucosa photos at a minimal cost, making it an affordable screening method in regions with a shortage of healthcare resources.</p></abstract>
<kwd-group>
<kwd>anemia</kwd>
<kwd>machine learning</kwd>
<kwd>classification</kwd>
<kwd>support vector machine (SVM)</kwd>
<kwd>decision tree</kwd>
</kwd-group>
<counts>
<fig-count count="15"/>
<table-count count="5"/>
<equation-count count="5"/>
<ref-count count="47"/>
<page-count count="12"/>
<word-count count="7944"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Medicine and Public Health</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Anemia is characterized by a reduction of hemoglobin-containing red blood cells in the blood. The criteria for anemia, as determined by the World Health Organization (WHO), are defined as a hemoglobin level in the blood below 13 g/dL in men, 12 g/dL in women, and below 11 g/dL in pregnant women (Conrad, <xref ref-type="bibr" rid="B9">2011</xref>). The most common causes of anemia are a decrease in red blood cell production or an increase in red blood cell destruction and loss, which is higher than normal (Brown, <xref ref-type="bibr" rid="B7">1991</xref>; Aapro et al., <xref ref-type="bibr" rid="B1">2002</xref>; Martinsson et al., <xref ref-type="bibr" rid="B27">2014</xref>). Additionally, the production of malformed red blood cells in some hereditary blood diseases can also cause anemia. This results in a decrease in the average red blood cell count in the blood. The gold standard for detecting anemia is by taking intravenous blood from a venous vein and analyzing this blood by hemogram (Prefumo et al., <xref ref-type="bibr" rid="B33">2019</xref>; An et al., <xref ref-type="bibr" rid="B4">2021</xref>; Milovanovic et al., <xref ref-type="bibr" rid="B29">2022</xref>). However, invasive procedures, particularly in pregnant and pediatric groups, are painful and difficult to coordinate (Bashiri et al., <xref ref-type="bibr" rid="B6">2003</xref>). The subject must also go to a clinic to receive the relevant procedure. In light of the coronavirus pandemic that began in 2019, performing these procedures in medicine and conventional follow-up mechanisms are no longer feasible. Non-invasive anemia follow-up may provide benefits in terms of patient comfort. Although anemia has many clinical side effects, it typically progresses with pallor of the skin. Therefore, it is more easily diagnosed in areas where the skin is thin, such as the conjunctiva, lips, tongue, and oral mucosa were Wang et al. employed a smartphone application that uses the camera and various illumination sources to noninvasive check blood hemoglobin content, Tamir et al. identified anemia by examining the pallor of the eye&#x00027;s front conjunctiva, by examining the color and metadata of smartphone photographs taken on the fingernail bed, Mannio et al. estimate hemoglobin levels and Selfie Anemia, a non-invasive hemoglobin estimate smartphone app that operates under regulated lighting conditions, was created by Noriega et al. (Wang et al., <xref ref-type="bibr" rid="B46">2016</xref>; Tamir et al., <xref ref-type="bibr" rid="B43">2017</xref>; Mannino et al., <xref ref-type="bibr" rid="B24">2018</xref>; Rojas et al., <xref ref-type="bibr" rid="B36">2019</xref>). Attempts have been made to detect anemia non-invasive, with a focus on the conjunctiva analysis using various methods (Dimauro et al., <xref ref-type="bibr" rid="B10">2020</xref>; Rahman et al., <xref ref-type="bibr" rid="B34">2020</xref>; Suner et al., <xref ref-type="bibr" rid="B42">2021</xref>). However, it is extremely difficult for patients to detect anemia from a conjunctiva image captured with a simple phone camera to be used for communication with the IoT.</p>
<p>The diagnosis of anemia, in non-invasive methods using images of the conjunctiva, palm and nail bed, has its advantages. It also faces some limitations. These limitations include challenges related to accuracy due to variations in skin color, the impact of factors, the presence of medical conditions, and diversity among patients. Additionally, there are concerns regarding data quality, privacy issues costs involved, approval processes, and the need for clinical validation. It is important to remember that any diagnostic approach like this should complement judgment and other tests to ensure reliable results, in real life healthcare settings. To address these limitations effectively requires research, thorough validation procedures, and careful consideration during implementation. Using lip mucosa images for diagnosis offers unique benefits. This method is non-invasive meaning it doesn&#x00027;t cause any pain or inconvenience to patients. It&#x00027;s easily accessible and suitable for affordable screening programs making it patient friendly. By analyzing the lip mucosa, we can avoid the discomfort and possible infection risks associated with blood tests. This approach is particularly appealing to people who have a fear of needles or those who live in healthcare settings, as it promotes acceptance and participation. Further highlighting its potential as a useful and effective diagnostic approach is its adaptability for continuous monitoring and application in certain groups, such as pediatric patients. To guarantee that lip mucosa analysis is helpful in detecting illnesses like anemia, however, it needs be extensively confirmed via scientific research and clinical trials.</p>
<p>In general, machine learning (ML) refers to computer techniques that automatically find approaches and parameters to arrive at the best solution to a problem, as opposed to being pre-programmed by a person to propose a predetermined solution. This learning process is classified as a branch of Artificial Intelligence (AI), which simulates a component of human intellect and can be used for intelligent goals. A crucial component of the ML methodology is the technique used to carry out classification, regression, clustering, or prescriptive modeling. These techniques can be separated into supervised and unsupervised strategies. This research is the first to use lip mucous images for anemia detection in the literature. The aim of this paper is to detect anemia using images of the lip mucous, where the skin tissue is thin, and to confirm the feasibility of non-invasive anemia detection in the home environment. This will be achieved by developing classical ML algorithms trained using the collected patient data and invasive blood values. Six widely used ML algorithms in medical applications, namely artificial neural network (ANN), decision tree (DT), k-nearest neighbors (KNN), logistic regression (LR), naive bayes (NB), and support vector machine (SVM) (Kambeitz et al., <xref ref-type="bibr" rid="B19">2015</xref>; Arbabshirani et al., <xref ref-type="bibr" rid="B5">2017</xref>), were employed to classify the collected data. Subsequently, important statistical metrics were used to evaluate the performance of the algorithms. Participant&#x00027;s data of age between 16 and 76 was used and just confirmed anaemia cases as well were chosen. Using the lips moucas images, for diagnosis. Two extracted features from lip moucas images (RGB and HSV), age, sex and haemoglobin level, the experiment was done to predict anaemia using machine learning algorithms.</p>
<p>The accuracy of anemia diagnosis based on lip mucosa analysis can be significantly impacted by a variety of factors related to lip appearance and conditions, including lip pallor caused by vitamin B12 deficiency, dark spots, uneven skin tone, and abnormalities like scaly or thick lips, lip sores, or leukoderma. These elements may alter the hue or create distortions in the pictures of the lips, which may cause misunderstandings. The development of reliable algorithms that take into consideration these lip-related alterations and distinguish between true anemia-related color changes and those resulting from other lip disorders is essential for effective anemia identification. Furthermore, to reduce the possibility of misdiagnosis and guarantee accurate findings, clinical judgment and a comprehensive patient evaluation should be used in conjunction with computerized diagnosis. The examination of the lip mucosa for the diagnosis of anemia has obstacles from smoking and the use of lipstick, which requires identifying possible problems with changes in lip color, changes in texture and the variability of data due to these variables. The data set must be supplemented with photos of smokers and lipstick users, and machine learning algorithms must be modified to account for these particular traits. These steps are necessary to improve accuracy and sensitivity. Even in populations with different lip diseases caused by smoking or lipstick usage, the reliability and usefulness of the diagnostic procedure can be guaranteed by carefully weighing these difficulties and implementing mitigation measures.</p>
<p>The study is important due to its utilization of AI through ML, which is widely used for medical prediction and diagnosis (Kononenko, <xref ref-type="bibr" rid="B20">2001</xref>; Jiang et al., <xref ref-type="bibr" rid="B17">2017</xref>; Christodoulou et al., <xref ref-type="bibr" rid="B8">2019</xref>; Alballa and Al-Turaiki, <xref ref-type="bibr" rid="B2">2021</xref>). Furthermore, this study offers a non-invasive method that is more convenient and accessible to patients than blood tests, especially in areas where these tests are not readily available. Early detection of anemia is crucial as it can prevent serious health consequences such as fatigue, weakness, and decreased immune function. Moreover, developing and training ML algorithms to analyze lip mucous images is a cost-effective screening method in areas where healthcare resources are limited. Furthermore, the objective analysis of lip mucous images by ML algorithms reduces the risk of human error or bias in the diagnosis of anemia. Lastly, the ML algorithm could be used to quickly and accurately analyze large volumes of lip mucous images, which can potentially increase the efficiency of anemia screening programs once validated. In comparison to previous research and conventional diagnostic techniques, the approach outlined in our work shows potential for greatly improving the accuracy and sensitivity of anemia identification. The use of machine learning algorithms with features selection, the inclusion of different data kinds and demographic data, a sizable and diverse dataset, and thorough evaluation utilizing performance measures are all credited with this increase. Furthermore, lip mucosa analysis&#x00027;s non-invasiveness, affordability, accessibility, and capability for continuous monitoring should result in improved patient compliance and faster anemia identification. Though rigorous validation, comparisons with other diagnostic techniques, and clinical trials are essential stages in proving its superiority in actual clinical practice, they are not sufficient by themselves to support these assertions.</p>
<p>The remaining parts of the article are planned as follows: Section 2 provides a review of anemia detection using non-invasive methods. Section 3 explains the methodology. Section 4 outlines and presents the results. Section 5 presents the discussion. The last section concludes the article and suggests directions for future research.</p>
</sec>
<sec id="s2">
<title>Related work</title>
<p>In the field of biomedical data classification, numerous studies have been published in the last decade, providing a foundation for the current research. This section will thus focus on the details of previously applied non-invasive methods for anemia detection, including conjunctiva and fingertip analysis, as well as the ML models employed in these studies.</p>
<p>Suner et al. aimed to detect anemia in patients in the emergency room at a training and research hospital (Suner et al., <xref ref-type="bibr" rid="B42">2021</xref>). In the first stage, images of both conjunctiva of 142 patients were taken with a smartphone. From each image, a region of interest was selected that targeted the palpebral conjunctiva. Image-based parameters were extracted and used in step-by-step regression analyses to develop a predictive model of predicted hemoglobin (HBc). In phase 2, a validation model was created with data from 202 new emergency room patients. The final model, based on all 344 patients, was tested for accuracy of anemia and transfusion thresholds.</p>
<p>Rojas et al. designed an application called selfienemia to estimate hemoglobin levels under controlled lighting conditions (Rojas et al., <xref ref-type="bibr" rid="B36">2019</xref>). After taking a photo and processing it in the app, a colorimetric analysis is performed using a mathematical model from the cloud service. A special camera was used outside the application for better control of external conditions in this prototype. Sixty-four tongue images and 64 conjunctival images were taken, and the results of the application were compared with traditional large blood count (CBC), which is considered the gold standard test for diagnosing anemia. In the analysis of tongue images, the results were 91.89% sensitive and 85.18% specific, and in the analysis of palpebral conjunctiva, the results were 91.89% sensitive and 70.34% specific.</p>
<p>Sevani et al. aimed to support the process of detecting anemia using conjunctival pallor with a smartphone camera (Sevani et al., <xref ref-type="bibr" rid="B40">2018</xref>). They applied the K-Means clustering method to analyze the pixels of conjunctival images represented by digital characters in RGB formats. They compared the test results obtained from this application with laboratory results, demonstrating that the method provided an accuracy of 90%. Mannino et al. estimated hemoglobin levels by analyzing the color and metadata of nail bed photos taken with a smartphone (Mannino et al., <xref ref-type="bibr" rid="B24">2018</xref>). In their study of 100 people, they were able to test anemia with a sensitivity of 97%.</p>
<p>Hasan et al. utilized video images of fingertip captured with a smartphone camera under a flashlight and trained ANN to predict HGB levels non-invasively (Hasan et al., <xref ref-type="bibr" rid="B15">2018</xref>). Red, green, and blue pixel densities were calculated in 100 blocks of fields in each frame in 10 seconds (300 frames) of recorded video by 75 adults, and this method was applied to all 300 frames. ANN was then used to develop a derived model for predicting hemoglobin levels. They found that there was a correlation of 0.93 between the model and gold standard hemoglobin levels in their sample of patients aged 20 to 56 years.</p>
<p>Tamir et al. developed an android application to capture a photo of the anterior eye conjunctiva with a smartphone camera in suitable lighting conditions with appropriate resolution (Tamir et al., <xref ref-type="bibr" rid="B43">2017</xref>). These images were then processed to obtain spectra of the conjunctival color and RGB components, which were compared to a threshold to determine anemia. In their study of 19 subjects whose hemoglobin levels were known, they compared the values of 15 people with the laboratory results and correctly identified them at a rate of 78.9%.</p>
<p>Li et al. proposed a novel method with dynamic spectrum, using a spectrograph with a computer to scan the transmission spectrum of fingertip (Li et al., <xref ref-type="bibr" rid="B21">2017</xref>). An average prediction correlation coefficient (R) of 0.8399 is achieved in their experiment. Using a light source, Wang et al. performed chromatic analysis of images taken of the patient&#x00027;s finger, using an application called HemaApp (Wang et al., <xref ref-type="bibr" rid="B46">2016</xref>). When they analytically evaluated 31 patients between the ages of 6 and 77, they obtained a correlation of 0.82 with the blood test. HemaApp showed a sensitivity of 85.7% and a specificity of 76.5% in anemia screening.</p>
<p>Most previous methods for detecting anemia have utilized data such as conjunctival images, human nails, and fingertip. ML has been used in some of these methods to develop both invasive and noninvasive approaches for detecting anemia. However, there is still a need to improve the performance of these methods by incorporating new data and employing more straightforward approaches. This study is particularly important as it is the first to utilize lip mucous images for predicting anemia using ML. The study&#x00027;s use of ML techniques, which has been widely employed in the medical field for prediction and diagnosis (Kononenko, <xref ref-type="bibr" rid="B20">2001</xref>; Jiang et al., <xref ref-type="bibr" rid="B17">2017</xref>; Christodoulou et al., <xref ref-type="bibr" rid="B8">2019</xref>; Alballa and Al-Turaiki, <xref ref-type="bibr" rid="B2">2021</xref>), is what provides it with its significance.</p>
</sec>
<sec sec-type="methods" id="s3">
<title>Methodology</title>
<p>The classification problem aims to detect anemia using a dataset of collected lip images. The study begined by building a dataset that includes two types of lip images: healthy and anemic. Data preparation was then performed, followed by the using of ML models for classification, which were evaluated to determine the best model.</p>
<sec>
<title> Participants</title>
<p>Following the Sakarya University Ethical Approval (E-71522473-05.01.04-74571-458), participants were recruited between November and December 2021. The study group consisted of participants residing in the Pamukova District of Sakarya Province. Data were collected from 138 participants, including 100 women and 38 men by a team of experienced medical doctors. Smokers images have not been included in this study. According to the WHO criteria, 23 women and 6 men were diagnosed with anemia (Organization, <xref ref-type="bibr" rid="B47">2011</xref>). The range of hemoglobin level 129 level of healthy individuals was 121 to 167 grams per liter (g/L). The study group was selected using the convenience sampling method (Patton, <xref ref-type="bibr" rid="B31">2005</xref>). The range of haemoglobin level for male was 133&#x02013;167 g/L with ages 16&#x02013;68 years old between and the range of haemoglobin level for female was 121&#x02013;158 g/L with ages between 11 and 76 years old. The range of anemia patient was 80&#x02013;130 grams per liter (g/L). Demographic variables such as gender and age ranges are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Participants demographic values.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="center" colspan="2"><bold>Variable</bold></th>
<th valign="top" align="center" colspan="2"><bold>Frequency</bold></th>
<th valign="top" align="center"><bold>Percentage of anemia patients</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<td valign="top" align="left" colspan="2"></td>
<td valign="top" align="center"><bold>Healthy</bold></td>
<td valign="top" align="center"><bold>Anemia</bold></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">Gender</td>
<td valign="top" align="center">Female</td>
<td valign="top" align="center">77</td>
<td valign="top" align="center">23</td>
<td valign="top" align="center">72.46</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Male</td>
<td valign="top" align="center">32</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">27.54</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">18&#x02013;30</td>
<td valign="top" align="center" colspan="2">53</td>
<td valign="top" align="center">38.4</td>
</tr>
 <tr>
<td valign="top" align="left">Age</td>
<td valign="top" align="center">30&#x02013;50</td>
<td valign="top" align="center" colspan="2">36</td>
<td valign="top" align="center">26</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">50</td>
<td valign="top" align="center" colspan="2">49</td>
<td valign="top" align="center">35.6</td>
</tr></tbody>
</table>
</table-wrap></sec>
<sec>
<title>Experimental Setup</title>
<p>The experimental setup shown in <xref ref-type="fig" rid="F1">Figure 1</xref> was designed to measure the facial features of the participants, it is designed to display only the lip area with the help of an adjustable frame. A Canon camera EOS 2000D was used for collecting the lip moucas images. The camera&#x00027;s flash was used for lighting. Examples of collected lip mucosa images for healthy and non healthy one are shown in <xref ref-type="fig" rid="F2">Figures 2</xref>, <xref ref-type="fig" rid="F3">3</xref> respectively.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>System setup: (1) camera, (2) frame.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Healthy person representation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0002.tif"/>
</fig>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Anemia patient.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0003.tif"/>
</fig></sec>
<sec>
<title>Software Setup</title>
<p>A custom Python application was developed to process and analyze the data collected from the participants. Firstly, the lip contour of each participant was obtained using corner detection, thresholding, and framing (see <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref> for lips pallor examples). Next, the digital image within the frame was converted to RGB and HSV formats. Finally, classical ML algorithms were employed to perform the classification task. These steps will be discussed in detail in this section.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Healthy lips pallor with Hb of 137 grams per liter (g/L).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Anemia patient lips pallor with Hb of 116 grams per liter (g/L).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0005.tif"/>
</fig></sec>
<sec>
<title>Data processing and preparing</title>
<p>Feature extraction and classification of the images and hemoglobin data obtained using the setup are explained in <xref ref-type="fig" rid="F6">Figure 6</xref>. A high-resolution image (4896x2752) of the participants was captured using a camera. The images were collected on the same day as the blood sampling. Thus, the data duration and the number of samples taken from each subject are close but not the same. Two different data types were obtained from participants&#x00027; images as features: RGB (Red, Green, Blue) red color values and HSV saturation values, along with age, sex, and hemoglobin levels, which were used for classification. Feature extraction identified different features of the image that were sent to the classification algorithm, thereby increasing the classification success.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Schematic of data processing and ML application.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0006.tif"/>
</fig></sec>
<sec>
<title>Data Normalization</title>
<p>In the absence of normalization adjustments, the larger scale variable will totally dominate ML algorithms&#x00027; attempts to predict trends. Numerous ML methods demand statistical rescaling of their input variables to prevent sacrificing numerical stability (Trentin, <xref ref-type="bibr" rid="B45">2015</xref>). In order to improve the model&#x00027;s fit to the supplied data, minimum maximum normalization techniques are used (Senvar and Sennaroglu, <xref ref-type="bibr" rid="B39">2016</xref>). According to it, all attributes are equally important in terms of size (Singh and Singh, <xref ref-type="bibr" rid="B41">2020</xref>). The unnormalized 145 data are linearly adjusted by the normalization technique to a defined lower and upper bound (Han et al., <xref ref-type="bibr" rid="B14">2011</xref>). Typically, the dataset is rescaled to lie between 0 and 1 or -1 and 1. In this study, the minimum maximum normalization technique with a [0,1] scale was examined. The five inputs were normalized to the value between [0 and 1]. Equation 1 demonstrates the process used to transform raw data into normalized data, where X stands for the real data, <italic>X</italic><sub><italic>min</italic></sub> represents the lowest value found in the dataset of all X values, and <italic>X</italic><sub><italic>max</italic></sub> reflects the highest X-value found in the dataset. Thus, <italic>X</italic><sub><italic>normalized</italic></sub> displays the normalized X value and ranges from 0 to 1.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>z</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>Machine Learning Algorithms</title>
<p>DT is a decision-making method that has a tree structure and Random Forest (RF) is an ensemble classifier that consists of many decision trees and outputs the class that is the mode of the class&#x00027;s output by individual trees. It improves predictive accuracy with average and reduces over fitting (Gnanapriya et al., <xref ref-type="bibr" rid="B12">2010</xref>; Schonlau and Zou, <xref ref-type="bibr" rid="B38">2020</xref>; Mansour et al., <xref ref-type="bibr" rid="B26">2023</xref>). In this study, the DT were constructed with a maximum depth of 1 for estimating the three joints moment. SVM is a popular supervised ML algorithm used for both classification and regression tasks. It is particularly effective in solving complex problems with high-dimensional data. SVM aims to find an optimal hyperplane that separates the data points of different classes, maximizing the margin between the classes.</p>
<p>For supervised classification and regression applications, the KNN algorithm is a non-parametric, instance-based learning technique. KNN is adaptable and useful for a variety of tasks because, unlike many other machine learning algorithms, it does not make firm assumptions about the distribution of the underlying data. The foundation of KNN is the idea that data points with comparable properties are more likely to fall into one class or display similar goal values. For supervised classification problems, the NB algorithm is a probabilistic ML technique. Although NB is straightforward and makes the &#x0201C;naive&#x0201D; premise of feature independence, it has been successful in a number of fields. In order to shed light on the NB algorithm&#x00027;s inner workings and demonstrate its adaptability for practical applications. A popular linear model in ML for binary and multi-class classification applications is LR. Despite being straightforward, it acts as a crucial building element for a number of sophisticated procedures.</p>
<p>Compared with previous ML techniques, ANN algorithm is widely applied for biomedical data classification. As lack of large data sets in the healthcare and interconnected complex relationships between the individual biological components have encouraged the scientific research community to integrate ANN models. It adapt within nonlinear boundaries and is efficient to provide better classification. ANN has the ability to learn from continuous data and update the model, which is not present in other ML algorithms such as the decision tree. An example of a back feed-forward neural network is shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. It is a simple classification algorithm where the information is routed from input to output. The backpropagation algorithm was a widely utilized technique for training multiple-layer perceptrons. The multilayered layer perceptron&#x00027;s weight values are adjusted by approaching the intricate intersection of input and output data. The structure of the model in this study was obtained when the input, output, and one hidden layer are 10-5-1 (number of neurons) Input, hidden, and output layers, respectively, used Relu, Relu, and sigmoid activation functions. The ANN model were applied with various hidden layer neurons utill it reaches its best results, it was trained from 5 to 10 neurons. The training cycle was 100 and Adam optimizer was used.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Feed-forward ANN.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0007.tif"/>
</fig>
<p>The performance of ML algorithms is critical to their usefulness and effectiveness in solving real-world problems (Hossin and Sulaiman, <xref ref-type="bibr" rid="B16">2015</xref>; Santafe et al., <xref ref-type="bibr" rid="B37">2015</xref>; Mansour et al., <xref ref-type="bibr" rid="B26">2023</xref>). High-performing algorithms can make more accurate predictions, process data more efficiently, scale to handle large datasets, and be more interpretable, leading to better decision-making and improved outcomes. Classification models are widely used in ML to predict outcomes based on a set of characteristics. To evaluate the performance of such models, several commonly used metrics are available. The choice of metric depends on the nature of the problem, class balance, and desired outcome of the model (Prati et al., <xref ref-type="bibr" rid="B32">2011</xref>; Santafe et al., <xref ref-type="bibr" rid="B37">2015</xref>; Tharwat, <xref ref-type="bibr" rid="B44">2021</xref>). Accuracy is a simple metric that measures the proportion of correct predictions out of all predictions made by the model (Folorunso et al., <xref ref-type="bibr" rid="B11">2022</xref>). However, it may not be the best choice when there is class imbalance in the data. Precision measures the proportion of true positives out of all positive predictions made by the model and is useful when the goal is to minimize false positives (Juba and Le, <xref ref-type="bibr" rid="B18">2019</xref>; Miao and Zhu, <xref ref-type="bibr" rid="B28">2022</xref>). Recall, on the other hand, measures the proportion of true positives out of all the actual positive examples in the data, and is useful when the goal is to minimize false negatives (Ali et al., <xref ref-type="bibr" rid="B3">2022</xref>; Miao and Zhu, <xref ref-type="bibr" rid="B28">2022</xref>). F1 score is the harmonic mean of precision and recall and is useful to balance the importance of both (Hossin and Sulaiman, <xref ref-type="bibr" rid="B16">2015</xref>). AUC measures the ability of the model to distinguish between positive and negative examples and is useful when identifying the best threshold to separate positive and negative examples (Tharwat, <xref ref-type="bibr" rid="B44">2021</xref>). Finally, the confusion matrix provides a more detailed view of the performance of the model than any single metric by displaying the number of true positives, true negatives, false positives, and false negatives for the given model (Haghighi et al., <xref ref-type="bibr" rid="B13">2018</xref>; Liang, <xref ref-type="bibr" rid="B22">2022</xref>). Overall, the performance evaluation of a ML model is essential to determine its effectiveness and identify areas for improvement. A combination of metrics should be used to evaluate the model, and the context of the problem should be considered when choosing a metric.</p>
</sec>
</sec>
<sec id="s4">
<title>Results and discussion</title>
<p>The experiments were performed using the Python programming language and its Keras and Tensorflow libraries. This research classifies anemia by applying the input characteristics produced with the extraction of RGB (Red value), HSV (Saturation), age, gender, hemoglobin levels. In the classification process, the spilt data function was applied to separate the data into train, val and test data where 60% (82) of the dataset was randomly assigned as train data, 20% (28) were randomly assigned as validate and 20% (28) as test data. The flow diagram describing the processing and classification is shown in <xref ref-type="fig" rid="F8">Figure 8</xref>, where the process starts with the collection of patient data and images, the cropping of the image, feature extraction and the preparation of the data, and finally classification and prediction of anemia cases and evaluation of ML models. In the classification process, six algorithms were applied; KNN, NB, LR, SVM, DT, and ANN algorithms, which are frequently applied in the anemia detection literature, were used. To evaluate the classification algorithms used in this study; different methods were used. These methods are accuracy, precision, recall, and F1 score.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Flowchart of data classification.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0008.tif"/>
</fig>
<p>The test data confusion matrix prediction evaluation for SVM, DT, ANN, KNN, NB and LR are shown in <xref ref-type="fig" rid="F10">Figures 10</xref>&#x02013;<xref ref-type="fig" rid="F15">15</xref>. Four main points to consider here, True positive (TP), False Positive (FP), True Negative (TN) and False Negative (FN). These points clarify the success of the algorithm where TP and TN are the correct predictions in the both positive and negative classes, and FP and FN are the false predictions respectively (<xref ref-type="fig" rid="F9">Figure 9</xref>). For SVM algorithm (<xref ref-type="fig" rid="F10">Figure 10</xref>), the TP was 19 and the TN was 2 from a total of 19 and 9 respectively. <xref ref-type="fig" rid="F11">Figure 11</xref> presents the confusion matrix for DT algorithm. For DT, the TP was 19 and the TN was 7 from a total of 19 and 9 respectively. <xref ref-type="fig" rid="F12">Figure 12</xref> presents the confusion matrix for ANN algorithm. For ANN, the TP was 23 and the TN was 3 from a total of 23 and 5 respectively. <xref ref-type="fig" rid="F13">Figure 13</xref> presents the confusion matrix for KNN, the TP was 22 and the TN was 4 from a total of 22 and 6 respectively. <xref ref-type="fig" rid="F14">Figure 14</xref> presents the confusion matrix for NB, the TP was 22 and the TN was 5 from a total of 22 and 6 respectively. Finally, <xref ref-type="fig" rid="F15">Figure 15</xref> presents the confusion matrix for LR, the TP was 22 and the TN was 0 from a total of 22 and 6 respectively.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0009.tif"/>
</fig>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>SVM confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0010.tif"/>
</fig>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p>DT confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0011.tif"/>
</fig>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p>ANN confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0012.tif"/>
</fig>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p>KNN confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0013.tif"/>
</fig>
<fig id="F14" position="float">
<label>Figure 14</label>
<caption><p>NB confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0014.tif"/>
</fig>
<fig id="F15" position="float">
<label>Figure 15</label>
<caption><p>LR confusion matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-06-1241899-g0015.tif"/>
</fig>
<p>The success rates of all algorithms are shown in <xref ref-type="table" rid="T2">Tables 2</xref>&#x02013;<xref ref-type="table" rid="T5">5</xref> where <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref> are the evaluation of both train and validation data. <xref ref-type="table" rid="T4">Tables 4</xref>, <xref ref-type="table" rid="T5">5</xref> present the macro and weighted average of the test data for both the positive and negative classes. In these tables and based on confusion matrices, four main parameters; accuracy (Equation 2), AUC (area under the ROC curve), precision (Equation 3), recall (Equation 4) and F1 score (Equation 5) which evaluate the six algorithms applied to classify anemia applying the input characteristic matrix of RGB (red value), HSV (saturation), age, sex, hemoglobin levels.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x0002A;</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x0002A;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Train data; accuracy, precision, recall and F1 score.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Algorithm</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1 score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">84</td>
<td valign="top" align="center">42</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">46</td>
</tr>
<tr>
<td valign="top" align="left">LR</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">82</td>
<td valign="top" align="center">87</td>
</tr>
<tr>
<td valign="top" align="left">DT</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left">ANN</td>
<td valign="top" align="center">79</td>
<td valign="top" align="center">40</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">44</td>
</tr>
<tr>
<td valign="top" align="left">KNN</td>
<td valign="top" align="center">90</td>
<td valign="top" align="center">91</td>
<td valign="top" align="center">79</td>
<td valign="top" align="center">83</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">82</td>
<td valign="top" align="center">87</td>
</tr></tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Val data; accuracy, precision, recall and F1 score.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Algorithm</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1 score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">75</td>
<td valign="top" align="center">38</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">43</td>
</tr>
<tr>
<td valign="top" align="left">LR</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">98</td>
<td valign="top" align="center">92</td>
<td valign="top" align="center">94</td>
</tr>
<tr>
<td valign="top" align="left">DT</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">86</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">89</td>
</tr>
<tr>
<td valign="top" align="left">ANN</td>
<td valign="top" align="center">75</td>
<td valign="top" align="center">38</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">43</td>
</tr>
<tr>
<td valign="top" align="left">KNN</td>
<td valign="top" align="center">89</td>
<td valign="top" align="center">86</td>
<td valign="top" align="center">81</td>
<td valign="top" align="center">83</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">98</td>
<td valign="top" align="center">92</td>
<td valign="top" align="center">94</td>
</tr></tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Test data macro averages; accuracy, precision, recall and F1 score.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Algorithm</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1 score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">75</td>
<td valign="top" align="center">86</td>
<td valign="top" align="center">87</td>
<td valign="top" align="center">61</td>
<td valign="top" align="center">60</td>
</tr>
<tr>
<td valign="top" align="left">LR</td>
<td valign="top" align="center">79</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">44</td>
</tr>
<tr>
<td valign="top" align="left">DT</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">95</td>
<td valign="top" align="center">89</td>
<td valign="top" align="center">95</td>
<td valign="top" align="center">91</td>
</tr>
<tr>
<td valign="top" align="left">ANN</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">80</td>
<td valign="top" align="center">85</td>
</tr>
<tr>
<td valign="top" align="left">KNN</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">83</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">83</td>
<td valign="top" align="center">88</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">91</td>
<td valign="top" align="center">98</td>
<td valign="top" align="center">92</td>
<td valign="top" align="center">94</td>
</tr></tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Test data weighted averages; accuracy, precision, recall and F1 score.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Algorithm</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1 score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">75</td>
<td valign="top" align="center">86</td>
<td valign="top" align="center">87</td>
<td valign="top" align="center">61</td>
<td valign="top" align="center">60</td>
</tr>
<tr>
<td valign="top" align="left">LR</td>
<td valign="top" align="center">79</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">62</td>
<td valign="top" align="center">79</td>
<td valign="top" align="center">69</td>
</tr>
<tr>
<td valign="top" align="left">DT</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">95</td>
<td valign="top" align="center">94</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">93</td>
</tr>
<tr>
<td valign="top" align="left">ANN</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">92</td>
</tr>
<tr>
<td valign="top" align="left">KNN</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">83</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">92</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">91</td>
<td valign="top" align="center">97</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">96</td>
</tr></tbody>
</table>
</table-wrap>
<p>When comparing the results obtained from the anemia classification performed using ML techniques, it was found that the disease was diagnosed with an accuracy success rate of 96% using NB, 93% using DT KNN and ANN, LM classifies the data with an accuracy success rate of 79% and SVM classifies the data with an accuracy success rate of 75%. The AUC were 50, 83, 86, 91, 95, and 96 for LR, KNN, SVM, NB, DT, and ANN respectively. Although this value seems high, higher percentages and more data-trained models are required to use it as a medical diagnosis.</p>
</sec>
<sec sec-type="discussion" id="s5">
<title>Discussion</title>
<p>In this study, extracted data from images of lip mucous were used to train a ML models to identify anemia. The results of this study are in agreement with those of previous investigations that used ML models to predict anemia using both invasive and non-invasive techniques. Accuracy, precision, recall, and F-score were used to assess how well ML models performed in predicting anemia. The models examined demonstrated high and considerable accuracy. For predicting anemia, NB reported the highest accuracy, the highest precision and F1 score and SVM reported the lowest scores.</p>
<p>Limited methods have been found in the literature to detect anemia using ML. For example, the K-means clustering technique was used to carry out conjunctival pallor image-based anemia detection, which demonstrated an accuracy of 90% when comparing the test results acquired from this application with laboratory results (Sevani et al., <xref ref-type="bibr" rid="B40">2018</xref>). ANN were used to assess images of fingertip captured with a smartphone camera to predict HGB levels non-invasively. As a result, a correlation of 0.93 was observed between the model and gold standard hemoglobin levels (Hasan et al., <xref ref-type="bibr" rid="B15">2018</xref>). Convolutional neural networks was techniques was used to carry out conjunctival image-based anemia detection with accuracy of 94% (Magdalena et al., <xref ref-type="bibr" rid="B23">2022</xref>). YOLO v5 was used to detect anemia using conjunctiva image collection with sensitivity of 71% and a specificity of 89% (Rivero-Palacio et al., <xref ref-type="bibr" rid="B35">2022</xref>). AlexNet was used to calculate total hemoglobin concentration by developing frequency-domain multidistance approach, based on a non-contact oximeter, provided data on total hemoglobin with accuracy of 87.50% (Moral and Bal, <xref ref-type="bibr" rid="B30">2020</xref>). Compared to these methods that used ML in general for anemia detection, the current study presents a non-invasive method that uses ML models to detect anemia with higher accuracy, reaching 99% using DT. This new method is simple and can be developed for the detection of real-time anemia. Several methods were found in the literature to detect anemia.</p>
<p>Anemia is a prevalent condition that affects millions of people worldwide, but it often goes undiagnosed until it becomes severe. Early detection of anemia is crucial for effective treatment, which is why the use of ML to detect anemia through lip mucosa image classification could be significant. This method provides a non-invasive and cost-effective alternative to traditional anemia screening methods such as blood tests, which can be invasive and costly, particularly in resource-limited settings. The automation of diagnosis through ML algorithms can reduce the need for expert human intervention and speed up the process of diagnosis and treatment. The potential for automation also makes this method scalable and easily accessible, allowing for widespread implementation of anemia detection tools. Our research suggests that employing ML approaches to detect anemia will aid in classifying the diagnosis, which will then help in the creation of efficient preventive measures. As a result, this research evaluates the predictive capability of several ML algorithms in addition to addressing the integration of cutting-edge technology for the prediction and diagnosis of low hemoglobin levels.</p>
<p>When compared to other human body parts, lip pallor, which is defined by the paleness or loss of natural color in the lips, provides significant advantages as a non-invasive diagnostic site. First of all, there is no need for specialist equipment or invasive procedures to examine the lips because they are quite obvious and accessible. Due to its accessibility, lip pallor is a sensible option for diagnostic techniques, enabling quick and little disruptive patient examinations. A rich vascular network with countless small blood vessels close to the surface is also present on the lips. Due to the vascular richness, variations in blood flow and oxygenation, which frequently show up as changes in lip color, may be quickly identified. The lips are perfect for identifying blood-related diseases like anemia because of the strong relationship between lip color and the circulatory system. This provides real-time information about a patient&#x00027;s health.</p>
<p>Lip pallor analysis is non-invasive, which increases patient comfort and compliance. The procedure is harmless and acceptable for people of all ages, whether it involves straightforward eye exams or more sophisticated imaging techniques. This technique also fits well with ethical standards and cultural norms because it is typically socially acceptable in all cultures to examine one&#x00027;s lips. Such acceptance may increase a patient&#x00027;s willingness to participate in lip-based diagnostic tests. The ability to capture high-resolution photographs of the lips thanks to advancements in imaging technology also makes it possible to analyze color variations and texture changes precisely. This degree of specificity is essential for identifying minute symptoms of diseases like anemia and guarantees that lip pallor analysis will always be a practical and affordable screening technique in medical settings, making it a crucial tool in the field of non-invasive diagnostics.</p>
<p>The contribution of this study to medical research is also significant, as the use of ML in medical research is still in its early stages. This study could provide a foundation for further research and development of ML based tools for anemia detection and diagnosis, potentially leading to even more accurate and effective diagnostic tools. The use of ML for anemia detection using lip mucosa image classification could have significant implications for healthcare, particularly in resource-limited settings where traditional screening methods may not be readily available. The potential for early detection, non-invasive and cost-effective screening, automation of diagnosis, contribution to medical research, and further development of diagnostic tools make this study a promising avenue for improving healthcare outcomes.</p>
</sec>
<sec sec-type="conclusions" id="s6">
<title>Conclusions</title>
<p>Images of the lip mucosa, which have thin skin tissue, were used in this study to identify anemia. Data from 138 patients, including 100 women and 38 men, were collected. Rgb red color values and hsv saturation values were obtained from participant images and used as features, along with age, sex, and hemoglobin levels, to perform classification. The efficacy of ML models in predicting anemia was tested using accuracy, precision, recall, and F score. The findings indicated that among the ML algorithms utilized, NB achieved the highest accuracy at 96%, while SVM received accuracy ratings of 75%. This research suggests that using ML to recognize anemia will aid in classifying the diagnosis, which would subsequently facilitate the development of effective preventive measures.</p>
<p>This work has established a link between anemia detection and lip mucous images, where the conjunctiva is used in general practices. This newly found method is easy to use, and it was demonstrated for the first time that it can be classified on the lip mucous. Although the accuracy of 93% appears high, more data-trained models and larger percentages are needed to use it as a diagnostic in medicine. In the future, augmented data will be used for an online classification along with deep neural networks and a mobile application.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec sec-type="ethics-statement" id="s8">
<title>Ethics statement</title>
<p>The studies involving humans were approved by Sakarya University Ethical Approval (E-71522473-05.01.04-74571-458). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.</p>
</sec>
<sec sec-type="author-contributions" id="s9">
<title>Author contributions</title>
<p>TD colleting the lip mucosa images and literature review. MM preparing the software and manuscript writing. MK preparing the images and feature extraction. CF writing the manuscript. SM writing the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s10">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aapro</surname> <given-names>M. S.</given-names></name> <name><surname>Cella</surname> <given-names>D.</given-names></name> <name><surname>Zagari</surname> <given-names>M.</given-names></name></person-group> (<year>2002</year>). <article-title>Age, anemia, and fatigue</article-title>. <source>Semin. Oncol</source>. <volume>29</volume>, <fpage>55</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1016/S0093-7754(02)70175-9</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alballa</surname> <given-names>N.</given-names></name> <name><surname>Al-Turaiki</surname> <given-names>I.</given-names></name></person-group> (<year>2021</year>). <article-title>Machine learning approaches in covid-19 diagnosis, mortality, and severity risk prediction: a review</article-title>. <source>Inform. Med</source>. 24, 100564. <pub-id pub-id-type="doi">10.1016/j.imu.2021.100564</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ali</surname> <given-names>I.</given-names></name> <name><surname>Mughal</surname> <given-names>N.</given-names></name> <name><surname>Khand</surname> <given-names>Z. H.</given-names></name> <name><surname>Ahmed</surname> <given-names>J.</given-names></name> <name><surname>Mujtaba</surname> <given-names>G.</given-names></name></person-group> (<year>2022</year>). <article-title>Resume classification system using natural language processing and machine learning techniques</article-title>. <source>Mehran Univ. Res. J Eng. Technol</source>. <volume>41</volume>, <fpage>65</fpage>&#x02013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.22581/muet1982.2201.07</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>An</surname> <given-names>R.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Man</surname> <given-names>Y.</given-names></name> <name><surname>Valentine</surname> <given-names>R. W.</given-names></name> <name><surname>Kucukal</surname> <given-names>E.</given-names></name> <name><surname>Goreke</surname> <given-names>U.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Emerging point-of-care technologies for anemia detection</article-title>. <source>Lab Chip</source> <volume>21</volume>, <fpage>1843</fpage>&#x02013;<lpage>1865</lpage>. <pub-id pub-id-type="doi">10.1039/D0LC01235A</pub-id><pub-id pub-id-type="pmid">33881041</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arbabshirani</surname> <given-names>M. R.</given-names></name> <name><surname>Plis</surname> <given-names>S.</given-names></name> <name><surname>Sui</surname> <given-names>J.</given-names></name> <name><surname>Calhoun</surname> <given-names>V. D.</given-names></name></person-group> (<year>2017</year>). <article-title>Single subject prediction of brain disorders in neuroimaging: promises and pitfalls</article-title>. <source>Neuroimage</source> <volume>145</volume>, <fpage>137</fpage>&#x02013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2016.02.079</pub-id><pub-id pub-id-type="pmid">27012503</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bashiri</surname> <given-names>A.</given-names></name> <name><surname>Burstein</surname> <given-names>E.</given-names></name> <name><surname>Sheiner</surname> <given-names>E.</given-names></name> <name><surname>Mazor</surname> <given-names>M.</given-names></name></person-group> (<year>2003</year>). <article-title>Anemia during pregnancy and treatment with intravenous iron: review of the literature</article-title>. <source>Eur. J. Obstet. Gynecol. Reprod. Biol</source>. <volume>110</volume>, <fpage>2</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1016/S0301-2115(03)00113-1</pub-id><pub-id pub-id-type="pmid">12932861</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>R. G.</given-names></name></person-group> (<year>1991</year>). <article-title>Determining the cause of anemia: general approach, with emphasis on microcytic hypochromic anemias</article-title>. <source>Postgrad. Med</source>. <volume>89</volume>, <fpage>161</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1080/00325481.1991.11700925</pub-id><pub-id pub-id-type="pmid">2020645</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Christodoulou</surname> <given-names>E.</given-names></name> <name><surname>Ma</surname> <given-names>J.</given-names></name> <name><surname>Collins</surname> <given-names>G. S.</given-names></name> <name><surname>Steyerberg</surname> <given-names>E. W.</given-names></name> <name><surname>Verbakel</surname> <given-names>J. Y.</given-names></name> <name><surname>Van Calster</surname> <given-names>B.</given-names></name></person-group> (<year>2019</year>). <article-title>A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models</article-title>. <source>J. Clin. Epidemiol</source>. <volume>110</volume>, <fpage>12</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1016/j.jclinepi.2019.02.004</pub-id><pub-id pub-id-type="pmid">30763612</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Conrad</surname> <given-names>M. E.</given-names></name></person-group> (<year>2011</year>). <source>Anemia</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Butterworths</publisher-name>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dimauro</surname> <given-names>G.</given-names></name> <name><surname>Caivano</surname> <given-names>D.</given-names></name> <name><surname>Di Pilato</surname> <given-names>P.</given-names></name> <name><surname>Dipalma</surname> <given-names>A.</given-names></name> <name><surname>Camporeale</surname> <given-names>M. G.</given-names></name></person-group> (<year>2020</year>). <article-title>A systematic mapping study on research in anemia assessment with non-invasive devices</article-title>. <source>Appl. Sci</source>. 10, 4804. <pub-id pub-id-type="doi">10.3390/app10144804</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Folorunso</surname> <given-names>S. O.</given-names></name> <name><surname>Awotunde</surname> <given-names>J. B.</given-names></name> <name><surname>Adeniyi</surname> <given-names>E. A.</given-names></name> <name><surname>Abiodun</surname> <given-names>K. M.</given-names></name> <name><surname>Ayo</surname> <given-names>F. E.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Heart disease classification using machine learning models,&#x0201D;</article-title> in <source>Informatics and Intelligent Applications: First International Conference, ICIIA 2021</source>. <publisher-loc>Ota, Nigeria</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>35</fpage>&#x02013;<lpage>49</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gnanapriya</surname> <given-names>S.</given-names></name> <name><surname>Suganya</surname> <given-names>R.</given-names></name> <name><surname>Devi</surname> <given-names>G. S.</given-names></name> <name><surname>Kumar</surname> <given-names>M. S.</given-names></name></person-group> (<year>2010</year>). <article-title>Data mining concepts and techniques</article-title>. <source>Data Knowl. Eng</source>. <volume>2</volume>, <fpage>256</fpage>&#x02013;<lpage>263</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haghighi</surname> <given-names>S.</given-names></name> <name><surname>Jasemi</surname> <given-names>M.</given-names></name> <name><surname>Hessabi</surname> <given-names>S.</given-names></name> <name><surname>Zolanvari</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Pycm: multiclass confusion matrix library in python</article-title>. <source>J. Open Sour. Software</source> <volume>3</volume>, <fpage>729</fpage>. <pub-id pub-id-type="doi">10.21105/joss.00729</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Pei</surname> <given-names>J.</given-names></name> <name><surname>Kamber</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <source>Data Mining: Concepts and Techniques</source>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Elsevier</publisher-name>.</citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hasan</surname> <given-names>M. K.</given-names></name> <name><surname>Haque</surname> <given-names>M. M.</given-names></name> <name><surname>Adib</surname> <given-names>R.</given-names></name> <name><surname>Tumpa</surname> <given-names>J. F.</given-names></name> <name><surname>Begum</surname> <given-names>A.</given-names></name> <name><surname>Love</surname> <given-names>R. R.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Smarthelp: Smartphone-based hemoglobin level prediction using an artificial neural network,&#x0201D;</article-title> in <source>AMIA Annual Symposium Proceedings</source>. <publisher-loc>Washington, D.C.</publisher-loc>: <publisher-name>American Medical Informatics Association, 535</publisher-name>.<pub-id pub-id-type="pmid">30815094</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hossin</surname> <given-names>M.</given-names></name> <name><surname>Sulaiman</surname> <given-names>M. N.</given-names></name></person-group> (<year>2015</year>). <article-title>A review on evaluation metrics for data classification evaluations</article-title>. <source>Int. J. Data Min. Knowl. Manag. Process</source> <volume>5</volume>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.5121/ijdkp.2015.5201</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>F.</given-names></name> <name><surname>Jiang</surname> <given-names>Y.</given-names></name> <name><surname>Zhi</surname> <given-names>H.</given-names></name> <name><surname>Dong</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Artificial intelligence in healthcare: past, present and future</article-title>. <source>Stroke Vascular Neurol</source>. <volume>2</volume>, <fpage>4</fpage>. <pub-id pub-id-type="doi">10.1136/svn-2017-000101</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Juba</surname> <given-names>B.</given-names></name> <name><surname>Le</surname> <given-names>H. S.</given-names></name></person-group> (<year>2019</year>). <article-title>Precision-recall versus accuracy and the role of large data sets</article-title>. <source>Proc. AAAI Conference Artif. Intell</source>. <volume>33</volume>, <fpage>4039</fpage>&#x02013;<lpage>4048</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.33014039</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kambeitz</surname> <given-names>J.</given-names></name> <name><surname>Kambeitz-Ilankovic</surname> <given-names>L.</given-names></name> <name><surname>Leucht</surname> <given-names>S.</given-names></name> <name><surname>Wood</surname> <given-names>S.</given-names></name> <name><surname>Davatzikos</surname> <given-names>C.</given-names></name> <name><surname>Malchow</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Detecting neuroimaging biomarkers for schizophrenia: a meta-analysis of multivariate pattern recognition studies</article-title>. <source>Neuropsychopharmacology</source> <volume>40</volume>, <fpage>1742</fpage>&#x02013;<lpage>1751</lpage>. <pub-id pub-id-type="doi">10.1038/npp.2015.22</pub-id><pub-id pub-id-type="pmid">25601228</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kononenko</surname> <given-names>I.</given-names></name></person-group> (<year>2001</year>). <article-title>Machine learning for medical diagnosis: history, state of the art and perspective</article-title>. <source>Artif. Intell. Med</source>. <volume>23</volume>, <fpage>89</fpage>&#x02013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1016/S0933-3657(01)00077-X</pub-id><pub-id pub-id-type="pmid">11470218</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Xu</surname> <given-names>S.</given-names></name> <name><surname>Zhou</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Lin</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Noninvasive hemoglobin measurement based on optimizing dynamic spectrum method</article-title>. <source>Spectrosc. Lett</source>. <volume>50</volume>, <fpage>164</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1080/00387010.2017.1302481</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <source>Confusion Matrix: Machine Learning. POGIL Activity Clearinghouse</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://pac.pogil.org/index.php/pac/article/view/304">https://pac.pogil.org/index.php/pac/article/view/304</ext-link></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Magdalena</surname> <given-names>R.</given-names></name> <name><surname>Saidah</surname> <given-names>S.</given-names></name> <name><surname>Da&#x00027;wan</surname> <given-names>I.</given-names></name> <name><surname>Ubaidah</surname> <given-names>S.</given-names></name> <name><surname>Fuadah</surname> <given-names>Y. N.</given-names></name> <name><surname>Herman</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Convolutional neural network for anemia detection based on conjunctiva palpebral images</article-title>. <source>Jurnal Teknik Informatika (Jutif)</source> <volume>3</volume>, <fpage>349</fpage>&#x02013;<lpage>354</lpage>. <pub-id pub-id-type="doi">10.20884/1.jutif.2022.3.2.197</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mannino</surname> <given-names>R. G.</given-names></name> <name><surname>Myers</surname> <given-names>D. R.</given-names></name> <name><surname>Tyburski</surname> <given-names>E. A.</given-names></name> <name><surname>Caruso</surname> <given-names>C.</given-names></name> <name><surname>Boudreaux</surname> <given-names>J.</given-names></name> <name><surname>Leong</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Smartphone app for non-invasive detection of anemia using only patient-sourced photos</article-title>. <source>Nat. Commun</source>. <volume>9</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-018-07262-2</pub-id><pub-id pub-id-type="pmid">30514831</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mansour</surname> <given-names>M.</given-names></name> <name><surname>Demirsoy</surname> <given-names>M. S.</given-names></name> <name><surname>Kutlu</surname> <given-names>M. C.</given-names></name></person-group> (<year>2023</year>). <article-title>Kidney segmentations using cnn models</article-title>. <source>J. Smart Syst. Res</source>. <volume>4</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.58769/joinssr.1175622</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mansour</surname> <given-names>M.</given-names></name> <name><surname>Serbest</surname> <given-names>K.</given-names></name> <name><surname>Kutlu</surname> <given-names>M.</given-names></name> <name><surname>Cilli</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Estimation of lower limb joint moments based on the inverse dynamics approach: a comparison of machine learning algorithms for rapid estimation,&#x0201D;</article-title> in <source>Medical</source> &#x00026; <italic>Biological Engineering</italic> &#x00026; <italic>Computing</italic>, <fpage>1</fpage>&#x02013;<lpage>24</lpage>.<pub-id pub-id-type="pmid">37561330</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martinsson</surname> <given-names>A.</given-names></name> <name><surname>Andersson</surname> <given-names>C.</given-names></name> <name><surname>Andell</surname> <given-names>P.</given-names></name> <name><surname>Koul</surname> <given-names>S.</given-names></name> <name><surname>Engstr&#x000F6;m</surname> <given-names>G.</given-names></name> <name><surname>Smith</surname> <given-names>J. G.</given-names></name></person-group> (<year>2014</year>). <article-title>Anemia in the general population: prevalence, clinical correlates and prognostic impact</article-title>. <source>Eur. J. Epidemiol</source>. <volume>29</volume>:<fpage>489</fpage>&#x02013;<lpage>498</lpage>. <pub-id pub-id-type="doi">10.1007/s10654-014-9929-9</pub-id><pub-id pub-id-type="pmid">24952166</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miao</surname> <given-names>J.</given-names></name> <name><surname>Zhu</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Precision-recall curve (prc) classification trees</article-title>. <source>Evol. Intell</source>. <volume>15</volume>, <fpage>1545</fpage>&#x02013;<lpage>1569</lpage>. <pub-id pub-id-type="doi">10.1007/s12065-021-00565-2</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Milovanovic</surname> <given-names>T.</given-names></name> <name><surname>Dragasevic</surname> <given-names>S.</given-names></name> <name><surname>Nikolic</surname> <given-names>A. N.</given-names></name> <name><surname>Markovic</surname> <given-names>A. P.</given-names></name> <name><surname>Lalosevic</surname> <given-names>M. S.</given-names></name> <name><surname>Popovic</surname> <given-names>D. D.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Anemia as a problem: Gp approach</article-title>. <source>Digest. Dis</source>. <volume>40</volume>, <fpage>370</fpage>&#x02013;<lpage>375</lpage>. <pub-id pub-id-type="doi">10.1159/000517579</pub-id><pub-id pub-id-type="pmid">34098557</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Moral</surname> <given-names>O. T.</given-names></name> <name><surname>Bal</surname> <given-names>U.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Non-contact total hemoglobin estimation using a deep learning model,&#x0201D;</article-title> in <source>2020 4th International Symposium on Multidisciplinary Studies and Innovative Technologies (ISMSIT)</source>. <publisher-loc>Istanbul</publisher-loc>: <publisher-name>IEEE</publisher-name>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Patton</surname> <given-names>M. Q.</given-names></name></person-group> (<year>2005</year>). <source>Qualitative Research</source>. <publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage</publisher-name>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prati</surname> <given-names>R. C.</given-names></name> <name><surname>Batista</surname> <given-names>G. E.</given-names></name> <name><surname>Monard</surname> <given-names>M. C.</given-names></name></person-group> (<year>2011</year>). <article-title>A survey on graphical methods for classification predictive performance evaluation</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <volume>23</volume>, <fpage>1601</fpage>&#x02013;<lpage>1618</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2011.59</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prefumo</surname> <given-names>F.</given-names></name> <name><surname>Fichera</surname> <given-names>A.</given-names></name> <name><surname>Fratelli</surname> <given-names>N.</given-names></name> <name><surname>Sartori</surname> <given-names>E.</given-names></name></person-group> (<year>2019</year>). <article-title>Fetal anemia: diagnosis and management</article-title>. <source>Best Pract. Res. Clin. Obstet. Gynaecol</source>. <volume>58</volume>, <fpage>2</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1016/j.bpobgyn.2019.01.001</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rahman</surname> <given-names>M.</given-names></name> <name><surname>Nakib</surname> <given-names>N.</given-names></name> <name><surname>Sadik</surname> <given-names>S. A.</given-names></name> <name><surname>Biswas</surname> <given-names>A.</given-names></name> <name><surname>Tazul</surname> <given-names>R. B.</given-names></name></person-group> (<year>2020</year>). <source>Point of Care Detection of Anemia in Non-Invasive Manner by Using Image Processing and Convolutional Neural Network With Mobile Devices</source> (PhD thesis). <publisher-loc>Dhaka</publisher-loc>: <publisher-name>Brac University</publisher-name>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rivero-Palacio</surname> <given-names>M.</given-names></name> <name><surname>Alfonso-Morales</surname> <given-names>W.</given-names></name> <name><surname>Caicedo-Bravo</surname> <given-names>E.</given-names></name></person-group> (<year>2022</year>). <article-title>Anemia detection using a full embedded mobile application with yolo algorithm</article-title>. <source>Commun. Comput. Inf. Sci</source>. <volume>1471</volume>, <fpage>3</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-91308-3_1</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rojas</surname> <given-names>P. M. W.</given-names></name> <name><surname>Noriega</surname> <given-names>L. A. M.</given-names></name> <name><surname>Silva</surname> <given-names>A. S.</given-names></name></person-group> (<year>2019</year>). <article-title>Hemoglobin screening using cloud based mobile photography applications</article-title>. <source>Ingenier&#x00027;&#x00131;a y Universidad</source> <volume>23</volume>, <fpage>2</fpage>. <pub-id pub-id-type="doi">10.11144/Javeriana.iyu23-2.hsuc</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santafe</surname> <given-names>G.</given-names></name> <name><surname>Inza</surname> <given-names>I.</given-names></name> <name><surname>Lozano</surname> <given-names>J. A.</given-names></name></person-group> (<year>2015</year>). <article-title>Dealing with the evaluation of supervised classification algorithms</article-title>. <source>Artif. Intell. Rev.</source> <volume>44</volume>, <fpage>467</fpage>&#x02013;<lpage>508</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-015-9433-y</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schonlau</surname> <given-names>M.</given-names></name> <name><surname>Zou</surname> <given-names>R. Y.</given-names></name></person-group> (<year>2020</year>). <article-title>The random forest algorithm for statistical learning</article-title>. <source>Stata J</source>. <volume>20</volume>, <fpage>3</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1177/1536867X20909688</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Senvar</surname> <given-names>O.</given-names></name> <name><surname>Sennaroglu</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>Comparing performances of clements, box-cox, johnson methods with weibull distributions for assessing process capability</article-title>. <source>J. Ind. Eng. Manag</source>. <volume>9</volume>, <fpage>634</fpage>&#x02013;<lpage>656</lpage>. <pub-id pub-id-type="doi">10.3926/jiem.1703</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sevani</surname> <given-names>N.</given-names></name> <name><surname>Persulessy</surname> <given-names>G.</given-names></name> <name><surname>Persulessy</surname> <given-names>G. B. V.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Detection anemia based on conjunctiva pallor level using <italic>k</italic>-means algorithm,&#x0201D;</article-title> in <source>IOP Conference Series: Materials Science and Engineering</source>. <publisher-loc>Bristol</publisher-loc>: <publisher-name>IOP Publishing</publisher-name>.</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>D.</given-names></name> <name><surname>Singh</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Investigating the impact of data normalization on classification performance</article-title>. <source>Appl. Soft Comput</source>. 97, 105524. <pub-id pub-id-type="doi">10.1016/j.asoc.2019.105524</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Suner</surname> <given-names>S.</given-names></name> <name><surname>Rayner</surname> <given-names>J.</given-names></name> <name><surname>Ozturan</surname> <given-names>I. U.</given-names></name> <name><surname>Hogan</surname> <given-names>G.</given-names></name> <name><surname>Meehan</surname> <given-names>C. P.</given-names></name> <name><surname>Chambers</surname> <given-names>A. B.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Prediction of anemia and estimation of hemoglobin concentration using a smartphone camera</article-title>. <source>PLoS ONE</source> <volume>16</volume>, <fpage>e0253495</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0253495</pub-id><pub-id pub-id-type="pmid">34260592</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tamir</surname> <given-names>A.</given-names></name> <name><surname>Jahan</surname> <given-names>C. S.</given-names></name> <name><surname>Saif</surname> <given-names>M. S.</given-names></name> <name><surname>Zaman</surname> <given-names>S. U.</given-names></name> <name><surname>Islam</surname> <given-names>M. M.</given-names></name> <name><surname>Khan</surname> <given-names>A. I.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Detection of anemia from image of the anterior conjunctiva of the eye by image processing and thresholding,&#x0201D;</article-title> in <source>2017 IEEE Region 10 Humanitarian Technology Conference (R10-HTC)</source>. <publisher-loc>Dhaka</publisher-loc>: <publisher-name>IEEE</publisher-name>, <fpage>697</fpage>&#x02013;<lpage>701</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tharwat</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Classification assessment methods</article-title>. <source>Appl. Comp. Informa</source>. <volume>17</volume>, <fpage>168</fpage>&#x02013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1016/j.aci.2018.08.003</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trentin</surname> <given-names>E.</given-names></name></person-group> (<year>2015</year>). <article-title>Maximum-likelihood normalization of features increases the robustness of neural-based spoken human-computer interaction</article-title>. <source>Pattern Recognit. Lett</source>. <volume>66</volume>, <fpage>71</fpage>&#x02013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2015.07.003</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>E. J.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Hawkins</surname> <given-names>D.</given-names></name> <name><surname>Gernsheimer</surname> <given-names>T.</given-names></name> <name><surname>Norby-Slycord</surname> <given-names>C.</given-names></name> <name><surname>Patel</surname> <given-names>S. N.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Hemaapp: noninvasive blood screening of hemoglobin using smartphone cameras,&#x0201D;</article-title> in <source>Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing</source>, <fpage>593</fpage>&#x02013;<lpage>604</lpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><collab>World Health Organization</collab></person-group>. (<year>2011</year>). <source>Haemoglobin Concentrations for the Diagnosis of Anaemia and Assessment of Severity</source>. <publisher-loc>Geneva</publisher-loc>: <publisher-name>World Health Organization</publisher-name>.</citation>
</ref>
</ref-list>
</back>
</article>