<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Nutr.</journal-id>
<journal-title>Frontiers in Nutrition</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Nutr.</abbrev-journal-title>
<issn pub-type="epub">2296-861X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnut.2021.669155</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Nutrition</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Predicting Obesity in Adults Using Machine Learning Techniques: An Analysis of Indonesian Basic Health Research 2018</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Thamrin</surname> <given-names>Sri Astuti</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1224520/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Arsyad</surname> <given-names>Dian Sidik</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1237624/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kuswanto</surname> <given-names>Hedi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1342297/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Lawi</surname> <given-names>Armin</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1342287/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Nasir</surname> <given-names>Sudirman</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1342391/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Statistics, Faculty of Mathematics and Natural Science, Hasanuddin University</institution>, <addr-line>Makassar</addr-line>, <country>Indonesia</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Epidemiology, Faculty of Public Health, Hasanuddin University</institution>, <addr-line>Makassar</addr-line>, <country>Indonesia</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Mathematics, Faculty of Mathematics and Natural Sciences, Hasanuddin University</institution>, <addr-line>Makassar</addr-line>, <country>Indonesia</country></aff>
<aff id="aff4"><sup>4</sup><institution>Department of Health Promotion, Faculty of Public Health, Hasanuddin University</institution>, <addr-line>Makassar</addr-line>, <country>Indonesia</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Francesco Sofi, Universit&#x000E0; degli Studi di Firenze, Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Jos&#x000E9; Mar&#x000ED;a Huerta, Instituto de Salud Carlos III (ISCIII), Spain; Rosa Casas Rodriguez, Institut de Recerca Biom&#x000E8;dica August Pi i Sunyer (IDIBAPS), Spain</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Sri Astuti Thamrin <email>tuti&#x00040;unhas.ac.id</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Nutritional Epidemiology, a section of the journal Frontiers in Nutrition</p></fn>
<fn fn-type="other" id="fn002"><p>&#x02020;These authors have contributed equally to this work and share first authorship</p></fn></author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>06</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>669155</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>02</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>04</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Thamrin, Arsyad, Kuswanto, Lawi and Nasir.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Thamrin, Arsyad, Kuswanto, Lawi and Nasir</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Obesity is strongly associated with multiple risk factors. It is significantly contributing to an increased risk of chronic disease morbidity and mortality worldwide. There are various challenges to better understand the association between risk factors and the occurrence of obesity. The traditional regression approach limits analysis to a small number of predictors and imposes assumptions of independence and linearity. Machine Learning (ML) methods are an alternative that provide information with a unique approach to the application stage of data analysis on obesity. This study aims to assess the ability of ML methods, namely Logistic Regression, Classification and Regression Trees (CART), and Na&#x000EF;ve Bayes to identify the presence of obesity using publicly available health data, using a novel approach with sophisticated ML methods to predict obesity as an attempt to go beyond traditional prediction models, and to compare the performance of three different methods. Meanwhile, the main objective of this study is to establish a set of risk factors for obesity in adults among the available study variables. Furthermore, we address data imbalance using Synthetic Minority Oversampling Technique (SMOTE) to predict obesity status based on risk factors available in the dataset. This study indicates that the Logistic Regression method shows the highest performance. Nevertheless, kappa coefficients show only moderate concordance between predicted and measured obesity. Location, marital status, age groups, education, sweet drinks, fatty/oily foods, grilled foods, preserved foods, seasoning powders, soft/carbonated drinks, alcoholic drinks, mental emotional disorders, diagnosed hypertension, physical activity, smoking, and fruit and vegetables consumptions are significant in predicting obesity status in adults. Identifying these risk factors could inform health authorities in designing or modifying existing policies for better controlling chronic diseases especially in relation to risk factors associated with obesity. Moreover, applying ML methods on publicly available health data, such as Indonesian Basic Health Research (RISKESDAS) is a promising strategy to fill the gap for a more robust understanding of the associations of multiple risk factors in predicting health outcomes.</p></abstract>
<kwd-group>
<kwd>classification</kwd>
<kwd>Logistic Regression</kwd>
<kwd>machine learning</kwd>
<kwd>Naive Bayes</kwd>
<kwd>obesity status</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="4"/>
<equation-count count="1"/>
<ref-count count="43"/>
<page-count count="15"/>
<word-count count="8910"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Obesity is a major health problem strongly associated with many chronic illnesses with negative effects and long-term consequences, not only for the patients but also their families. In Southeast Asia, problems related to nutrition or malnutrition are a double burden because the number of cases of malnutrition and malnourishment is still relatively high and the number of cases of obesity has also increased significantly over time (<xref ref-type="bibr" rid="B1">1</xref>).</p>
<p>Data from the 2013 national-level survey of Indonesian Basic Health Research (RISKESDAS) showed the prevalence of obesity in Indonesia has increased over the years. Obesity among adult men was 13.9% in 2007, 7.8% in 2010, and 19.7% in 2013, whereas for adult women the prevalence was 14.8% in 2007, 15.5% in 2010, and increased drastically to 32.9% in 2013 (<xref ref-type="bibr" rid="B2">2</xref>). By 2018, the same survey (RISKESDAS 2018) showed that the prevalence of obesity in men and women had decreased slightly to 14.5 and 29.3%, respectively (<xref ref-type="bibr" rid="B3">3</xref>).</p>
<p>Risk factors for obesity have been studied extensively, and in general, they are divided into several categories: demographic and socio-economic factors (gender, age, education, income, marital status, and urban areas) (<xref ref-type="bibr" rid="B4">4</xref>&#x02013;<xref ref-type="bibr" rid="B6">6</xref>); lifestyle factors (consumption of fast food, stress, smoking, alcoholic drinks, and low level of physical activity) (<xref ref-type="bibr" rid="B6">6</xref>, <xref ref-type="bibr" rid="B7">7</xref>); and genetic factors (obese parents) (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B5">5</xref>). Among these risk factors, some can be changed or modified, while others cannot. Identifying modifiable risk factors for obesity at the individual and the population level is urgently required in order to implement an effective risk reduction strategy. Numerous studies have explored better approaches to predicting obesity using available data. A novel method recently introduced to answer this question uses Machine Learning (ML), which is currently one of the most popular topics in the scientific community for large-scale datasets.</p>
<p>Epidemiological data modeling using ML approaches is becoming increasingly popular in the published scientific literature. These methods have the potential to improve our understanding of general health regarding disease distribution, detection, and the identification of risk factors for health problems, and thus, opportunities for intervention. Various ML methods and algorithms have been applied to various aspects of health data including obesity (<xref ref-type="bibr" rid="B8">8</xref>). In the case of obesity, it is essential to develop a precise data classification to facilitate the process of finding predictive risk factors from the given data, in efforts to control these risk factors and eventually to decrease morbidity and mortality linked to obesity.</p>
<p>For the purpose of obesity prevention, ML has been used to predict the probability of obesity based on data encoding adherence to dietary recommendations and several other factors (<xref ref-type="bibr" rid="B9">9</xref>). The ML has also been applied for the prediction of obesity in children using electronic health records before the age of 2 (<xref ref-type="bibr" rid="B10">10</xref>); prediction of obesogenic environments for children (<xref ref-type="bibr" rid="B11">11</xref>); and for the aggregation of metabolomics, lipidomics, and other clinical data to modeling drug dose responses (<xref ref-type="bibr" rid="B12">12</xref>).</p>
<p>Based on previous research, ML approaches can increase the risk prediction of health outcomes compared to conventional approaches (<xref ref-type="bibr" rid="B13">13</xref>). Prediction of obesity using ML has been investigated by many researchers: Zhang et al. (<xref ref-type="bibr" rid="B14">14</xref>), Adnan et al. (<xref ref-type="bibr" rid="B15">15</xref>), Toschke et al. (<xref ref-type="bibr" rid="B16">16</xref>), Golino et al. (<xref ref-type="bibr" rid="B17">17</xref>), Dugan et al. (<xref ref-type="bibr" rid="B10">10</xref>), Zheng and Ruggiero (<xref ref-type="bibr" rid="B18">18</xref>), Chatterjee et al. (<xref ref-type="bibr" rid="B19">19</xref>), Singh and Tawfik (<xref ref-type="bibr" rid="B20">20</xref>), and Colmenarejo (<xref ref-type="bibr" rid="B21">21</xref>). The ML approach provides an alternative in providing information with a unique approach at the application stage of data analysis on obesity which is important in providing a better predictive solution to the likelihood of obesity (<xref ref-type="bibr" rid="B22">22</xref>).</p></sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and Methods</title>
<sec>
<title>Data Source</title>
<p>The dataset used to develop the classification model in this study is publicly available data from an Indonesia national scale survey with a cross-sectional and non-intervention design, the RISKESDAS survey, which was conducted by the Indonesian Ministry of Health. The RISKESDAS report is a community-based health survey whose indicators can be generalized with variables described from the national level down to the district/city level. It is conducted every 5 years across 34 provinces and 514 districts/cities in order to track important indicators of public health status, diseases risk factors, and to evaluate healthcare services delivery programs. The methodology and detailed protocols of the survey are described elsewhere (<xref ref-type="bibr" rid="B3">3</xref>). Briefly, the target sample for this study is 300,000 households from 30,000 Census Block (CBs) in 34 provinces and 514 district-cities throughout Indonesia. The sampling frame lists are provided by the Central Bureau of Statistics (BPS) using a two-stage sampling method. In the first stage, 180,000 CBs (25%) were selected from 720,000 CBs from the national socio-economic survey (SUSENAS) as a sampling frame using a proportionate to population size (PPS) method and stratified by prosperity level, continued by systematically selecting 30,000 CBs from 180,000 CBs priorly selected and stratified by urban and rural for each district or city. In the second stage, 10 households were selected systematically using implicit stratification for the education level of the head of household to maintain variation of education among households. Household members who were eligible according to the inclusion criteria were invited to participate in the interview.</p>
<p>The dataset can be accessed by request at the Institute of Health Research and Development of the Indonesian Ministry of Health (<ext-link ext-link-type="uri" xlink:href="https://www.litbang.kemkes.go.id/layanan-permintaan-data-riset/">https://www.litbang.kemkes.go.id/layanan-permintaan-data-riset/</ext-link>).</p></sec>
<sec>
<title>Pre-processing Data</title>
<sec>
<title>Data Cleaning or Filtering</title>
<p>The sample used in this study included all the data from the RISKESDAS dataset for individuals aged 18 or above; in total there was data for 634,709 respondents. We conducted data cleaning by excluding all records with incomplete or missing values for the variable/feature Body Mass Index (BMI), a core feature used to categorize obesity status. The number of samples included for the analysis process after cleaning was 618,898 records. Data cleaning was performed by using the <italic>dplyr</italic> package of R version 3.5.1 to perform filtering (<xref ref-type="bibr" rid="B23">23</xref>).</p></sec>
<sec>
<title>Feature Selection</title>
<p>After removing missing values, we proceeded to variable or feature selection. Variable selection is a process of reducing the data dimensions to reduce processing time as well as computation costs (<xref ref-type="bibr" rid="B24">24</xref>). We selected a subset of variables that contributed significantly to the target class to improve the overall predictive performance of the classification using the Chi-Square (&#x003C7;<sup>2</sup>) test between obesity status with each of the variables and including those with a <italic>p</italic>-value &#x0003C;0.05. All features that met these criteria (a total of 21 features) were selected for developing the classification model. These variables or features were location (X1), marital status (X2), age group (X3), education (X4), work category (X5), sugary foods (X6), sweet drinks (X7), salty foods (X8), fatty/oily foods (X9), grilled foods (X10), preserved foods (X11), seasoning powders (X12), soft/carbonated drinks (X13), energy drinks (X14), instant foods (X15), alcoholic drinks (X16), mental-emotional disorders (X17), diagnosed hypertension (X18), physical activity (X19), smoking (X20), and fruit and vegetables consumptions (X21). A list of these features and how it was generated from the questionnaire (for composited and calculated feature, i.e., obesity, fruit and vegetables consumption, physical activity, and mental-emotional disorders) can be found in the <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1</xref>. The process of developing a classification model was carried out by using the R Statistical Software version 3.5.1 (<xref ref-type="bibr" rid="B25">25</xref>).</p></sec></sec>
<sec>
<title>Dealing With Imbalanced Datasets</title>
<p>Data imbalance occurs when there are one or more classes that dominate the whole data as major classes, and other classes are rare occurrences or minor classes. Imbalanced data will produce a good classification prediction accuracy against the major class, but in the minor class, the resulting accuracy is poor.</p>
<p>The Synthetic Minority Oversampling Technique (SMOTE) was introduced by Chawla et al. (<xref ref-type="bibr" rid="B26">26</xref>) and Chawla (<xref ref-type="bibr" rid="B27">27</xref>), as a way of dealing with the effect of the lack of information on minority classes in a data set. SMOTE is an algorithm with an oversampling approach, which generates artificial data for minority data classes (<xref ref-type="bibr" rid="B28">28</xref>) so that the proportions of major and minor data classes are more balanced (<xref ref-type="bibr" rid="B29">29</xref>). Artificial data or synthetic data are made based on the <italic>k</italic>-nearest neighbor. All attributes used in this study were categorical features so that the calculation of the distance between the minor class samples was carried out using the Modify Value Difference Metric (MVDM) method (<xref ref-type="bibr" rid="B30">30</xref>). In this method, several steps are taken, namely calculating the distance between two observations at a nominal scale and choosing the majority category between the minority class observations with its <italic>k</italic>-closest neighbors for a nominal value, and if the same value occurs, it is chosen randomly. Furthermore, the selected value is a new observation. In this study, the SMOTE technique with oversampling of 200% and 300% was used which resulted in two new datasets.</p></sec>
<sec>
<title>Machine Learning Classification Methods</title>
<sec>
<title>Logistic Regression</title>
<p>One of the basic linear models developed with a probabilistic approach to classification problems is Logistic Regression (<xref ref-type="bibr" rid="B31">31</xref>) and is one of the supervised learning models widely used in ML. Logistic Regression can be seen as a development of Linear Regression models with a logistic function for data with a target in the form of classes (<xref ref-type="bibr" rid="B32">32</xref>) as follows:</p>
<disp-formula id="E1"><mml:math id="M1"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>x</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M2"><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the <italic>D-</italic>dimensional data, <inline-formula><mml:math id="M3"><mml:mi>&#x003B2;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are the weight parameters, &#x003B2;<sub>0</sub> is the bias parameter, and &#x003C3; is a logistic function that is shaped as <inline-formula><mml:math id="M4"><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p>
<p>The weights of &#x003B2; can be obtained by using probabilistic concepts. For example, if <italic>y</italic><sub><italic>n</italic></sub> &#x0003D; <italic>y</italic>(<italic>x</italic><sub><italic>n</italic></sub>) and <italic>t</italic><sub><italic>n</italic></sub> &#x02208; {0, 1} are an independent identical distribution. The joint probabilistic or likelihood function for all the data can be expressed by the Bernoulli distribution <italic>p</italic>(<italic>t</italic>|&#x003B2;), where <inline-formula><mml:math id="M5"><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Therefore, the Logistic Regression learning and bias (&#x003B2;) is to maximize <italic>p</italic>(<italic>t</italic>&#x02228;&#x003B2;). The learning method for determining the weight and bias (&#x003B2;) parameters is known as the maximum likelihood method. Generally, the solution to the maximum likelihood problem is done by minimizing the negative of the logarithm of the likelihood function, namely <inline-formula><mml:math id="M6"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow></mml:munder><mml:mtext>&#x000A0;</mml:mtext><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>E</italic>(&#x003B2;) &#x0003D; &#x02212;ln(<italic>p</italic>(<italic>t</italic>&#x02228;&#x003B2;)). Logistic Regression models can use regularization techniques to solve the problem of overfitting by adding the weight norm ||w|| in the error function, namely <inline-formula><mml:math id="M7"><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mi>C</mml:mi><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, where C &#x0003E; 0 is the inverse parameter of the regulation.</p>
<p>Simultaneous and partial parameter testing is performed to examine the role of predictor variables in the model. Simultaneous parameter testing uses the G test.</p></sec>
<sec>
<title>Classification and Regression Trees</title>
<p>Breiman et al. (<xref ref-type="bibr" rid="B33">33</xref>) proposes a new algorithm for tree arrangement, namely Classification and Regression Tree (CART). CART is a non-parametric statistical method used for classification analysis, both for categorical and continuous response variables, and for explanatory variables which may consist of nominal, ordinal, or continuous features. The resulting tree model depends on the scale of the response attribute. CART generates a classification tree if the response variables are categorical, and generates a regression tree if the response variables are continuous (<xref ref-type="bibr" rid="B33">33</xref>).</p>
<p>The tree structure in the CART method is obtained through a binary recursive partitioning algorithm against its explanatory variables (<xref ref-type="bibr" rid="B31">31</xref>, <xref ref-type="bibr" rid="B32">32</xref>). The binding is carried out by dividing the data set into two subclusters called nodes. The impurity value at node <italic>t</italic> is a measurement of the heterogeneity level of a class from a particular node in the classification tree. The process of forming a classification tree is carried out in three stages; selecting a classifier, determining the final node, and marking the class label (<xref ref-type="bibr" rid="B31">31</xref>). In selecting the classifier, each partitioning depends on the value that comes from only one explanatory variable. For categorical variables, the partitioning that occurs comes from all the possible partitioning based on the formation of two subgroups that are mutually exclusive (disjoint). In addition, in solving classification tree problems, the Gini Splitting Rule (also known as the Gini Index) is the most common rule to be used (<xref ref-type="bibr" rid="B32">32</xref>). Then, the partitioning evaluation is performed using the goodness of split &#x003C6;(<italic>s, t</italic>) of the <italic>s</italic> partition at <italic>t</italic> node. The partitioning function is defined as decreased heterogeneity. A sort that produces a higher value is a better sort because it reduces the impurity value more significantly. If the resulting node is of a non-homogeneous class, the same procedure will be repeated until the tree <inline-formula><mml:math id="M8"><mml:mi>&#x003C6;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>&#x003C6;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:munder><mml:mi>&#x003C6;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Determination of child nodes is carried out recursively by using the same method as determining the main node.</p>
<p>After selecting the classifier, the end node is determined. The minimum number of cases in a node is generally five. If this is fulfilled, tree development will be stopped and continued with the marking of class labels. Class label marking at the end node is carried out based on the highest number rule. The process of forming classification trees stops when there is only one observation in each child node. One of the ways to get the optimal tree is by consecutively pruning the tree that is less important. In random pruning, the observations are divided into two parts, namely training data <italic>L</italic><sub>1</sub> and test data <italic>L</italic><sub>2</sub>. Through the pruning process, a row of trees is formed from <italic>L</italic><sub>1</sub>. Next, <italic>L</italic><sub>2</sub> is used to form the total proportion of misclassification (<italic>R</italic>|<italic>ts</italic>(<italic>G</italic>)). The optimal tree that meets the criteria as <inline-formula><mml:math id="M9"><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo class="qopname">min</mml:mo><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></sec>
<sec>
<title>Na&#x000EF;ve Bayesian</title>
<p>Na&#x000EF;ve Bayesian classification is a statistical approach which attempts to predict the probability of each class (<xref ref-type="bibr" rid="B14">14</xref>). The advantage of this Bayes grouping is that it has a high level of accuracy and speed when using large data sets. Na&#x000EF;ve Bayesian grouping assumes that the values of the variables on the class labels are independent of other attribute values, which can facilitate the calculation (<xref ref-type="bibr" rid="B10">10</xref>, <xref ref-type="bibr" rid="B34">34</xref>).</p>
<p>Na&#x000EF;ve Bayesian Classification is achieved by applying the Bayes rule to calculate the probability of each attribute and predicting the class based on the highest prior probability (<xref ref-type="bibr" rid="B34">34</xref>).</p></sec></sec>
<sec>
<title>Model Validation</title>
<p>The validation process in this study used <italic>k</italic>-fold cross-validation (<xref ref-type="bibr" rid="B35">35</xref>). Cross-Validation (CV) divides the dataset into two parts: one part is used as the training data and the other is used as testing data. In this study, the data were divided into 10 parts, 90% of which was used as training and the rest was used for testing. This process was done repeatedly, a maximum of 10 times, until all data records were part of the testing data. This process is also known as the 10-fold CV. The 10-fold CV process has been used in several previous health care- and medical-related studies (<xref ref-type="bibr" rid="B36">36</xref>).</p></sec>
<sec>
<title>Evaluation of Classification Performance</title>
<p>Measuring accuracy is a diagnostic step to test the level of performance of an algorithm against the dataset used. A matrix, known as the confusion matrix, is used to evaluate the learning algorithm (<xref ref-type="bibr" rid="B37">37</xref>). Each column in the matrix shows the number of observations in the predicted class. The rows in the matrix represent the actual number of observations in the class.</p>
<p>In ML, the term metric refers to a value that can be used to represent the performance of the resulting model. In classification modeling, the model output is a label/class. There are several metrics that are commonly used, namely accuracy, precision, sensitivity, specificity, recall, F1-score, kappa, and <italic>F</italic><sub>&#x003B2;</sub>. In terms of the confusion matrix, accuracy is the ratio of the number of diagonal elements to the total number of matrix elements. The accuracy of the method is only considered adequate when the comparison of the actual number of data labels is nearly identical with the confusion matrix. If the comparison is imbalanced, then other metrics can be used. Precision is an appropriate metric when false positives are to be avoided. Sensitivity can be interpreted as the degree of reliability of the model to detect data labeled positive correctly. Sensitivity is an appropriate metric when false negatives are to be avoided (high risk). Specificity is the degree of model reliability for detecting data labeled negative correctly. This metric is closely related to sensitivity. This metric is appropriate when the true negative rate is to be maximized. To minimize both (false positive and false negative) outcomes at the same time, precision and sensitivity need to be summarized by using the F1-score. Recall is a valid choice of evaluation metric when we want to capture as many positives (obese) as possible. In this study, we want to be sure that the sample we catch is obese (precision) and we also want to capture as many obese (recall) as possible. The F1-score manages this trade-off. However, the main problem with the F1-score is that it gives equal weight to precision and recall. Sometimes we may need to include domain knowledge in our evaluations where we want more recall or more precision. To solve this, we can create a weighted F1 metric, where beta (&#x003B2;) sets the balance between precision and recall. This is called <italic>F</italic><sub>&#x003B2;</sub>. In this study, we used &#x003B2; = 0.5 to measure more weight on precision and less weight on recall.</p>
<p>Kappa is used to test the inter reliability. Kappa values range from 0 to 1.0 which can be divided into several classifications, namely 0&#x02013;0.20 (slight), 0.21&#x02013;0.40 (fair), 0.41&#x02013;0.60 (moderate), 0.61&#x02013;0.80 (substantial), and 0.81&#x02013;1.0 (perfect) (<xref ref-type="bibr" rid="B38">38</xref>).</p>
<p>The Area Under ROC Curve, also known as AUC, has a range between 0.5 (50%) and 1 (100%). The interpretation of AUC values can be classified into five different sections, namely 0.5&#x02013;0.6 (false accuracy), 0.6&#x02013;0.7 (poor accuracy), 0.7&#x02013;0.8 (moderate accuracy), 0.8&#x02013;0.9 (high accuracy), and 0.9&#x02013;1 (very high level of accuracy) (<xref ref-type="bibr" rid="B39">39</xref>).</p></sec></sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>An overview of the explanatory variables contained in the obesity data of the Indonesia RISKESDAS 2018 survey is given in <xref ref-type="table" rid="T1">Table 1</xref>. As can be seen from <xref ref-type="table" rid="T1">Table 1</xref>, out of 618,898 respondents, there are 134,709 (21.77%) people who are classified as obese, 484,189 (78.23%) people are non-obese. In <xref ref-type="table" rid="T1">Table 1</xref>, it can also be seen that the number of obese (21.77%) and non-obese classes (78.23%) seems imbalanced. Based on <xref ref-type="table" rid="T1">Table 1</xref>, the respondents in this study lived in rural areas (56.71%), married (76.31%), aged 35&#x02013;39 years (12.53%), finished senior high school (25.43%), unemployed (27.79%), consumed sugary foods 1&#x02013;2 times per week (28.63%), drank sweet drinks one time per day (31.57%), consumed salty foods 1&#x02013;2 times per week (27.54%), consumed fatty/oily foods 1&#x02013;2 times per week (26.61%), consumed grilled foods more than 3 times per month (32.68%), never consumed preserved foods (56.70%), consumed seasoning powders less that one time per day (36.74%), never drank soft/carbonated drinks (72.19%), never drank energy drinks (81.58%), experienced no mental emotional disorders (90.13%), consumed instant foods 1&#x02013;2 times per week (35.57%), drank non-alcoholic drinks (95.11%), diagnosed with no hypertension (50.97%), not adequate physical activity (88.09%), not a smoker (62.30%), and consumed inadequate fruit and vegetables (95.26%). This general description of the obesity data can be seen in detail in <xref ref-type="table" rid="T1">Table 1</xref>. Moreover, the obesity status description can be seen in detail in the <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 2</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>General description of obesity data from Indonesian RISKESDAS 2018.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Variables</bold></th>
<th valign="top" align="left"><bold>Categories</bold></th>
<th valign="top" align="center"><bold>Frequency</bold></th>
<th valign="top" align="center"><bold>Percentage</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Obesity status (Y)</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="center">484,189</td>
<td valign="top" align="center">78.23</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="center">134,709</td>
<td valign="top" align="center">21.77</td>
</tr>
<tr>
<td valign="top" align="left">Location (X1)</td>
<td valign="top" align="left">Urban</td>
<td valign="top" align="center">267,913</td>
<td valign="top" align="center">43.29</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Rural</td>
<td valign="top" align="center">350,985</td>
<td valign="top" align="center">56.71</td>
</tr>
<tr>
<td valign="top" align="left">Marital status (X2)</td>
<td valign="top" align="left">Not married</td>
<td valign="top" align="center">84,792</td>
<td valign="top" align="center">13.70</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Married</td>
<td valign="top" align="center">472,269</td>
<td valign="top" align="center">76.31</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Divorced</td>
<td valign="top" align="center">14,333</td>
<td valign="top" align="center">2.32</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Widowed</td>
<td valign="top" align="center">47,504</td>
<td valign="top" align="center">7.68</td>
</tr>
<tr>
<td valign="top" align="left">Age groups (X3)</td>
<td valign="top" align="left">18&#x02013;24 years</td>
<td valign="top" align="center">69,532</td>
<td valign="top" align="center">11.23</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">25&#x02013;29 years</td>
<td valign="top" align="center">60,380</td>
<td valign="top" align="center">9.76</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">30&#x02013;34 years</td>
<td valign="top" align="center">68,683</td>
<td valign="top" align="center">11.10</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">35&#x02013;39 years</td>
<td valign="top" align="center">77,538</td>
<td valign="top" align="center">12.53</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">40&#x02013;44 years</td>
<td valign="top" align="center">73,775</td>
<td valign="top" align="center">11.92</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">45&#x02013;49 years</td>
<td valign="top" align="center">70,503</td>
<td valign="top" align="center">11.39</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">50&#x02013;54 years</td>
<td valign="top" align="center">58,618</td>
<td valign="top" align="center">9.47</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">55&#x02013;59 years</td>
<td valign="top" align="center">49,632</td>
<td valign="top" align="center">8.02</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">60&#x02013;64 years</td>
<td valign="top" align="center">35,471</td>
<td valign="top" align="center">5.73</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003E;64 years</td>
<td valign="top" align="center">54,766</td>
<td valign="top" align="center">8.85</td>
</tr>
<tr>
<td valign="top" align="left">Education (X4)</td>
<td valign="top" align="left">Not/Never schooled</td>
<td valign="top" align="center">40,861</td>
<td valign="top" align="center">6.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Not finished basic school</td>
<td valign="top" align="center">84,637</td>
<td valign="top" align="center">13.68</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished basic school</td>
<td valign="top" align="center">157,391</td>
<td valign="top" align="center">25.43</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished Junior High School</td>
<td valign="top" align="center">104,435</td>
<td valign="top" align="center">16.87</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished Senior High School</td>
<td valign="top" align="center">170,246</td>
<td valign="top" align="center">27.51</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished Academy/College</td>
<td valign="top" align="center">20,005</td>
<td valign="top" align="center">3.23</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished higher education</td>
<td valign="top" align="center">41,323</td>
<td valign="top" align="center">6.68</td>
</tr>
<tr>
<td valign="top" align="left">Work types (X5)</td>
<td valign="top" align="left">Not working</td>
<td valign="top" align="center">171,984</td>
<td valign="top" align="center">27.79</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">School</td>
<td valign="top" align="center">12,238</td>
<td valign="top" align="center">1.98</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Government employee</td>
<td valign="top" align="center">27,703</td>
<td valign="top" align="center">4.48</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Private employee</td>
<td valign="top" align="center">50,049</td>
<td valign="top" align="center">8.09</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Entrepreneur</td>
<td valign="top" align="center">91,011</td>
<td valign="top" align="center">14.71</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Farmer</td>
<td valign="top" align="center">163,009</td>
<td valign="top" align="center">26.34</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Fisherman</td>
<td valign="top" align="center">8,344</td>
<td valign="top" align="center">1.35</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Daily waged labors</td>
<td valign="top" align="center">52,379</td>
<td valign="top" align="center">8.46</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Others</td>
<td valign="top" align="center">42,181</td>
<td valign="top" align="center">6.82</td>
</tr>
<tr>
<td valign="top" align="left">Sugary foods (X6)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">82,775</td>
<td valign="top" align="center">13.37</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">125,754</td>
<td valign="top" align="center">20.32</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">138,685</td>
<td valign="top" align="center">22.41</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">177,173</td>
<td valign="top" align="center">28.63</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">62,972</td>
<td valign="top" align="center">10.17</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">31,539</td>
<td valign="top" align="center">5.10</td>
</tr>
<tr>
<td valign="top" align="left">Sweet drinks (X7)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">176,096</td>
<td valign="top" align="center">28.45</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">195,361</td>
<td valign="top" align="center">31.57</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">87,827</td>
<td valign="top" align="center">14.19</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">95,409</td>
<td valign="top" align="center">15.42</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">33,666</td>
<td valign="top" align="center">5.44</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">30,539</td>
<td valign="top" align="center">4.93</td>
</tr>
<tr>
<td valign="top" align="left">Salty foods (X8)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">64,660</td>
<td valign="top" align="center">10.45</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">78,744</td>
<td valign="top" align="center">12.72</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">105,363</td>
<td valign="top" align="center">17.02</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">170,442</td>
<td valign="top" align="center">27.54</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">107,318</td>
<td valign="top" align="center">17.34</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">92,371</td>
<td valign="top" align="center">14.93</td>
</tr>
<tr>
<td valign="top" align="left">Fatty/Oily foods (X9)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">103,634</td>
<td valign="top" align="center">16.74</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">113,057</td>
<td valign="top" align="center">18.27</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">133,552</td>
<td valign="top" align="center">21.58</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">164,703</td>
<td valign="top" align="center">26.61</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">72,739</td>
<td valign="top" align="center">11.75</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">31,213</td>
<td valign="top" align="center">5.04</td>
</tr>
<tr>
<td valign="top" align="left">Grilled foods (X10)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">12,948</td>
<td valign="top" align="center">2.09</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">22,189</td>
<td valign="top" align="center">3.59</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">63,967</td>
<td valign="top" align="center">10.34</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">161,356</td>
<td valign="top" align="center">26.07</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">202,251</td>
<td valign="top" align="center">32.68</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">156,187</td>
<td valign="top" align="center">25.24</td>
</tr>
<tr>
<td valign="top" align="left">Preserved foods (X11)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">6,310</td>
<td valign="top" align="center">1.02</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">12,024</td>
<td valign="top" align="center">1.94</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">31,993</td>
<td valign="top" align="center">5.17</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">72,618</td>
<td valign="top" align="center">11.73</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">145,068</td>
<td valign="top" align="center">23.44</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">350,885</td>
<td valign="top" align="center">56.70</td>
</tr>
<tr>
<td valign="top" align="left">Seasonings powders (X12)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">227,357</td>
<td valign="top" align="center">36.74</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">226,628</td>
<td valign="top" align="center">36.62</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">42,598</td>
<td valign="top" align="center">6.88</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">34,030</td>
<td valign="top" align="center">5.50</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">20,887</td>
<td valign="top" align="center">3.37</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">67,398</td>
<td valign="top" align="center">10.89</td>
</tr>
<tr>
<td valign="top" align="left">Soft/Carbonated drinks (X13)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">3,689</td>
<td valign="top" align="center">0.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">7,857</td>
<td valign="top" align="center">1.27</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">16,470</td>
<td valign="top" align="center">2.66</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">43,686</td>
<td valign="top" align="center">7.06</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">100,398</td>
<td valign="top" align="center">16.22</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">446,798</td>
<td valign="top" align="center">72.19</td>
</tr>
<tr>
<td valign="top" align="left">Energy drinks (X14)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">3,654</td>
<td valign="top" align="center">0.59</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">7,761</td>
<td valign="top" align="center">1.25</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">12,888</td>
<td valign="top" align="center">2.08</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">31,045</td>
<td valign="top" align="center">5.02</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">58,659</td>
<td valign="top" align="center">9.48</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">504,891</td>
<td valign="top" align="center">81.58</td>
</tr>
<tr>
<td valign="top" align="left">Instant foods (X15)</td>
<td valign="top" align="left">&#x0003E;1 time per day</td>
<td valign="top" align="center">12,144</td>
<td valign="top" align="center">1.96</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1 time per day</td>
<td valign="top" align="center">28,943</td>
<td valign="top" align="center">4.68</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">108,287</td>
<td valign="top" align="center">17.50</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">220,125</td>
<td valign="top" align="center">35.57</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">149,066</td>
<td valign="top" align="center">24.09</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">100,333</td>
<td valign="top" align="center">16.21</td>
</tr>
<tr>
<td valign="top" align="left">Alcoholic drinks (X16)</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">30,240</td>
<td valign="top" align="center">4.89</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">No</td>
<td valign="top" align="center">588,658</td>
<td valign="top" align="center">95.11</td>
</tr>
<tr>
<td valign="top" align="left">Mental-emotional disorders (X17)</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">61,092</td>
<td valign="top" align="center">9.87</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">No</td>
<td valign="top" align="center">557,806</td>
<td valign="top" align="center">90.13</td>
</tr>
<tr>
<td valign="top" align="left">Diagnosed hypertension (X18)</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">55,640</td>
<td valign="top" align="center">8.99</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">No</td>
<td valign="top" align="center">315,467</td>
<td valign="top" align="center">50.97</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Unknown</td>
<td valign="top" align="center">247,791</td>
<td valign="top" align="center">40.04</td>
</tr>
<tr>
<td valign="top" align="left">Physical activity (X19)</td>
<td valign="top" align="left">Adequate</td>
<td valign="top" align="center">73,736</td>
<td valign="top" align="center">11.91</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Not adequate</td>
<td valign="top" align="center">545,162</td>
<td valign="top" align="center">88.09</td>
</tr>
<tr>
<td valign="top" align="left">Smoking (X20)</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">233,306</td>
<td valign="top" align="center">37.70</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">No</td>
<td valign="top" align="center">385,592</td>
<td valign="top" align="center">62.30</td>
</tr>
<tr>
<td valign="top" align="left">Fruit and vegetables consumptions (X21)</td>
<td valign="top" align="left">Adequate</td>
<td valign="top" align="center">29,321</td>
<td valign="top" align="center">4.74</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Not adequate</td>
<td valign="top" align="center">589,577</td>
<td valign="top" align="center">95.26</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To overcome the oversampling of the prediction of this obesity status classification due to class imbalance in the dataset (<xref ref-type="table" rid="T1">Table 1</xref>), the SMOTE technique was used. In this study, the SMOTE technique used two different percentages, namely 200% and 300%. SMOTE with 300% can improve minor class data better (from 21.77%, in the original dataset, to 47.3%). As a result, the comparison between major class (non-obese) and minor class (obese) is balanced, namely 47.3% and 52.7%, respectively. The new dataset resulting from the SMOTE technique with 300% was used to build a classification model and prediction of obesity risk factors.</p>
<p>Using the three models (Logistic Regression model, CART, and Na&#x000EF;ve Bayes), 10-fold CV was carried out to train and see which model performed better in predicting test set points on all data (<xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref>). This is also to ensure that all these new data resulting from the SMOTE technique are not bias in the result.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Comparison of classification accuracy with 10-fold CV based on the obesity test data using three models with confusion matrix.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>ML methods</bold></th>
<th valign="top" align="left"><bold>Classification prediction</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 1 Test</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 2 Test</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 3 Test</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 4 Test</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 5 Test</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="center" colspan="10" style="border-bottom: thin solid #000000;"><bold>Real circumstances</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="left"><bold>Non-obese</bold></th>
<th valign="top" align="left"><bold>Obese</bold></th>
<th valign="top" align="left"><bold>Non-obese</bold></th>
<th valign="top" align="left"><bold>Obese</bold></th>
<th valign="top" align="left"><bold>Non-obese</bold></th>
<th valign="top" align="left"><bold>Obese</bold></th>
<th valign="top" align="left"><bold>Non-obese</bold></th>
<th valign="top" align="left"><bold>Obese</bold></th>
<th valign="top" align="left"><bold>Non-obese</bold></th>
<th valign="top" align="left"><bold>Obese</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CART</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="left">360,554</td>
<td valign="top" align="left">193,472</td>
<td valign="top" align="left">360,260</td>
<td valign="top" align="left">193,579</td>
<td valign="top" align="left">360,791</td>
<td valign="top" align="left">193,595</td>
<td valign="top" align="left">360,325</td>
<td valign="top" align="left">193,504</td>
<td valign="top" align="left">360,459</td>
<td valign="top" align="left">193,685</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="left">75,298</td>
<td valign="top" align="left">291,411</td>
<td valign="top" align="left">75,283</td>
<td valign="top" align="left">291,744</td>
<td valign="top" align="left">75,227</td>
<td valign="top" align="left">291,362</td>
<td valign="top" align="left">75,294</td>
<td valign="top" align="left">291,335</td>
<td valign="top" align="left">75,401</td>
<td valign="top" align="left">291,611</td>
</tr>
<tr>
<td valign="top" align="left">Na&#x000EF;ve-Bayes</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="left">314,384</td>
<td valign="top" align="left">141,264</td>
<td valign="top" align="left">313,957</td>
<td valign="top" align="left">141,209</td>
<td valign="top" align="left">314,357</td>
<td valign="top" align="left">141,167</td>
<td valign="top" align="left">314,080</td>
<td valign="top" align="left">141,106</td>
<td valign="top" align="left">314,273</td>
<td valign="top" align="left">141,413</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="left">121,468</td>
<td valign="top" align="left">343,619</td>
<td valign="top" align="left">121,586</td>
<td valign="top" align="left">344,114</td>
<td valign="top" align="left">121,661</td>
<td valign="top" align="left">343,790</td>
<td valign="top" align="left">121,539</td>
<td valign="top" align="left">343,733</td>
<td valign="top" align="left">121,587</td>
<td valign="top" align="left">343,883</td>
</tr>
<tr>
<td valign="top" align="left">Logistic Regression</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="left">320,456</td>
<td valign="top" align="left">140,260</td>
<td valign="top" align="left">319,952</td>
<td valign="top" align="left">140,279</td>
<td valign="top" align="left">320,628</td>
<td valign="top" align="left">140,336</td>
<td valign="top" align="left">320,202</td>
<td valign="top" align="left">140,144</td>
<td valign="top" align="left">320,285</td>
<td valign="top" align="left">140,474</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="left">115,396</td>
<td valign="top" align="left">344,623</td>
<td valign="top" align="left">115,591</td>
<td valign="top" align="left">345,044</td>
<td valign="top" align="left">115,390</td>
<td valign="top" align="left">344,621</td>
<td valign="top" align="left">115,417</td>
<td valign="top" align="left">344,695</td>
<td valign="top" align="left">115,575</td>
<td valign="top" align="left">344,822</td>
</tr> <tr>
<td valign="top" align="left"><bold>ML methods</bold></td>
<td valign="top" align="left"><bold>Classification prediction</bold></td>
<td valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 6 Test</bold></td>
<td valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 7 Test</bold></td>
<td valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 8 Test</bold></td>
<td valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 9 Test</bold></td>
<td valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Fold 10 Test</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center" colspan="10" style="border-bottom: thin solid #000000;"><bold>Real circumstances</bold></td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td/>
<td valign="top" align="left"><bold>Non-obese</bold></td>
<td valign="top" align="left"><bold>Obese</bold></td>
<td valign="top" align="left"><bold>Non-obese</bold></td>
<td valign="top" align="left"><bold>Obese</bold></td>
<td valign="top" align="left"><bold>Non-obese</bold></td>
<td valign="top" align="left"><bold>Obese</bold></td>
<td valign="top" align="left"><bold>Non-obese</bold></td>
<td valign="top" align="left"><bold>Obese</bold></td>
<td valign="top" align="left"><bold>Non-obese</bold></td>
<td valign="top" align="left"><bold>Obese</bold></td>
</tr> <tr>
<td valign="top" align="left">CART</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="left">360,531</td>
<td valign="top" align="left">193,271</td>
<td valign="top" align="left">360,426</td>
<td valign="top" align="left">193,360</td>
<td valign="top" align="left">360,177</td>
<td valign="top" align="left">193,275</td>
<td valign="top" align="left">360,566</td>
<td valign="top" align="left">193,586</td>
<td valign="top" align="left">360,411</td>
<td valign="top" align="left">193,430</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="left">75,312</td>
<td valign="top" align="left">291,645</td>
<td valign="top" align="left">75,410</td>
<td valign="top" align="left">291,447</td>
<td valign="top" align="left">75,351</td>
<td valign="top" align="left">291,331</td>
<td valign="top" align="left">75,308</td>
<td valign="top" align="left">291,504</td>
<td valign="top" align="left">75,317</td>
<td valign="top" align="left">291,377</td>
</tr>
<tr>
<td valign="top" align="left">Na&#x000EF;ve-Bayes</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="left">314,356</td>
<td valign="top" align="left">141,221</td>
<td valign="top" align="left">314,273</td>
<td valign="top" align="left">141,183</td>
<td valign="top" align="left">314,030</td>
<td valign="top" align="left">141,113</td>
<td valign="top" align="left">314,239</td>
<td valign="top" align="left">141,296</td>
<td valign="top" align="left">314,234</td>
<td valign="top" align="left">141,345</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="left">121,487</td>
<td valign="top" align="left">343,695</td>
<td valign="top" align="left">121,563</td>
<td valign="top" align="left">343,624</td>
<td valign="top" align="left">121,498</td>
<td valign="top" align="left">343,493</td>
<td valign="top" align="left">121,635</td>
<td valign="top" align="left">343,794</td>
<td valign="top" align="left">121,494</td>
<td valign="top" align="left">343,462</td>
</tr>
<tr>
<td valign="top" align="left">Logistic Regression</td>
<td valign="top" align="left">Non-obese</td>
<td valign="top" align="left">320,479</td>
<td valign="top" align="left">140,281</td>
<td valign="top" align="left">320,423</td>
<td valign="top" align="left">140,220</td>
<td valign="top" align="left">320,206</td>
<td valign="top" align="left">140,253</td>
<td valign="top" align="left">320,464</td>
<td valign="top" align="left">140,277</td>
<td valign="top" align="left">320,355</td>
<td valign="top" align="left">140,328</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Obese</td>
<td valign="top" align="left">115,364</td>
<td valign="top" align="left">344,635</td>
<td valign="top" align="left">115,413</td>
<td valign="top" align="left">344,587</td>
<td valign="top" align="left">115,322</td>
<td valign="top" align="left">344,353</td>
<td valign="top" align="left">115,410</td>
<td valign="top" align="left">344,813</td>
<td valign="top" align="left">115,373</td>
<td valign="top" align="left">344,479</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Evaluation of classification prediction performance with 10-fold CV based on the obesity test data using 3 ML methods.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>ML methods</bold></th>
<th valign="top" align="left"><bold>Test</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
<th valign="top" align="center"><bold>Sensitivity (%)</bold></th>
<th valign="top" align="center"><bold>Specificity (%)</bold></th>
<th valign="top" align="center"><bold>Precision (%)</bold></th>
<th valign="top" align="center"><bold>F1-Score (%)</bold></th>
<th valign="top" align="center"><bold>Kappa (%)</bold></th>
<th valign="top" align="center"><bold>AUC (%)</bold></th>
<th valign="top" align="center"><bold><italic>F<sub><bold>&#x003B2;</bold></sub></italic> <sub><bold> &#x0003D; 0.5</bold></sub> (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CART</td>
<td valign="top" align="left">1-Fold</td>
<td valign="top" align="center">70.81</td>
<td valign="top" align="center"><bold>82.72</bold></td>
<td valign="top" align="center">60.10</td>
<td valign="top" align="center">65.08</td>
<td valign="top" align="center"><bold>72.85</bold></td>
<td valign="top" align="center">42.24</td>
<td valign="top" align="center">74.57</td>
<td valign="top" align="center">67.98</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">2-Fold</td>
<td valign="top" align="center">70.80</td>
<td valign="top" align="center"><bold>82.72</bold></td>
<td valign="top" align="center">60.11</td>
<td valign="top" align="center">65.05</td>
<td valign="top" align="center"><bold>72.83</bold></td>
<td valign="top" align="center">42.24</td>
<td valign="top" align="center">74.56</td>
<td valign="top" align="center">67.95</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3-Fold</td>
<td valign="top" align="center">70.81</td>
<td valign="top" align="center"><bold>82.75</bold></td>
<td valign="top" align="center">60.08</td>
<td valign="top" align="center">65.08</td>
<td valign="top" align="center"><bold>72.86</bold></td>
<td valign="top" align="center">42.25</td>
<td valign="top" align="center">74.56</td>
<td valign="top" align="center">67.98</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">4-Fold</td>
<td valign="top" align="center">70.80</td>
<td valign="top" align="center"><bold>82.72</bold></td>
<td valign="top" align="center">60.09</td>
<td valign="top" align="center">65.06</td>
<td valign="top" align="center"><bold>72.83</bold></td>
<td valign="top" align="center">42.22</td>
<td valign="top" align="center">74.55</td>
<td valign="top" align="center">67.96</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">5-Fold</td>
<td valign="top" align="center">70.79</td>
<td valign="top" align="center"><bold>82.70</bold></td>
<td valign="top" align="center">60.09</td>
<td valign="top" align="center">65.05</td>
<td valign="top" align="center"><bold>72.82</bold></td>
<td valign="top" align="center">42.21</td>
<td valign="top" align="center">74.54</td>
<td valign="top" align="center">67.95</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">6-Fold</td>
<td valign="top" align="center">70.83</td>
<td valign="top" align="center"><bold>82.72</bold></td>
<td valign="top" align="center">60.14</td>
<td valign="top" align="center">65.10</td>
<td valign="top" align="center"><bold>72.86</bold></td>
<td valign="top" align="center">42.28</td>
<td valign="top" align="center">74.55</td>
<td valign="top" align="center">68.00</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">7-Fold</td>
<td valign="top" align="center">70.81</td>
<td valign="top" align="center"><bold>82.70</bold></td>
<td valign="top" align="center">60.12</td>
<td valign="top" align="center">65.08</td>
<td valign="top" align="center"><bold>72.84</bold></td>
<td valign="top" align="center">42.24</td>
<td valign="top" align="center">74.55</td>
<td valign="top" align="center">67.98</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">8-Fold</td>
<td valign="top" align="center">70.81</td>
<td valign="top" align="center"><bold>82.70</bold></td>
<td valign="top" align="center">60.12</td>
<td valign="top" align="center">65.08</td>
<td valign="top" align="center"><bold>72.84</bold></td>
<td valign="top" align="center">42.24</td>
<td valign="top" align="center">74.56</td>
<td valign="top" align="center">67.97</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">9-Fold</td>
<td valign="top" align="center">70.80</td>
<td valign="top" align="center"><bold>82.72</bold></td>
<td valign="top" align="center">60.09</td>
<td valign="top" align="center">65.07</td>
<td valign="top" align="center"><bold>72.84</bold></td>
<td valign="top" align="center">42.23</td>
<td valign="top" align="center">74.56</td>
<td valign="top" align="center">67.97</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">10-Fold</td>
<td valign="top" align="center">70.81</td>
<td valign="top" align="center"><bold>82.71</bold></td>
<td valign="top" align="center">60.10</td>
<td valign="top" align="center">65.07</td>
<td valign="top" align="center"><bold>72.84</bold></td>
<td valign="top" align="center">42.24</td>
<td valign="top" align="center">74.54</td>
<td valign="top" align="center">67.97</td>
</tr>
<tr>
<td valign="top" align="left">Na&#x000EF;ve-Bayes</td>
<td valign="top" align="left">1-Fold</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center">72.13</td>
<td valign="top" align="center">70.87</td>
<td valign="top" align="center">69.00</td>
<td valign="top" align="center">70.53</td>
<td valign="top" align="center">42.90</td>
<td valign="top" align="center">78.47</td>
<td valign="top" align="center">69.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">2-Fold</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center">72.08</td>
<td valign="top" align="center">70.90</td>
<td valign="top" align="center">68.98</td>
<td valign="top" align="center">70.50</td>
<td valign="top" align="center">42.89</td>
<td valign="top" align="center">78.47</td>
<td valign="top" align="center">69.58</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3-Fold</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center">72.10</td>
<td valign="top" align="center">70.89</td>
<td valign="top" align="center">69.01</td>
<td valign="top" align="center">70.52</td>
<td valign="top" align="center">42.89</td>
<td valign="top" align="center">78.47</td>
<td valign="top" align="center">69.61</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">4-Fold</td>
<td valign="top" align="center">71.47</td>
<td valign="top" align="center">72.10</td>
<td valign="top" align="center">70.90</td>
<td valign="top" align="center">69.00</td>
<td valign="top" align="center">70.52</td>
<td valign="top" align="center">42.90</td>
<td valign="top" align="center">78.47</td>
<td valign="top" align="center">69.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">5-Fold</td>
<td valign="top" align="center">71.45</td>
<td valign="top" align="center">72.10</td>
<td valign="top" align="center">70.86</td>
<td valign="top" align="center">68.97</td>
<td valign="top" align="center">70.50</td>
<td valign="top" align="center">42.87</td>
<td valign="top" align="center">78.45</td>
<td valign="top" align="center">69.57</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">6-Fold</td>
<td valign="top" align="center">71.47</td>
<td valign="top" align="center">72.13</td>
<td valign="top" align="center">70.88</td>
<td valign="top" align="center">69.00</td>
<td valign="top" align="center">70.53</td>
<td valign="top" align="center">42.90</td>
<td valign="top" align="center">78.48</td>
<td valign="top" align="center">69.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">7-Fold</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center">72.11</td>
<td valign="top" align="center">70.88</td>
<td valign="top" align="center">69.00</td>
<td valign="top" align="center">70.52</td>
<td valign="top" align="center">42.89</td>
<td valign="top" align="center">78.46</td>
<td valign="top" align="center">69.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">8-Fold</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center">72.10</td>
<td valign="top" align="center">70.88</td>
<td valign="top" align="center">69.00</td>
<td valign="top" align="center">70.52</td>
<td valign="top" align="center">42.89</td>
<td valign="top" align="center">78.45</td>
<td valign="top" align="center">69.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">9-Fold</td>
<td valign="top" align="center">71.45</td>
<td valign="top" align="center">72.09</td>
<td valign="top" align="center">70.87</td>
<td valign="top" align="center">68.98</td>
<td valign="top" align="center">70.50</td>
<td valign="top" align="center">42.87</td>
<td valign="top" align="center">78.48</td>
<td valign="top" align="center">69.58</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">10-Fold</td>
<td valign="top" align="center">71.45</td>
<td valign="top" align="center">72.12</td>
<td valign="top" align="center">70.85</td>
<td valign="top" align="center">68.97</td>
<td valign="top" align="center">70.51</td>
<td valign="top" align="center">42.86</td>
<td valign="top" align="center">78.47</td>
<td valign="top" align="center">69.58</td>
</tr>
<tr>
<td valign="top" align="left">Logistic Regression</td>
<td valign="top" align="left">1-Fold</td>
<td valign="top" align="center"><bold>72.23</bold></td>
<td valign="top" align="center">73.52</td>
<td valign="top" align="center"><bold>71.07</bold></td>
<td valign="top" align="center"><bold>69.56</bold></td>
<td valign="top" align="center">71.49</td>
<td valign="top" align="center"><bold>44.47</bold></td>
<td valign="top" align="center"><bold>79.80</bold></td>
<td valign="top" align="center"><bold>70.32</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">2-Fold</td>
<td valign="top" align="center"><bold>72.21</bold></td>
<td valign="top" align="center">73.46</td>
<td valign="top" align="center"><bold>71.10</bold></td>
<td valign="top" align="center"><bold>69.52</bold></td>
<td valign="top" align="center">71.44</td>
<td valign="top" align="center"><bold>44.43</bold></td>
<td valign="top" align="center"><bold>79.79</bold></td>
<td valign="top" align="center"><bold>70.27</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3-Fold</td>
<td valign="top" align="center"><bold>72.23</bold></td>
<td valign="top" align="center">73.54</td>
<td valign="top" align="center"><bold>71.06</bold></td>
<td valign="top" align="center"><bold>69.56</bold></td>
<td valign="top" align="center">71.49</td>
<td valign="top" align="center"><bold>44.47</bold></td>
<td valign="top" align="center"><bold>79.80</bold></td>
<td valign="top" align="center"><bold>70.32</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">4-Fold</td>
<td valign="top" align="center"><bold>72.24</bold></td>
<td valign="top" align="center">73.51</td>
<td valign="top" align="center"><bold>71.09</bold></td>
<td valign="top" align="center"><bold>69.56</bold></td>
<td valign="top" align="center">71.48</td>
<td valign="top" align="center"><bold>44.47</bold></td>
<td valign="top" align="center"><bold>79.80</bold></td>
<td valign="top" align="center"><bold>70.31</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">5-Fold</td>
<td valign="top" align="center"><bold>72.20</bold></td>
<td valign="top" align="center">73.48</td>
<td valign="top" align="center"><bold>71.05</bold></td>
<td valign="top" align="center"><bold>69.51</bold></td>
<td valign="top" align="center">71.44</td>
<td valign="top" align="center"><bold>44.41</bold></td>
<td valign="top" align="center"><bold>79.77</bold></td>
<td valign="top" align="center"><bold>70.27</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">6-Fold</td>
<td valign="top" align="center"><bold>72.24</bold></td>
<td valign="top" align="center">73.53</td>
<td valign="top" align="center"><bold>71.07</bold></td>
<td valign="top" align="center"><bold>69.55</bold></td>
<td valign="top" align="center">71.49</td>
<td valign="top" align="center"><bold>44.47</bold></td>
<td valign="top" align="center"><bold>79.80</bold></td>
<td valign="top" align="center"><bold>70.31</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">7-Fold</td>
<td valign="top" align="center"><bold>72.23</bold></td>
<td valign="top" align="center">73.52</td>
<td valign="top" align="center"><bold>71.08</bold></td>
<td valign="top" align="center"><bold>69.56</bold></td>
<td valign="top" align="center">71.48</td>
<td valign="top" align="center"><bold>44.47</bold></td>
<td valign="top" align="center"><bold>79.78</bold></td>
<td valign="top" align="center"><bold>70.32</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">8-Fold</td>
<td valign="top" align="center"><bold>72.22</bold></td>
<td valign="top" align="center">73.52</td>
<td valign="top" align="center"><bold>71.06</bold></td>
<td valign="top" align="center"><bold>69.54</bold></td>
<td valign="top" align="center">71.48</td>
<td valign="top" align="center"><bold>44.45</bold></td>
<td valign="top" align="center"><bold>79.78</bold></td>
<td valign="top" align="center"><bold>70.30</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">9-Fold</td>
<td valign="top" align="center"><bold>72.24</bold></td>
<td valign="top" align="center">73.52</td>
<td valign="top" align="center"><bold>71.08</bold></td>
<td valign="top" align="center"><bold>69.55</bold></td>
<td valign="top" align="center">71.48</td>
<td valign="top" align="center"><bold>44.48</bold></td>
<td valign="top" align="center"><bold>79.81</bold></td>
<td valign="top" align="center"><bold>70.31</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">10-Fold</td>
<td valign="top" align="center"><bold>72.22</bold></td>
<td valign="top" align="center">73.52</td>
<td valign="top" align="center"><bold>71.05</bold></td>
<td valign="top" align="center"><bold>69.54</bold></td>
<td valign="top" align="center">71.48</td>
<td valign="top" align="center"><bold>44.45</bold></td>
<td valign="top" align="center"><bold>79.79</bold></td>
<td valign="top" align="center"><bold>70.30</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold values shows in which aspect does the ML methods performed best</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>The prediction performance for the classification of obesity status from these methods is also assessed based on accuracy, sensitivity, specificity, precision, recall, F1-score, kappa, and <italic>F</italic><sub>&#x003B2;</sub>. The measurement results of these metrics based on the 10-fold CV using ML methods for the obesity data set can be seen in <xref ref-type="table" rid="T3">Table 3</xref>. Based on <xref ref-type="table" rid="T3">Table 3</xref>, the classification prediction using the Logistic Regression method achieves the best performance based on the accuracy metric (72%), specificity (71%), precision (69%), Kappa (44%), and <italic>F</italic><sub>&#x003B2;</sub> (70%). Classification prediction by the CART method achieves the highest sensitivity (82%) and the highest F1-score (72%).</p>
<p><xref ref-type="fig" rid="F1">Figures 1</xref>&#x02013;<xref ref-type="fig" rid="F3">3</xref> show AUC performance of the respective classification methods with 10-fold CV. The results show that the Logistic Regression classifier has the highest average AUC values (0.798) (<xref ref-type="fig" rid="F3">Figure 3</xref>). In addition to comparing the AUC values obtained, the accuracy, sensitivity, specificity, precision, F1-Score, and <italic>F</italic><sub>&#x003B2;</sub> values of each method can also be considered. The AUC is a classification threshold invariant metric that measures the predictive quality of a model regardless of which classification threshold is selected.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>AUC performance of the classification methods with 10-fold CV using the CART method.</p></caption>
<graphic xlink:href="fnut-08-669155-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>AUC performance on the classification method with the 10-fold CV using the Na&#x000EF;ve Bayes method.</p></caption>
<graphic xlink:href="fnut-08-669155-g0002.tif"/>
</fig>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>AUC performance on the classification method with the 10-fold CV using the Logistic Regression method.</p></caption>
<graphic xlink:href="fnut-08-669155-g0003.tif"/>
</fig>
<p>After calculating the classification performance for correctly determining the obesity status for each of the 3 different models, it is also necessary to estimate a set of risk factors for obesity among the available study variables. Based on the evaluation of classification prediction performance, the Logistic Regression method had the better performance compared with the CART method and the Na&#x000EF;ve Bayes method. Overall, fold 6 out of 10-fold CV showed the best accuracy for the classification performance of the obesity status. Partial testing of parameters of the Logistic Regression model using the Wald test showed that all explanatory variables qualify as factors that can affect the obesity status (<xref ref-type="table" rid="T4">Table 4</xref>). From <xref ref-type="table" rid="T4">Table 4</xref>, the variables that have the greatest effect on the obesity status in adults (<italic>p</italic>-value &#x0003C;0.05) included location (X1), marital status (X2), age groups (X3), education (X4), sweet drinks (X7), fatty/oily foods (X9), grilled foods (X10), preserved foods (X11), seasoning powders (X12), soft/carbonated drinks (X13), alcoholic drinks (X16), mental emotional disorders (X17), diagnosed hypertension (X18), physical activity (X19), smoking (X20), and fruit and vegetables consumptions (X21).</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Estimation of the Logistic Regression parameters based on fold 6 out of the 10-fold CV for obesity dataset in Indonesian RISKESDAS 2018 survey.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Descriptive of variables</bold></th>
<th valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Fold 6 out of 10-fold CV Test</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="center"><bold>&#x003B2;</bold></th>
<th valign="top" align="center"><bold>SE</bold></th>
<th valign="top" align="center"><bold>Wald</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-Value</bold></th>
<th valign="top" align="center"><bold>Odd Ratio</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="2">Constant</td>
<td valign="top" align="center">6.510</td>
<td valign="top" align="center">0.046</td>
<td valign="top" align="center">142.754</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">671.976</td>
</tr>
<tr>
<td valign="top" align="left">Location (X1)</td>
<td valign="top" align="left">Rural</td>
<td valign="top" align="center">&#x02212;0.305</td>
<td valign="top" align="center">0.005</td>
<td valign="top" align="center">&#x02212;59.121</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.737</td>
</tr>
<tr>
<td valign="top" align="left">Marital status (X2)</td>
<td valign="top" align="left">Married</td>
<td valign="top" align="center">&#x02212;0.363</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">&#x02212;50.033</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.695</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Divorced</td>
<td valign="top" align="center">0.271</td>
<td valign="top" align="center">0.015</td>
<td valign="top" align="center">18.000</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.311</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Widowed</td>
<td valign="top" align="center">0.289</td>
<td valign="top" align="center">0.012</td>
<td valign="top" align="center">24.963</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.335</td>
</tr>
<tr>
<td valign="top" align="left">Age groups (X3)</td>
<td valign="top" align="left">25&#x02013;29 years</td>
<td valign="top" align="center">0.488</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">46.674</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.630</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">30&#x02013;34 years</td>
<td valign="top" align="center">0.560</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">52.679</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.750</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">35&#x02013;39 years</td>
<td valign="top" align="center">0.680</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">64.375</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.975</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">40&#x02013;44 years</td>
<td valign="top" align="center">0.746</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">69.255</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">2.110</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">45&#x02013;49 years</td>
<td valign="top" align="center">0.741</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">67.743</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">2.097</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">50&#x02013;54 years</td>
<td valign="top" align="center">0.549</td>
<td valign="top" align="center">0.012</td>
<td valign="top" align="center">46.783</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.731</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">55&#x02013;59 years</td>
<td valign="top" align="center">0.333</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">26.349</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.396</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">60&#x02013;64 years</td>
<td valign="top" align="center">0.304</td>
<td valign="top" align="center">0.014</td>
<td valign="top" align="center">21.859</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.355</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003E;64 years</td>
<td valign="top" align="center">&#x02212;0.457</td>
<td valign="top" align="center">0.014</td>
<td valign="top" align="center">&#x02212;32.580</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.633</td>
</tr>
<tr>
<td valign="top" align="left">Education (X4)</td>
<td valign="top" align="left">Not finished basic school</td>
<td valign="top" align="center">0.313</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">24.156</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.367</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished basic school</td>
<td valign="top" align="center">0.361</td>
<td valign="top" align="center">0.012</td>
<td valign="top" align="center">29.692</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.435</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished Junior High School</td>
<td valign="top" align="center">0.456</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">35.808</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.577</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished Senior High School</td>
<td valign="top" align="center">0.469</td>
<td valign="top" align="center">0.012</td>
<td valign="top" align="center">38.083</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.598</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished Academy/College</td>
<td valign="top" align="center">0.502</td>
<td valign="top" align="center">0.018</td>
<td valign="top" align="center">28.496</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.652</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Finished higher education</td>
<td valign="top" align="center">0.506</td>
<td valign="top" align="center">0.015</td>
<td valign="top" align="center">33.432</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.659</td>
</tr>
<tr>
<td valign="top" align="left">Work types (X5)</td>
<td valign="top" align="left">School</td>
<td valign="top" align="center">&#x02212;0.356</td>
<td valign="top" align="center">0.018</td>
<td valign="top" align="center">&#x02212;19.850</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.700</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Government employee</td>
<td valign="top" align="center">0.197</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">15.224</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.218</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Private employee</td>
<td valign="top" align="center">&#x02212;0.117</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">&#x02212;12.055</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.889</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Entrepreneur</td>
<td valign="top" align="center">0.069</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">8.797</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.072</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Farmer</td>
<td valign="top" align="center">&#x02212;0.548</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">&#x02212;74.090</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.578</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Fisherman</td>
<td valign="top" align="center">&#x02212;0.838</td>
<td valign="top" align="center">0.024</td>
<td valign="top" align="center">&#x02212;35.437</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.432</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Daily waged labors</td>
<td valign="top" align="center">&#x02212;0.389</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">&#x02212;39.463</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.678</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Others</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">0.987</td>
<td valign="top" align="center">0.324</td>
<td valign="top" align="center">1.010</td>
</tr>
<tr>
<td valign="top" align="left">Sugary foods (X6)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">&#x02212;0.135</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;15.096</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.874</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">&#x02212;0.141</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;15.938</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.869</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.158</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;18.457</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.854</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">1.189</td>
<td valign="top" align="center">0.234</td>
<td valign="top" align="center">1.013</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.101</td>
<td valign="top" align="center">0.014</td>
<td valign="top" align="center">&#x02212;7.308</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.904</td>
</tr>
<tr>
<td valign="top" align="left">Sweet drinks (X7)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">0.094</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">13.815</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.099</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">0.148</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">17.454</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.159</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">0.189</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">22.735</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.208</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">0.313</td>
<td valign="top" align="center">0.012</td>
<td valign="top" align="center">26.572</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.368</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">0.297</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">23.106</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.346</td>
</tr>
<tr>
<td valign="top" align="left">Salty foods (X8)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">0.070</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">6.824</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.073</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">&#x02212;0.077</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">&#x02212;7.773</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.926</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.113</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;12.268</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.893</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">&#x02212;0.056</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">&#x02212;5.640</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.946</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.016</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">&#x02212;1.568</td>
<td valign="top" align="center">0.117</td>
<td valign="top" align="center">0.984</td>
</tr>
<tr>
<td valign="top" align="left">Fatty/Oily foods (X9)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">&#x02212;0.092</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;10.707</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.913</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">&#x02212;0.158</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">&#x02212;19.229</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.854</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.165</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">&#x02212;20.722</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.848</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">&#x02212;0.184</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">&#x02212;18.937</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.832</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.495</td>
<td valign="top" align="center">0.014</td>
<td valign="top" align="center">&#x02212;35.457</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.609</td>
</tr>
<tr>
<td valign="top" align="left">Grilled foods (X10)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">&#x02212;0.184</td>
<td valign="top" align="center">0.019</td>
<td valign="top" align="center">&#x02212;9.749</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.832</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">&#x02212;0.311</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">&#x02212;18.881</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.733</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.419</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">&#x02212;26.825</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.658</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">&#x02212;0.430</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">&#x02212;27.690</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.651</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.452</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">&#x02212;28.697</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.636</td>
</tr>
<tr>
<td valign="top" align="left">Preserved foods (X11)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">&#x02212;0.465</td>
<td valign="top" align="center">0.025</td>
<td valign="top" align="center">&#x02212;18.674</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.628</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">&#x02212;0.550</td>
<td valign="top" align="center">0.022</td>
<td valign="top" align="center">&#x02212;25.115</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.577</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.597</td>
<td valign="top" align="center">0.021</td>
<td valign="top" align="center">&#x02212;28.800</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.551</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">&#x02212;0.694</td>
<td valign="top" align="center">0.020</td>
<td valign="top" align="center">&#x02212;34.273</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.499</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.856</td>
<td valign="top" align="center">0.020</td>
<td valign="top" align="center">&#x02212;42.964</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.425</td>
</tr>
<tr>
<td valign="top" align="left">Seasonings powders (X12)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">0.117</td>
<td valign="top" align="center">0.006</td>
<td valign="top" align="center">19.308</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.124</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">0.276</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">27.709</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.318</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">0.229</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">20.837</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.257</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">0.582</td>
<td valign="top" align="center">0.013</td>
<td valign="top" align="center">46.073</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.789</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">0.399</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">47.027</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.491</td>
</tr>
<tr>
<td valign="top" align="left">Soft/Carbonated drinks (X13)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">0.313</td>
<td valign="top" align="center">0.032</td>
<td valign="top" align="center">9.805</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.368</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">0.156</td>
<td valign="top" align="center">0.029</td>
<td valign="top" align="center">5.284</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.169</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">0.073</td>
<td valign="top" align="center">0.028</td>
<td valign="top" align="center">2.621</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">1.076</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">&#x02212;0.158</td>
<td valign="top" align="center">0.027</td>
<td valign="top" align="center">&#x02212;5.753</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.854</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.457</td>
<td valign="top" align="center">0.027</td>
<td valign="top" align="center">&#x02212;16.900</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.633</td>
</tr>
<tr>
<td valign="top" align="left">Energy drinks (X14)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">0.046</td>
<td valign="top" align="center">0.031</td>
<td valign="top" align="center">1.476</td>
<td valign="top" align="center">0.140</td>
<td valign="top" align="center">1.047</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">0.020</td>
<td valign="top" align="center">0.029</td>
<td valign="top" align="center">0.681</td>
<td valign="top" align="center">0.496</td>
<td valign="top" align="center">1.020</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.032</td>
<td valign="top" align="center">0.027</td>
<td valign="top" align="center">&#x02212;1.185</td>
<td valign="top" align="center">0.236</td>
<td valign="top" align="center">0.968</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">&#x02212;0.095</td>
<td valign="top" align="center">0.027</td>
<td valign="top" align="center">&#x02212;3.549</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.909</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.713</td>
<td valign="top" align="center">0.026</td>
<td valign="top" align="center">&#x02212;27.394</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.490</td>
</tr>
<tr>
<td valign="top" align="left">Instant foods (X15)</td>
<td valign="top" align="left">1 times per day</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">0.019</td>
<td valign="top" align="center">0.512</td>
<td valign="top" align="center">0.609</td>
<td valign="top" align="center">1.010</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3&#x02013;6 times per week</td>
<td valign="top" align="center">0.048</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">2.767</td>
<td valign="top" align="center">0.006</td>
<td valign="top" align="center">1.049</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">1&#x02013;2 times per week</td>
<td valign="top" align="center">&#x02212;0.063</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">&#x02212;3.710</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.939</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">&#x0003C;3 times per month</td>
<td valign="top" align="center">0.084</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">4.901</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.088</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Never</td>
<td valign="top" align="center">&#x02212;0.009</td>
<td valign="top" align="center">0.018</td>
<td valign="top" align="center">&#x02212;0.533</td>
<td valign="top" align="center">0.594</td>
<td valign="top" align="center">0.991</td>
</tr>
<tr>
<td valign="top" align="left">Alcoholic drinks (X16)</td>
<td valign="top" align="left">No</td>
<td valign="top" align="center">&#x02212;1.576</td>
<td valign="top" align="center">0.008</td>
<td valign="top" align="center">&#x02212;190.048</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.207</td>
</tr>
<tr>
<td valign="top" align="left">Mental-emotional disorders (X17)</td>
<td valign="top" align="left">No</td>
<td valign="top" align="center">&#x02212;1.029</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">&#x02212;150.755</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.357</td>
</tr>
<tr>
<td valign="top" align="left">Diagnosed hypertension (X18)</td>
<td valign="top" align="left">No</td>
<td valign="top" align="center">&#x02212;0.867</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;100.728</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.420</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Unknown</td>
<td valign="top" align="center">&#x02212;0.982</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;110.600</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.375</td>
</tr>
<tr>
<td valign="top" align="left">Physical activity (X19)</td>
<td valign="top" align="left">Not adequate</td>
<td valign="top" align="center">&#x02212;0.852</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">&#x02212;128.275</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.427</td>
</tr>
<tr>
<td valign="top" align="left">Smoking (X20)</td>
<td valign="top" align="left">No</td>
<td valign="top" align="center">0.219</td>
<td valign="top" align="center">0.005</td>
<td valign="top" align="center">41.165</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">1.244</td>
</tr>
<tr>
<td valign="top" align="left">Fruit and vegetables consumptions (X21)</td>
<td valign="top" align="left">Not adequate</td>
<td valign="top" align="center">&#x02212;1.248</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">&#x02212;135.504</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.287</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition to the Logistic Regression method, prediction of obesity classification also used CART and Na&#x000EF;ve Bayes methods. From <xref ref-type="fig" rid="F4">Figure 4</xref>, it can be seen that the characteristics of the variables that influence the occurrence of obesity in the Indonesia RISKESDAS 2018 are significant variables that function as the main partitioning of all the trees produced. In this case, the main partitioning variables for 10% test data with fold 6 out of the 10-fold CV are alcoholic drinks (X16). The order of important variables in this CART model are alcoholic drinks (X16), energy drinks (X14), soft/carbonated drinks (X13), mental-emotional disorders (X17), fruit and vegetables consumptions (X21), diagnosed hypertension (X18), physical activity (X19), and marital status (X2).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Obesity data classification tree for fold 6 out of the 10-fold CV for CART model based on the variables of alcoholic drinks (X16), energy drinks (X14), soft/carbonated drinks (X13), mental-emotional disorders (X17), Fruit and Vegetables Consumptions (X21), diagnosed hypertension (X18), Physical Activity (X19), and Marital Status (X2).</p></caption>
<graphic xlink:href="fnut-08-669155-g0004.tif"/>
</fig>
<p>Obesity prediction using the Na&#x000EF;ve Bayes model was also done by looking for values of <italic>P</italic>(<italic>C</italic><sub><italic>i</italic></sub>) for the obese class and <italic>P</italic>(<italic>C</italic><sub><italic>j</italic></sub>)for the non-obese class. In this case, the value of <italic>i</italic> = 1 and the value of <italic>j</italic> = 2. The probability value for each variable on the class label is presented in detail in the <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 3</xref>.</p></sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>We have conducted a study to establish a set of risk factors for obesity in adults among the available study variables using ML methods using publicly available data on RISKESDAS (RISKESDAS 2018). In this study, three methods (Logistic Regression, CART, and Na&#x000EF;ve Bayes) were used in the ML approach to select a method that produces predictions with high accuracy. The result revealed that the Logistic Regression method shows a better accuracy compared to the other methods with AUC = 0.798 using 21 variables, namely location (X1), marital status (X2), age groups (X3), education (X4), work types (X5), sugary foods (X6), sweet drinks (X7), fatty/oily foods (X9), grilled foods (X10), preserved foods (X11), seasoning powders (X12), soft/carbonated drinks (X13), energy drinks (X14), instant foods (X15), alcoholic drinks (X16), mental emotional disorders (X17), diagnosed hypertension (X18), physical activity (X19), smoking (X20), and fruit and vegetables consumptions (X21).</p>
<p>With the accelerated economic growth and lifestyle changes around the world, including in Indonesia, it is important to evaluate and build predictive models for obesity using common risk factors. Based on RISKESDAS 2013 and 2018, Indonesia as a middle-income country seems to underestimate the significance of actual obesity cases even though there has been a significant increase in cases. As shown in this study, the 21 selected measures play a prominent role in increasing the risk for obesity in adults. This is in parallel with some previous studies. In their study, Roemling and Qaim (<xref ref-type="bibr" rid="B4">4</xref>) found that obesity risk in Indonesia occurred both in rural and urban areas and was closely associated with food consumption pattern changes coupled with physical activity decreases. Rachmi et al. (<xref ref-type="bibr" rid="B5">5</xref>) showed that the increasing prevalence of overweight children, adolescents, and adults in Indonesia over the past two decades coincides with higher numbers of obesity in urban areas. Similarly, Oddo et al. (<xref ref-type="bibr" rid="B6">6</xref>) demonstrated that there were more obesity cases in rural areas compared to the past even though the overall case numbers are still higher in urban areas in Indonesia. They also showed that highly processed foods are mostly consumed and decreased physical activities have led to the higher prevalence of obesity. Dewi et al. (<xref ref-type="bibr" rid="B7">7</xref>) found that the consumption of oil and fat, animal source foods, and low physical activities are some of the significant determinants of obesity in Indonesia. Emery et al. (<xref ref-type="bibr" rid="B40">40</xref>) revealed that there was a relationship between less healthy food consumption with obesity. Sinha and Jastreboff (<xref ref-type="bibr" rid="B41">41</xref>) found that eating habits and the increased consumption of food result from stress. Koski and Naukkarinen (<xref ref-type="bibr" rid="B42">42</xref>) strengthened the fact that the development of obesity is significantly due to persistent stress. The difference in confounding factors involved in the analysis is one of the reasons for the differences found in this study with previous studies.</p>
<p>In this study, we employed the metrics for accuracy, sensitivity, specificity, precision, recall, F1-score, kappa, and <italic>F</italic><sub>&#x003B2;</sub> with 10-fold CV for performance evaluation of the three classification methods. The results obtained are the prediction of the classification with 10-fold CV using the Logistic Regression method, which achieved the best performance as assessed by the accuracy metric (72%), specificity (71%), precision (69%), kappa (44%), and <italic>F</italic><sub>&#x003B2; &#x0003D; 0.5</sub> (70%). Classification prediction by the CART method achieved the highest sensitivity (82%), and F1-score (72%). The Na&#x000EF;ve Bayes method had an accuracy of 71% and a <italic>F</italic><sub>&#x003B2; &#x0003D; 0.5</sub> of 69%.</p>
<p>In general, this ML approach is an alternative to the classical methods used so far (<xref ref-type="bibr" rid="B22">22</xref>). Using ML methods on public health data can help to improve predictions and find a rich structure among available data and increase understanding of complex problems in public health, including risk factors for obesity with ML. The ML method could inform the design of more appropriate health policies and programs to address Non-Communicable Diseases, most notably in predicting obesity incidence/prevalence, and in turn, reducing severity as well as the cost of treating obesity and obesity-related condition which eventually could improve the health and well-being of the population. Apart from that, the ML method as shown in the current study could be utilized to identify the most significant risk factors for predicting obesity status can be applied to publicly available data, such as RISKESDAS data.</p>
<p>In general, RISKESDAS provides an overview of Indonesian health indicators, such as health status, health services, health behavior, and environmental health. RISKESDAS is supposedly the best data available on health in Indonesia but its main limitation is the fact that the purpose and nature of RISKESDAS are based on a periodic study (every 5 years) examining a broad range of health issues and health behaviors. This then results in a data set that lacks depth.</p>
<p>In Indonesia, policies on obesity prevention and control in adults are related to limiting consumption of fats and oils, sugary foods and carbohydrates, and increasing vegetable intake are carried out through the Health Community Movement, known as GERMAS and the Food Label with the inclusion of sugar, salt, and fat content on food labels (<xref ref-type="bibr" rid="B7">7</xref>). Yet, these efforts seem to be ineffective as the increase in the proportion of obesity remains relatively high. The findings of this study in predicting the risk factor for obesity among the available study variables on RISKESDAS 2018 can then convince the policy makers in Indonesia (primarily the government) to put more attention into the pressing obesity problems. As a result, the effectiveness of existing program policies could be further improved and the financing of the health care system can be made more efficient (<xref ref-type="bibr" rid="B43">43</xref>).</p>
<p>This study provides an overview of the methods available for predicting risk factors for obesity in adults among the available study variables in Indonesia. Several factors that might influence obesity (e.g., sex, dietary quality, clinical and physiological, wealth, genetic and cultural influences) were not included in this study, and thereby, the relationship between these factors and obesity cannot be explained further. Further research needs to be carried out using large datasets with individual subjects to confirm the results of this study and to describe the variation in the results for individual regions.</p></sec>
<sec sec-type="conclusions" id="s5">
<title>Conclusion</title>
<p>The Logistic Regression method showed better results on the accuracy, specificity, precision, kappa, and <italic>F</italic><sub>&#x003B2;</sub> metrics. Meanwhile, the CART method showed better results on the sensitivity, recall, and F1-score. For the 10-fold CV, the Logistic Regression method had the highest AUC performance which was 0.798. Then, from the Logistic Regression method, it can also be seen that the variables that affect the prediction of obesity status in adults are location, marital status, age groups, education, sweet drinks, fatty/oily foods, grilled foods, preserved foods, seasoning powders, soft/carbonated drinks, alcoholic drinks, mental emotional disorders, diagnosed hypertension, physical activity, smoking, and fruit and vegetables consumptions. The constructed obesity classification model can evaluate and predict the risk of obesity using ML methods for the population of Indonesia which can then be applied to publicly available open data, such as the RISKESDAS survey data. In general, this study has been able to establish a set of risk factors for obesity in adults among the available study variables. However, more studies should be done to further improve the quality of predictions by exploring other ML models. In the future work, we will validate the results with other relevant groups. Additionally, we will also evaluate differences in the prediction of obesity status at the district/city or province level in Indonesia with regional disaggregation.</p></sec>
<sec sec-type="data-availability-statement" id="s6">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="https://www.litbang.kemkes.go.id/layanan-permintaan-data-riset">https://www.litbang.kemkes.go.id/layanan-permintaan-data-riset</ext-link>.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>ST contributed to the concept and design of the study, carried out the statistical analysis, and wrote the manuscript. DA interpreted the data, analyzed, and wrote the manuscript. HK collected the necessary data and carried out the statistical analysis. AL interpreted the data and analyzed the manuscript. SN analyzed and wrote the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</body>
<back>
<ack><p>ST would like to thank the Ministry of Research and Technology/National Research and Innovation Agency for funding this research through the PDUPT Scheme for the 2020 fiscal year. In addition, the authors would also like to thank the Ministry of Health through the Community Research and Development Agency for providing access to the Indonesian RISKESDAS survey data.</p>
</ack>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fnut.2021.669155/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fnut.2021.669155/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/></sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="web"><person-group person-group-type="author"><collab>ASEAN/UNICEF/WHO Regional Report</collab></person-group>. <source>World Health Statistics 2016: Monitoring Health for the SDGs, Sustainable Development Goals</source>. (<year>2016</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.who.int/about/licensing/copyright_form/en/index.html">https://www.who.int/about/licensing/copyright_form/en/index.html</ext-link></citation></ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="web"><person-group person-group-type="author"><collab>Institute of Health Research and Development</collab></person-group>. <source>Basic Health Research Reports</source>. (<year>2013</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.litbang.kemkes.go.id/laporan-riset-kesehatan-dasar-riskesdas/">https://www.litbang.kemkes.go.id/laporan-riset-kesehatan-dasar-riskesdas/</ext-link></citation></ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="web"><person-group person-group-type="author"><collab>Institute of Health Research and Development</collab></person-group>. <source>Basic Health Research Reports</source>. (<year>2018</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="http://labdata.litbang.kemkes.go.id/images/download/laporan/RKD/2013/Laporan_riskesdas_2013_final.pdf">http://labdata.litbang.kemkes.go.id/images/download/laporan/RKD/2013/Laporan_riskesdas_2013_final.pdf</ext-link></citation></ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roemling</surname> <given-names>C</given-names></name> <name><surname>Qaim</surname> <given-names>M</given-names></name></person-group>. <article-title>Obesity trends and determinants in Indonesia</article-title>. <source>Appetite.</source> (<year>2012</year>) <volume>58</volume>:<fpage>1005</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1016/j.appet.2012.02.053</pub-id></citation></ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rachmi</surname> <given-names>CN</given-names></name> <name><surname>Li</surname> <given-names>M</given-names></name> <name><surname>Alison Baur</surname> <given-names>L</given-names></name></person-group>. <article-title>Overweight and obesity in Indonesia: prevalence and risk factors-a literature review</article-title>. <source>Public Health.</source> (<year>2017</year>) <volume>147</volume>:<fpage>20</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/j.puhe.2017.02.002</pub-id><pub-id pub-id-type="pmid">28404492</pub-id></citation></ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oddo</surname> <given-names>VM</given-names></name> <name><surname>Maehara</surname> <given-names>M</given-names></name> <name><surname>Rah</surname> <given-names>JH</given-names></name></person-group>. <article-title>Overweight in Indonesia: an observational study of trends and risk factors among adults and children</article-title>. <source>BMJ Open.</source> (<year>2019</year>) <volume>9</volume>:<fpage>e031198</fpage>. <pub-id pub-id-type="doi">10.1136/bmjopen-2019-031198</pub-id><pub-id pub-id-type="pmid">31562157</pub-id></citation></ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dewi</surname> <given-names>NU</given-names></name> <name><surname>Tanziha</surname> <given-names>I</given-names></name> <name><surname>Solechah</surname> <given-names>SA</given-names></name> <name><surname>Bohari</surname> <given-names>B</given-names></name></person-group>. <article-title>Obesity determinants and the policy implications for the prevention and management of obesity in Indonesia</article-title>. <source>Curr Res Nutr Food Sci J.</source> (<year>2020</year>) <volume>8</volume>:<fpage>942</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.12944/CRNFS.8.3.22</pub-id></citation></ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wiemken</surname> <given-names>TL</given-names></name> <name><surname>Kelley</surname> <given-names>RR</given-names></name></person-group>. <article-title>Machine learning in epidemiology and health outcomes research</article-title>. <source>Annu Rev Public Health.</source> (<year>2020</year>) <volume>41</volume>:<fpage>21</fpage>&#x02013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-publhealth-040119-094437</pub-id><pub-id pub-id-type="pmid">31577910</pub-id></citation></ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Giabbanelli</surname> <given-names>PJ</given-names></name> <name><surname>Adams</surname> <given-names>J</given-names></name></person-group>. <article-title>Identifying small groups of foods that can predict achievement of key dietary recommendations: data mining of the UK National Diet and Nutrition Survey, 2008&#x02013;12</article-title>. <source>Public Health Nutr.</source> (<year>2016</year>) <volume>19</volume>:<fpage>1543</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1017/S1368980016000185</pub-id><pub-id pub-id-type="pmid">26879185</pub-id></citation></ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dugan</surname> <given-names>TM</given-names></name> <name><surname>Mukhopadhyay</surname> <given-names>S</given-names></name> <name><surname>Carroll</surname> <given-names>A</given-names></name> <name><surname>Downs</surname> <given-names>S</given-names></name></person-group>. <article-title>Machine learning techniques for prediction of early childhood obesity</article-title>. <source>Appl Clin Inform.</source> (<year>2015</year>) <volume>6</volume>:<fpage>506</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.4338/ACI-2015-03-RA-0036</pub-id><pub-id pub-id-type="pmid">26448795</pub-id></citation></ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nau</surname> <given-names>C</given-names></name> <name><surname>Ellis</surname> <given-names>H</given-names></name> <name><surname>Huang</surname> <given-names>H</given-names></name> <name><surname>Schwartz</surname> <given-names>BS</given-names></name> <name><surname>Hirsch</surname> <given-names>A</given-names></name> <name><surname>Bailey-Davis</surname> <given-names>L</given-names></name> <etal/></person-group>. <article-title>Exploring the forest instead of the trees: an innovative method for defining obesogenic and obesoprotective environments</article-title>. <source>Health Place.</source> (<year>2015</year>) <volume>35</volume>:<fpage>136</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1016/j.healthplace.2015.08.002</pub-id><pub-id pub-id-type="pmid">26398219</pub-id></citation></ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Acharjee</surname> <given-names>A</given-names></name> <name><surname>Ament</surname> <given-names>Z</given-names></name> <name><surname>West</surname> <given-names>JA</given-names></name> <name><surname>Stanley</surname> <given-names>E</given-names></name> <name><surname>Griffin</surname> <given-names>JL</given-names></name></person-group>. <article-title>Integration of metabolomics, lipidomics and clinical data using a machine learning method</article-title>. <source>BMC Bioinformatics.</source> (<year>2016</year>) <volume>17</volume>:<fpage>440</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-016-1292-2</pub-id><pub-id pub-id-type="pmid">28185575</pub-id></citation></ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Selya</surname> <given-names>AS</given-names></name> <name><surname>Anshutz</surname> <given-names>D</given-names></name></person-group>. <article-title>Machine learning for the classification of obesity from dietary and physical activity patterns. In: Giabbanelli P, Mago V, Papageorgiou E, editors</article-title>. <source>Advanced Data Analytics in Health</source>. <publisher-name>Springer</publisher-name> (<year>2018</year>). p. <fpage>77</fpage>&#x02013;<lpage>97</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://doi-org-443.webvpn.fjmu.edu.cn/10.1007/978-3-319-77911-9_5">http://doi-org-443.webvpn.fjmu.edu.cn/10.1007/978-3-319-77911-9_5</ext-link></citation></ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>S</given-names></name> <name><surname>Tjortjis</surname> <given-names>C</given-names></name> <name><surname>Zeng</surname> <given-names>X</given-names></name> <name><surname>Qiao</surname> <given-names>H</given-names></name> <name><surname>Buchan</surname> <given-names>I</given-names></name> <name><surname>Keane</surname> <given-names>J</given-names></name></person-group>. <article-title>Comparing data mining methods with logistic regression in childhood obesity prediction</article-title>. <source>Inform Syst Front.</source> (<year>2009</year>) <volume>11</volume>:<fpage>449</fpage>&#x02013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1007/s10796-009-9157-0</pub-id></citation></ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Adnan</surname> <given-names>MHBM</given-names></name> <name><surname>Husain</surname> <given-names>W</given-names></name> <name><surname>Rashid</surname> <given-names>NA</given-names></name></person-group>. <article-title>Parameter identification and selection for childhood obesity prediction using data mining</article-title>. In: <source>2nd International Conference on Management and Artificial Intelligence</source>. <publisher-loc>Singapore</publisher-loc> (<year>2012</year>). p. <fpage>7</fpage>.</citation></ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toschke</surname> <given-names>AM</given-names></name> <name><surname>Beyerlein</surname> <given-names>A</given-names></name> <name><surname>Von Kries</surname> <given-names>R</given-names></name></person-group>. <article-title>Children at high risk for overweight: a classification and regression trees analysis approach</article-title>. <source>Obes Res.</source> (<year>2005</year>) <volume>13</volume>:<fpage>1270</fpage>&#x02013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1038/oby.2005.151</pub-id><pub-id pub-id-type="pmid">16076998</pub-id></citation></ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Golino</surname> <given-names>HF</given-names></name> <name><surname>Amaral</surname> <given-names>LSB</given-names></name> <name><surname>Duarte</surname> <given-names>SFP</given-names></name> <name><surname>Gomes</surname> <given-names>CMA</given-names></name> <name><surname>Soares</surname> <given-names>J</given-names></name> <name><surname>Reis</surname> <given-names>LA</given-names></name> <etal/></person-group>. <article-title>Predicting increased blood pressure using machine learning</article-title>. <source>J Obes.</source> (<year>2014</year>) <volume>2014</volume>:<fpage>637635</fpage>. <pub-id pub-id-type="doi">10.1155/2014/637635</pub-id><pub-id pub-id-type="pmid">24669313</pub-id></citation></ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>Z</given-names></name> <name><surname>Ruggiero</surname> <given-names>K</given-names></name></person-group>. <article-title>Using machine learning to predict obesity in high school students</article-title>. In: <source>2017 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)</source>. <publisher-loc>Kansas</publisher-loc> (<year>2017</year>). p. <fpage>2132</fpage>&#x02013;<lpage>2138</lpage>. <pub-id pub-id-type="doi">10.1109/BIBM.2017.8217988</pub-id></citation></ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chatterjee</surname> <given-names>A</given-names></name> <name><surname>Gerdes</surname> <given-names>MW</given-names></name> <name><surname>Martinez</surname> <given-names>SG</given-names></name></person-group>. <article-title>Identification of risk factors associated with obesity and overweight&#x02013;a machine learning overview</article-title>. <source>Sensors.</source> (<year>2020</year>) <volume>20</volume>:<fpage>2734</fpage>. <pub-id pub-id-type="doi">10.3390/s20092734</pub-id><pub-id pub-id-type="pmid">32403349</pub-id></citation></ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>B</given-names></name> <name><surname>Tawfik</surname> <given-names>H</given-names></name></person-group>. <article-title>Machine learning approach for the early prediction of the risk of overweight and obesity in young people</article-title>. <source>Comput Sci ICCS 2020.</source> (<year>2020</year>). <volume>12140</volume>:<fpage>523</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-50423-6_39</pub-id></citation></ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Colmenarejo</surname> <given-names>G</given-names></name></person-group>. <article-title>Machine learning models to predict childhood and adolescent obesity: a review</article-title>. <source>Nutrients.</source> (<year>2020</year>) <volume>12</volume>:<fpage>2466</fpage>. <pub-id pub-id-type="doi">10.3390/nu12082466</pub-id><pub-id pub-id-type="pmid">32824342</pub-id></citation></ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>DeGregory</surname> <given-names>KW</given-names></name> <name><surname>Kuiper</surname> <given-names>P</given-names></name> <name><surname>DeSilvio</surname> <given-names>T</given-names></name> <name><surname>Pleuss</surname> <given-names>JD</given-names></name> <name><surname>Miller</surname> <given-names>R</given-names></name> <name><surname>Roginski</surname> <given-names>JW</given-names></name> <etal/></person-group>. <article-title>A review of machine learning in obesity</article-title>. <source>Obes Rev.</source> (<year>2018</year>) <volume>19</volume>:<fpage>668</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1111/obr.12667</pub-id></citation></ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wickham</surname> <given-names>H</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>R</given-names></name> <name><surname>Henry</surname> <given-names>L</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>K</given-names></name></person-group>. <source>dplyr: A Grammar of Data Manipulation</source>. R package version 0.7.6 (<year>2018</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://cran.r-project.org/package=dplyr">https://cran.r-project.org/package=dplyr</ext-link></citation></ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blum</surname> <given-names>AL</given-names></name> <name><surname>Langley</surname> <given-names>P</given-names></name></person-group>. <article-title>Selection of relevant features and examples in machine learning</article-title>. <source>Artif Intell.</source> (<year>1997</year>) <volume>97</volume>:<fpage>245</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/S0004-3702(97)00063-5</pub-id></citation></ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="book"><person-group person-group-type="author"><collab>R Core Team</collab></person-group>. <source>R: A Language and Environment for Statistical Computing</source>. <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Foundation for Statistical Computing</publisher-name> (<year>2020</year>).</citation></ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chawla</surname> <given-names>NV</given-names></name> <name><surname>Bowyer</surname> <given-names>KW</given-names></name> <name><surname>Hall</surname> <given-names>LO</given-names></name> <name><surname>Kegelmeyer</surname> <given-names>WP</given-names></name></person-group>. <article-title>SMOTE: synthetic minority over-sampling technique</article-title>. <source>J Artif Intell Res.</source> (<year>2002</year>) <volume>16</volume>:<fpage>321</fpage>&#x02013;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.1613/jair.953</pub-id><pub-id pub-id-type="pmid">24088532</pub-id></citation></ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chawla</surname> <given-names>NV</given-names></name></person-group>. <article-title>Data mining for imbalanced datasets: an overview. In: Maimon O, Rokach L, editors</article-title>. <source>Data Mining and Knowledge Discovery Handbook</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2005</year>). p. <fpage>853</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1007/0-387-25465-X_40</pub-id></citation></ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blagus</surname> <given-names>R</given-names></name> <name><surname>Lusa</surname> <given-names>L</given-names></name></person-group>. <article-title>Joint use of over- and under-sampling techniques and cross-validation for the development and assessment of prediction models</article-title>. <source>BMC Bioinformatics.</source> (<year>2015</year>) <volume>16</volume>:<fpage>363</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-015-0784-9</pub-id><pub-id pub-id-type="pmid">26537827</pub-id></citation></ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alghamdi</surname> <given-names>M</given-names></name> <name><surname>Al-Mallah</surname> <given-names>M</given-names></name> <name><surname>Keteyian</surname> <given-names>S</given-names></name> <name><surname>Brawner</surname> <given-names>C</given-names></name> <name><surname>Ehrman</surname> <given-names>J</given-names></name> <name><surname>Sakr</surname> <given-names>S</given-names></name></person-group>. <article-title>Predicting diabetes mellitus using SMOTE and ensemble machine learning approach: the Henry Ford ExercIse Testing (FIT) project</article-title>. <source>PLoS ONE.</source> (<year>2017</year>) <volume>12</volume>:<fpage>e0179805</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0179805</pub-id><pub-id pub-id-type="pmid">28738059</pub-id></citation></ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cost</surname> <given-names>S</given-names></name> <name><surname>Salzberg</surname> <given-names>S</given-names></name></person-group>. <article-title>A weighted nearest neighbor algorithm for learning with symbolic features</article-title>. <source>Mach Learn.</source> (<year>1993</year>) <volume>10</volume>:<fpage>57</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1023/A:1022664626993</pub-id></citation></ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T</given-names></name> <name><surname>Tibshirani</surname> <given-names>R</given-names></name> <name><surname>Friedman</surname> <given-names>J</given-names></name></person-group>. <source>The Elements of Statistical Learning: Data Mining, Inference, and Prediction</source>. <publisher-name>Springer-Verlag</publisher-name> (<year>2009</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://web.stanford.edu/&#x0007E;hastie/ElemStatLearn/">https://web.stanford.edu/&#x0007E;hastie/ElemStatLearn/</ext-link></citation></ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Bishop</surname> <given-names>C</given-names></name></person-group>. <source>Pattern Recognition and Machine Learning</source>. <publisher-name>Springer-Verlag New York</publisher-name> (<year>2006</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.springer.com/gp/book/9780387310732">https://www.springer.com/gp/book/9780387310732</ext-link></citation></ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L</given-names></name> <name><surname>Friedman</surname> <given-names>JH</given-names></name> <name><surname>Olshen</surname> <given-names>RA</given-names></name> <name><surname>Stone</surname> <given-names>CJ</given-names></name></person-group>. <source>Classification and Regression Trees</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>Chapman &#x00026; Hall; CRC</publisher-name> (<year>1984</year>).</citation></ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>J</given-names></name> <name><surname>Kamber</surname> <given-names>M</given-names></name> <name><surname>Pei</surname> <given-names>J</given-names></name></person-group>. <source>Data mining: Concepts and Techniques</source>, <edition>3rd ed</edition>. <publisher-name>Morgan Kaufmann Publishers</publisher-name> (<year>2012</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="http://myweb.sabanciuniv.edu/rdehkharghani/files/2016/02/The-Morgan-Kaufmann-Series-in-Data-Management-Systems-Jiawei-Han-Micheline-Kamber-Jian-Pei-Data-Mining.-Concepts-and-Techniques-3rd-Edition-Morgan-Kaufmann-2011.pdf">http://myweb.sabanciuniv.edu/rdehkharghani/files/2016/02/The-Morgan-Kaufmann-Series-in-Data-Management-Systems-Jiawei-Han-Micheline-Kamber-Jian-Pei-Data-Mining.-Concepts-and-Techniques-3rd-Edition-Morgan-Kaufmann-2011.pdf</ext-link></citation></ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Refaeilzadeh</surname> <given-names>P</given-names></name> <name><surname>Tang</surname> <given-names>L</given-names></name> <name><surname>Liu</surname> <given-names>H</given-names></name></person-group>. <article-title>Cross-validation. In: LIU L, &#x000D6;ZSU MT, editors</article-title>. <source>Encyclopedia of Database Systems</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2009</year>). p. <fpage>24</fpage>. <pub-id pub-id-type="doi">10.1007/978-0-387-39940-9</pub-id></citation></ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>B</given-names></name> <name><surname>Fang</surname> <given-names>L</given-names></name> <name><surname>Liu</surname> <given-names>F</given-names></name> <name><surname>Wang</surname> <given-names>X</given-names></name> <name><surname>Chen</surname> <given-names>J</given-names></name> <name><surname>Chou</surname> <given-names>KC</given-names></name></person-group>. <article-title>Identification of real microRNA precursors with a pseudo structure status composition approach</article-title>. <source>PLoS ONE.</source> (<year>2015</year>) <volume>10</volume>:<fpage>e0121501</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0121501</pub-id><pub-id pub-id-type="pmid">25821974</pub-id></citation></ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nguyen</surname> <given-names>GH</given-names></name> <name><surname>Bouzerdoum</surname> <given-names>A</given-names></name> <name><surname>Phung</surname> <given-names>SL</given-names></name></person-group>. <source>Learning Pattern Classification Tasks with Imbalanced Data Sets</source>. <publisher-loc>London</publisher-loc>: <publisher-name>IntechOpen</publisher-name> (<year>2009</year>). <pub-id pub-id-type="doi">10.5772/7544</pub-id></citation></ref>
<ref id="B38">
<label>38.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Landis</surname> <given-names>J</given-names></name> <name><surname>Koch</surname> <given-names>G</given-names></name></person-group>. <article-title>The measurement of observer agreement for categorical data</article-title>. <source>Biometrics.</source> (<year>1977</year>) <volume>33</volume>:<fpage>159</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.2307/2529310</pub-id><pub-id pub-id-type="pmid">843571</pub-id></citation></ref>
<ref id="B39">
<label>39.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Brefeld</surname> <given-names>U</given-names></name> <name><surname>Scheffer</surname> <given-names>T</given-names></name></person-group>. <article-title>AUC maximizing support vector learning. In: Ferri C, Lachiche N, Macskassy S, Rakotomamonjy A, editors</article-title>. <source>Proceedings of the 2nd Workshop on ROC Analysis in Machine Learning (ROCML 2005).</source> (<year>2005</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.59.7864&#x00026;rep=rep1&#x00026;type=pdf">https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.59.7864&#x00026;rep=rep1&#x00026;type=pdf</ext-link></citation></ref>
<ref id="B40">
<label>40.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Emery</surname> <given-names>CF</given-names></name> <name><surname>Olson</surname> <given-names>KL</given-names></name> <name><surname>Lee</surname> <given-names>VS</given-names></name> <name><surname>Habash</surname> <given-names>DL</given-names></name> <name><surname>Nasar</surname> <given-names>JL</given-names></name> <name><surname>Bodine</surname> <given-names>A</given-names></name></person-group>. <article-title>Home environment and psychosocial predictors of obesity status among community-residing men and women</article-title>. <source>Int J Obes.</source> (<year>2015</year>) <volume>39</volume>:<fpage>1401</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1038/ijo.2015.70</pub-id><pub-id pub-id-type="pmid">25916909</pub-id></citation></ref>
<ref id="B41">
<label>41.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sinha</surname> <given-names>R</given-names></name> <name><surname>Jastreboff</surname> <given-names>AM</given-names></name></person-group>. <article-title>Stress as a common risk factor for obesity and addiction</article-title>. <source>Biol Psychiatry.</source> (<year>2013</year>) <volume>73</volume>:<fpage>827</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1016/j.biopsych.2013.01.032</pub-id><pub-id pub-id-type="pmid">23541000</pub-id></citation></ref>
<ref id="B42">
<label>42.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koski</surname> <given-names>M</given-names></name> <name><surname>Naukkarinen</surname> <given-names>H</given-names></name></person-group>. <article-title>The relationship between stress and severe obesity: a case-control study</article-title>. <source>Biomed Hub.</source> (<year>2017</year>) <volume>2</volume>:<fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1159/000458771</pub-id><pub-id pub-id-type="pmid">31988895</pub-id></citation></ref>
<ref id="B43">
<label>43.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>H</given-names></name> <name><surname>Chris</surname> <given-names>K</given-names></name> <name><surname>Junxiu</surname> <given-names>L</given-names></name> <name><surname>Yujin</surname> <given-names>L</given-names></name> <name><surname>Jonathan</surname> <given-names>PS</given-names></name> <name><surname>Brendan</surname> <given-names>C</given-names></name></person-group>. <article-title>Cost-effectiveness of the US food and drug administration added sugar labeling policy for improving diet and health</article-title>. <source>Circulation.</source> (<year>2019</year>) <volume>139</volume>:<fpage>2613</fpage>&#x02013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1161/CIRCULATIONAHA.118.036751</pub-id><pub-id pub-id-type="pmid">30982338</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This research was funded by the Ministry of Research, Technology/National Research, and Innovation Agency of Indonesia through Grant PDUPT Hasanuddin University in 2020 with the number 1516/UN4.22/PT.01.03/2020.</p>
</fn>
</fn-group>
</back>
</article> 
