<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Big Data</journal-id>
<journal-title>Frontiers in Big Data</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Big Data</abbrev-journal-title>
<issn pub-type="epub">2624-909X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdata.2024.1402384</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Big Data</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Random kernel k-nearest neighbors regression</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes" equal-contrib="yes">
<name><surname>Srisuradetchai</surname> <given-names>Patchanok</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2689528/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" equal-contrib="yes">
<name><surname>Suksrikran</surname> <given-names>Korn</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2689753/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/resources/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>Department of Mathematics and Statistics, Thammasat University</institution>, <addr-line>Pathum Thani</addr-line>, <country>Thailand</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Dongpo Xu, Northeast Normal University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Debo Cheng, University of South Australia, Australia</p>
<p>Alladoumbaye Ngueilbaye, Shenzhen University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Patchanok Srisuradetchai <email>patchanok&#x00040;mathstat.sci.tu.ac.th</email></corresp>
<fn fn-type="equal" id="fn001"><p>&#x02020;These authors have contributed equally to this work and share first authorship</p></fn></author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>07</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>7</volume>
<elocation-id>1402384</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>03</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>05</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2024 Srisuradetchai and Suksrikran.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Srisuradetchai and Suksrikran</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The k-nearest neighbors (KNN) regression method, known for its nonparametric nature, is highly valued for its simplicity and its effectiveness in handling complex structured data, particularly in big data contexts. However, this method is susceptible to overfitting and fit discontinuity, which present significant challenges. This paper introduces the random kernel k-nearest neighbors (RK-KNN) regression as a novel approach that is well-suited for big data applications. It integrates kernel smoothing with bootstrap sampling to enhance prediction accuracy and the robustness of the model. This method aggregates multiple predictions using random sampling from the training dataset and selects subsets of input variables for kernel KNN (K-KNN). A comprehensive evaluation of RK-KNN on 15 diverse datasets, employing various kernel functions including Gaussian and Epanechnikov, demonstrates its superior performance. When compared to standard KNN and the random KNN (R-KNN) models, it significantly reduces the root mean square error (RMSE) and mean absolute error, as well as improving R-squared values. The RK-KNN variant that employs a specific kernel function yielding the lowest RMSE will be benchmarked against state-of-the-art methods, including support vector regression, artificial neural networks, and random forests.</p></abstract>
<kwd-group>
<kwd>bootstrapping</kwd>
<kwd>feature selection</kwd>
<kwd>k-nearest neighbors regression</kwd>
<kwd>kernel k-nearest neighbors</kwd>
<kwd>state-of-the-art (SOTA)</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="3"/>
<equation-count count="14"/>
<ref-count count="56"/>
<page-count count="14"/>
<word-count count="7553"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Machine Learning and Artificial Intelligence</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>The recent increase in machine learning research has highlighted the significance of ensemble techniques and regression models, which have demonstrated enhanced predictive capabilities. This trend is observable across a wide range of domains and use cases, as evidenced by the current research landscape. Li et al. (<xref ref-type="bibr" rid="B31">2023</xref>) conducted a comprehensive study in the field of agriculture, analyzing meteorological patterns and soybean yield statistics from various counties and weather stations within China&#x00027;s primary soybean cultivation regions. They utilized a stacking ensemble framework to construct a predictive model for soybean yield estimation, employing algorithms such as k-nearest neighbor (KNN), random forest (RF), and support vector regression (SVR). Jiang et al. (<xref ref-type="bibr" rid="B29">2023</xref>) developed a stacking ensemble model that integrates RF, KNN regression, gradient boosting regression (GBR), and a meta-learner, specifically linear regression (LR), to predict greenhouse gas emissions from irrigated rice farms. Bian and Huang (<xref ref-type="bibr" rid="B8">2024</xref>) developed a novel fuzzy modeling approach using an enhanced evidence theory integrated with KNN for dynamic and accurate air pollution estimation.</p>
<p>In the energy sector, El-Kenawy et al. (<xref ref-type="bibr" rid="B16">2021</xref>) introduced an improved ensemble model for predicting solar radiation levels. This model operates in two stages: data preparation and ensemble training. It is enhanced through KNN regression, and its effectiveness is evaluated using a dataset from Kaggle. Compared to existing benchmarks, the unique advantages of this model are evident. In a related study, Chung et al. (<xref ref-type="bibr" rid="B12">2019</xref>) explored various machine learning techniques to predict charging patterns, analyzing factors such as duration and energy consumption from historical data. They developed the Ensemble Predicting Algorithm (EPA) by integrating diverse techniques to enhance predictive accuracy. Sharma and Lakshmi (<xref ref-type="bibr" rid="B40">2023</xref>) proposed a model that initially segments the values of the target variable into multiple categories. Then, a unified KNN model, which merges both weighted attribute KNN and distance-weighted KNN, is applied. The weighting for each attribute is determined through information gain. This model is employed to predict the target variable&#x00027;s value for each test instance. Their primary aim was to use various KNN-focused models to increase the accuracy of air pollutant level predictions. Cheng et al. (<xref ref-type="bibr" rid="B11">2014</xref>) introduced a novel KNN methodology based on sparse learning, designed to address the limitations of previous KNN approaches, such as using a fixed <italic>k</italic> value for all test instances and overlooking sample correlations. This strategy adjusts test samples and uses training samples to identify the optimal <italic>k</italic> value for each instance. Subsequently, the refined KNN method, with the optimized <italic>k</italic> value, is applied to various tasks, including categorization, regression, and imputation of missing values.</p>
<p>Song and Choi (<xref ref-type="bibr" rid="B41">2023</xref>) introduced innovative integrated models within the finance industry, aimed at forecasting both short-term and long-term closing prices of major stock market indices: DAX, DOW, and S&#x00026;P500. They proposed an enhancement involving the calculation of the mean of the highest and lowest prices of these indices to improve accuracy. In a separate domain, Dimopoulos et al. (<xref ref-type="bibr" rid="B15">2018</xref>) conducted a comparative study on the effectiveness of machine learning vs. traditional risk ratings in estimating the risk of cardiovascular disease.</p>
<p>KNN regressions have also been discovered for environmental research. Jafar et al. (<xref ref-type="bibr" rid="B28">2023</xref>) conducted a study to compare the effectiveness of multiple linear regression with 19 different machine learning techniques. These algorithms included regression, decision trees, and boosting mechanisms. The analyzed models included LR, least angle regression (LAR), Bayesian ridge chain (BR), ridge regression (Ridge), KNN, extra tree regression, and the notably robust XGBoost. In a related effort, Srisuradetchai and Panichkitkosolkul (<xref ref-type="bibr" rid="B44">2022</xref>) employed an ensemble machine learning approach that incorporated KNN, MLR, RF, SVR, and other algorithms to predict PM2.5 levels in Bangkok. This ensemble learning method was further applied by Srisuradetchai et al. (<xref ref-type="bibr" rid="B45">2023</xref>) to forecast daily new confirmations of COVID-19 cases.</p>
<p>KNN regression has been enhanced through its combination with other algorithms. Ghavami et al. (<xref ref-type="bibr" rid="B21">2023</xref>) introduced an innovative ensemble prediction technique named COA-KNN, which integrates the Coyote optimization algorithm (COA) with KNN to enhance the accuracy of fatigue and rutting predictions in reclaimed asphalt pavement mixtures. When compared to established prediction models, including RF, GB, decision tree regression (DTR), and MLR, COA-KNN demonstrated superior performance across various metrics. Similarly, Song et al. (<xref ref-type="bibr" rid="B42">2018</xref>) developed a potent regression learning approach termed the distance-weighted KNN algorithm. This algorithm aims to elucidate the nonlinear relationships between input structural parameters and resultant motor performances.</p>
<p>In the expanding field of KNN classification, particularly in the context of big data, Bermejo and Cabestany (<xref ref-type="bibr" rid="B7">2000</xref>) pioneered an adaptive soft KNN classifier that estimates posterior class probabilities, showcasing improved handwritten character recognition. Meanwhile, Deng et al. (<xref ref-type="bibr" rid="B14">2016</xref>) optimized KNN classification for large datasets using a hybrid approach of k-means clustering and KNN classification. Ingram and Munzner (<xref ref-type="bibr" rid="B27">2015</xref>) proposed the Q-SNE algorithm, a dimensionality reduction technique tailored for document data, significantly enhancing the layout quality of large document collections. Similarly, Pramanik et al. (<xref ref-type="bibr" rid="B34">2021</xref>) reviewed the applications and challenges of big data classification, discussing the imperative of systematic data processing for knowledge discovery and decision-making. Saadatfar et al. (<xref ref-type="bibr" rid="B38">2020</xref>) addressed the computational challenges of applying KNN to big data by clustering data into smaller, manageable partitions. Abdalla and Amer (<xref ref-type="bibr" rid="B1">2022</xref>) introduced NCP-KNN, a variation that reduces search complexity and excels in high-dimensional classification, promising efficiency for large datasets. Finally, Ukey et al. (<xref ref-type="bibr" rid="B51">2023</xref>) delivered a comprehensive survey on exact KNN queries over high-dimensional data.</p>
<p>Kernel functions are employed in KNN, as demonstrated by Zheng and Cao (<xref ref-type="bibr" rid="B56">2008</xref>), who explored the use of kernel functions in KNN for Holter waveform classification. Enriquez et al. (<xref ref-type="bibr" rid="B17">2019</xref>) devised and examined a methodology for identifying faults in power transformers using a KNN classifier with a weighted classification distance. Rubio et al. (<xref ref-type="bibr" rid="B37">2009</xref>) introduced a parallel implementation of the sequential kernel-weighted KNN algorithm in Matlab, specifically designed for cluster platforms. Ali et al. (<xref ref-type="bibr" rid="B3">2020</xref>) developed a group model utilizing the KNN algorithm, employing samples and random features to generate predictions by pooling various models. Bay (<xref ref-type="bibr" rid="B5">1999</xref>) also explored a similar concept, aiming to enhance nearest neighbor classifiers through the utilization of a combination of multiple models, each emphasizing random features. However, these studies, including research conducted by Garc&#x000ED;a-Pedrajas and Ortiz-Boyer (<xref ref-type="bibr" rid="B20">2009</xref>), Steele (<xref ref-type="bibr" rid="B46">2009</xref>), and Li et al. (<xref ref-type="bibr" rid="B32">2014</xref>), primarily aimed to enhance classifiers by utilizing a random subset of input variables without considering the utilization of kernel functions. For the KNN time series model, Srisuradetchai (<xref ref-type="bibr" rid="B43">2023</xref>) proposed a new approach for interval forecasting that combines the KNN time series model with bootstrapping.</p>
<p>This study enhances random KNN regression by incorporating kernel methods. While traditional random KNN regression is effective with various data types, it may not detect intricate patterns that are crucial for accurate predictions. The method introduced here, named Random Kernel KNN regression (RK-KNN), employs random feature selection, bootstraps data samples, and applies kernel functions to weight distances. This paper evaluates RK-KNN across 15 datasets and compares its performance with state-of-the-art methods, including random forest, support vector regression, and artificial neural networks.</p></sec>
<sec id="s2">
<title>2 Theoretical background</title>
<sec>
<title>2.1 Kernel functions</title>
<p>Kernel functions are used to weigh the contributions of each point based on its distance from the query point. While traditional KNN uses uniform weights, kernel functions allow these weights to vary, often improving performance. Here are some widely used kernels that can be applied in KNN regression (Sch&#x000F6;lkopf and Smola, <xref ref-type="bibr" rid="B39">2001</xref>; Tsybakov, <xref ref-type="bibr" rid="B50">2009</xref>; Beitollahi et al., <xref ref-type="bibr" rid="B6">2022</xref>):</p>
<list list-type="bullet">
<list-item><p>Gaussian (Radial Basis Function) kernel:</p></list-item>
</list>
<p>Perhaps the most popular kernel, the Gaussian kernel, has a bell-shaped curve and can assign weights to points in the input space based on their distance from the query point, with this influence rapidly declining as the distance increases, as shown in <xref ref-type="disp-formula" rid="E1">Equation (1)</xref>.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C3;<sup>2</sup> is the standard deviation (bandwidth).</p>
<list list-type="bullet">
<list-item><p>Epanechnikov kernel:</p></list-item></list>
<p>This kernel is parabolic and is often used because of its computational efficiency. It assigns more weight to nearby points than to points further away, but unlike the Gaussian kernel, it becomes zero beyond a certain distance, as defined in <xref ref-type="disp-formula" rid="E2">Equation (2)</xref>.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and <italic>K</italic>(<italic>x, x</italic>&#x02032;) &#x0003D; 0 otherwise, where <italic>h</italic> is the bandwidth.</p>
<list list-type="bullet">
<list-item><p>Uniform kernel:</p></list-item></list>
<p>The uniform kernel gives equal weight to all points within a certain range of the query point and no weight to points outside this range. It is the simplest form of kernel and is equivalent to the traditional KNN method when used with a fixed radius, as expressed in <xref ref-type="disp-formula" rid="E3">Equation (3)</xref>.</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>h</mml:mi></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and <italic>K</italic>(<italic>x, x</italic>&#x02032;) &#x0003D; 0 otherwise, where <italic>h</italic> is the bandwidth.</p>
<list list-type="bullet">
<list-item><p>Triangular kernel:</p></list-item></list>
<p>The triangular kernel assigns weights that decrease linearly with distance from the query point. It is zero beyond the kernel&#x00027;s bandwidth, as shown in <xref ref-type="disp-formula" rid="E4">Equation (4)</xref>.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and <italic>K</italic>(<italic>x, x</italic>&#x02032;) &#x0003D; 0 otherwise, where <italic>h</italic> is the bandwidth.</p>
<list list-type="bullet">
<list-item><p>Quartic (Biweight) kernel:</p></list-item></list>
<p>This kernel is similar to the Epanechnikov kernel but assigns weight with a smooth, bell-shaped curve, which reaches zero at the kernel&#x00027;s bandwidth, as defined in <xref ref-type="disp-formula" rid="E5">Equation (5)</xref>.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>15</mml:mn></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and <italic>K</italic>(<italic>x, x</italic>&#x02032;) &#x0003D; 0 otherwise, where <italic>h</italic> is the bandwidth.</p>
<list list-type="bullet">
<list-item><p>Tricube kernel:</p></list-item></list>
<p>The tricube kernel is a higher-order kernel with compact support, meaning it assigns a weight of zero to any point outside a certain range of the query point. It is smoother and has heavier tails than the quartic kernel, according to <xref ref-type="disp-formula" rid="E6">Equation (6)</xref>.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>70</mml:mn></mml:mrow><mml:mrow><mml:mn>81</mml:mn></mml:mrow></mml:mfrac><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and <italic>K</italic>(<italic>x, x</italic>&#x02032;) &#x0003D; 0 otherwise, where <italic>h</italic> is the bandwidth.</p>
<p>All kernel functions, as plotted in <xref ref-type="fig" rid="F1">Figure 1</xref>, have a bandwidth of one. The Gaussian kernel is depicted as a smooth curve peaking at the center. The Epanechnikov kernel displays a parabolic shape that cuts off at the bandwidth&#x00027;s edge. The uniform kernel provides equal weight within a fixed bandwidth and falls to zero beyond it. The triangular kernel&#x00027;s weight decreases linearly with distance, ending at the bandwidth limit. The quartic kernel features a bell shape that smoothly tapers to zero, while the tricube kernel has a more pronounced peak with a faster decline.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Comparison of different kernel functions, all centered at zero and using a bandwidth of one.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0001.tif"/>
</fig>
</sec>
<sec>
<title>2.2 K-nearest neighbor regression</title>
<p>KNN regression is a type of non-parametric method used for predicting the continuous outcome of a new data point based on the outcomes of its nearest neighbors in the feature space. It does not make any assumptions about the underlying data distribution and is particularly useful when dealing with complex data structures (Hastie et al., <xref ref-type="bibr" rid="B23">2009</xref>). Given a dataset with <italic>n</italic> points, (<italic>x</italic><sub>1</sub>, <italic>y</italic><sub>1</sub>), (<italic>x</italic><sub>2</sub>, <italic>y</italic><sub>2</sub>), ..., (<italic>x</italic><sub><italic>n</italic></sub>, <italic>y</italic><sub><italic>n</italic></sub>), where each <italic>x</italic><sub><italic>i</italic></sub> represents a vector of features and each <italic>y</italic><sub><italic>i</italic></sub> represents the corresponding continuous outcome, KNN regression predicts the outcome &#x00177; for a new data point <italic>x</italic> based on the outcomes of its <italic>k</italic> nearest neighbors in the feature space. The mathematical formulation of KNN regression includes (Altman, <xref ref-type="bibr" rid="B4">1992</xref>):</p>
<list list-type="bullet">
<list-item><p>Distance metric: the first step in KNN regression is to determine the &#x0201C;closeness&#x0201D; of data points in the feature space, which requires a distance metric. The most common choice is the Euclidean distance, though other metrics like Manhattan or Minkowski can also be used. The distance <italic>d</italic> between two data points <italic>x</italic> and <italic>x</italic><sub><italic>i</italic></sub> is calculated by using <xref ref-type="disp-formula" rid="E7">Equation (7)</xref> for Euclidean distance.</p>
<p><disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></list-item>
<list-item><p>Finding neighbors: for a given data point <italic>x</italic>, find the <italic>k</italic> points in the dataset that are closest to <italic>x</italic> based on the distance metric. These points are termed KNN.</p></list-item>
<list-item><p>Prediction: The predicted outcome &#x00177; is calculated as the average of the outcomes of the k-nearest neighbors. Mathematically, this can be represented as:</p>
<p><disp-formula id="E8"><label>(8)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x00177;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In <xref ref-type="disp-formula" rid="E8">Equation (8)</xref>, <italic>N</italic><sub><italic>k</italic></sub> contains the indices of the <italic>k</italic> closest (in <italic>l</italic><sub>2</sub> distance) of <italic>x</italic><sub>1</sub>, ..., <italic>x</italic><sub><italic>n</italic></sub> to <italic>x</italic>.</p>
</list-item>
</list>
<p><xref ref-type="fig" rid="F2">Figure 2</xref> illustrates the example of the KNN regression with <italic>k</italic> = 10, where the dataset was synthetically generated from model <italic>Y</italic> &#x0003D; sin(<italic>x</italic>)&#x0002B; sin(2<italic>x</italic>)&#x0002B;&#x003B5;. It can be observed that KNN regression makes no assumptions about linearity and fits the data well. The predictions for new data points are based on the average outcomes of the 9 nearest points from the training data.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>KNN regression model with <italic>k</italic> = 9 applied to synthetically generated data.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0002.tif"/>
</fig>
</sec>
<sec>
<title>2.3 Kernel k-nearest neighbor regression</title>
<p>Kernel <italic>k</italic>-Nearest Neighbor (K-KNN) regression extends the conventional KNN regression algorithm, an instance-based learning method, by incorporating kernel functions. This integration allows the algorithm to weigh the contributions of each point&#x00027;s neighbors based on their distance, effectively smoothing out predictions and improving the model&#x00027;s ability to handle complex, non-linear relationships (Tan et al., <xref ref-type="bibr" rid="B47">2020</xref>; Yao et al., <xref ref-type="bibr" rid="B55">2021</xref>). Given a dataset with <italic>n</italic> points (<italic>x</italic><sub>1</sub>, <italic>y</italic><sub>1</sub>), (<italic>x</italic><sub>2</sub>, <italic>y</italic><sub>2</sub>), ..., (<italic>x</italic><sub><italic>n</italic></sub>, <italic>y</italic><sub><italic>n</italic></sub>), where each <italic>x</italic><sub><italic>i</italic></sub> represents a vector of features and each <italic>y</italic><sub><italic>i</italic></sub> represents the corresponding continuous outcome, the prediction &#x00177; for a new data point <italic>x</italic> is calculated not just by averaging the outcomes of its <italic>k</italic> but by taking a weighted average, where the weights are determined by a kernel function based on the distance between <italic>x</italic> and each <italic>x</italic><sub><italic>i</italic></sub>.</p>
<p>The kernel function <italic>K</italic> in <xref ref-type="disp-formula" rid="E9">Equation (9)</xref> applied in this context is a symmetric function that satisfies certain mathematical conditions (like positivity and integrability) with the general form <italic>K</italic>:&#x0211D;<sup><italic>d</italic></sup> &#x02192; &#x0211D;, where <italic>d</italic> is the dimension of the input space. The kernel function <italic>K</italic> must satisfy</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x0222B;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>d</mml:mi><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x0222B;</mml:mo><mml:mi>t</mml:mi><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>d</mml:mi><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mn>0</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:mo>&#x0222B;</mml:mo><mml:msup><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>d</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The choice of kernel function can significantly influence the regression outcome, as different kernels impose different structures on the data (Hofmann et al., <xref ref-type="bibr" rid="B26">2008</xref>). The prediction &#x00177; in K-KNN regression is then given by:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x00177;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="false"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="false"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In <xref ref-type="disp-formula" rid="E10">Equation (10)</xref> <italic>K</italic>(<italic>x, x</italic><sub><italic>i</italic></sub>) is the kernel function evaluating the similarity (or smoothness) between the target point <italic>x</italic> and each neighbor <italic>x</italic><sub><italic>i</italic></sub>.</p>
</sec>
<sec>
<title>2.4 Cross-validation for optimal parameters</title>
<p>It is imperative to determine the optimal <italic>k</italic> for the neighbors and the best-suited bandwidth for the kernel function in the context of each bootstrap sample. This step ensures that the model is not just fitted to the training data but also generalizes well to unseen data.</p>
<p>Utilizing &#x003BD;&#x02212;fold cross-validation, the original training set is randomly partitioned into &#x003BD; equal-sized subsamples. Of the &#x003BD; subsamples, a single subsample is retained as the validation data for testing the model, and the remaining &#x003BD;&#x02212;1 subsamples are used as training data. The cross-validation process is then repeated &#x003BD; times (the folds), with each of the &#x003BD; subsamples used exactly once as the validation data. For each fold and each candidate combination of parameters (specific <italic>k</italic> and bandwidth), the model is trained, and the prediction error (e.g., RMSE) on the validation fold is computed. The average error across all &#x003BD; folds is then calculated for each combination (Wong and Yang, <xref ref-type="bibr" rid="B53">2017</xref>; Wong and Yeh, <xref ref-type="bibr" rid="B54">2020</xref>).</p></sec>
</sec>
<sec id="s3">
<title>3 Proposed method</title>
<p>Combining bootstrap sampling, choosing features at random, and using kernel methods in a KNN model is employed to make standard KNN better at prediction. Given a training dataset <italic>LD</italic>(<italic>X</italic>; <italic>Y</italic>), where <italic>X</italic> is a <italic>p</italic>-dimensional feature matrix with <italic>n</italic> observations and <italic>Y</italic> is the corresponding response variable, the objective is to predict the response &#x00177; for a new observation <italic>x</italic><sub>0</sub> in the test dataset. Note that in KNN regression, it is essential to preprocess all predictors to ensure they are unitless. The step for random kernel KNN (RK-KNN) regression is as follows:</p>
<list list-type="simple">
<list-item><p>1) <italic>Bootstrap Sampling for KNN</italic></p></list-item>
</list>
<p>Bootstrap sampling is integral to ensemble methodologies, particularly bagging. It involves generating <italic>B</italic> unique datasets from the original training data, <italic>D</italic>, each termed <italic>D</italic><sub><italic>b</italic></sub> (where <italic>b</italic> &#x0003D; 1, 2, &#x02026;, <italic>B</italic>) by sampling <italic>n</italic> observations with replacement. In mathematical notation,</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In this study, <italic>B</italic> is set to 1,000.</p>
<list list-type="simple">
<list-item><p>3) <italic>Random Feature Selection</italic></p></list-item>
</list>
<p>Incorporating a feature randomness aspect akin to Random Forests, each bootstrap sample <italic>D</italic><sub><italic>b</italic></sub>, as shown in <xref ref-type="disp-formula" rid="E11">Equation (11)</xref>, undergoes a feature selection process where only a random subset of <italic>d</italic> features (where <italic>d</italic>&#x0003C;<italic>p</italic>) is considered for model training. During the training phase for each <italic>D</italic><sub><italic>b</italic></sub>, the algorithm does not utilize the full feature set. Instead, it randomly selects a subset, contributing to model diversity within the ensemble (Breiman, <xref ref-type="bibr" rid="B10">2001</xref>). In this study, <italic>d</italic> is set to <italic>p</italic>/2, <italic>p</italic>/5, and <inline-formula><mml:math id="M13"><mml:msqrt><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula>, and the best <italic>d</italic> is determined from one that yields the lowest RMSE or MAE or highest R<sup>2</sup>.</p>
<list list-type="simple">
<list-item><p>3) <italic>Kernel Enhancement in KNN</italic></p></list-item>
</list>
<p>We add the Gaussian, Epanechnikov, uniform, triangular, quartic, and tricube kernels to a standard KNN in this paper. Within KNN, kernel functions can adjust neighbor contributions, giving more weight to nearer neighbors. Suppose <italic>N</italic><sub><italic>k</italic></sub>(<italic>x</italic><sub>0</sub>) denotes the set of <italic>k</italic> nearest neighbors to a query point <italic>x</italic><sub>0</sub>, determined using a subset of <italic>d</italic> features. The kernel-weighted response estimate is given by:</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x00177;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="false"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="false"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In <xref ref-type="disp-formula" rid="E12">Equation (12)</xref>, <italic>K</italic>(<italic>x</italic><sub>0</sub>, <italic>x</italic><sub><italic>i</italic></sub>) is the kernel function evaluating the closeness of points <italic>x</italic><sub>0</sub> and <italic>x</italic><sub><italic>i</italic></sub>, and <italic>y</italic><sub><italic>i</italic></sub> are the response values of the neighbors. Note that all <italic>x</italic><sub><italic>i</italic></sub> are needed to be rescaled to be in [0, 1]. This scaling not only helps remove unit dominance but also easily helps determine the bandwidth value of the kernel functions.</p>
<list list-type="simple">
<list-item><p>4) <italic>Determining Optimal k and Bandwidth</italic></p></list-item>
</list>
<p>The optimal <italic>k</italic> and bandwidth parameters are those that minimize the average prediction error estimated through a 5-fold cross-validation. Let&#x00027;s denote the set of candidate <italic>k</italic> values as {<italic>k</italic><sub>1</sub>, <italic>k</italic><sub>2</sub>, ..., <italic>k</italic><sub><italic>r</italic></sub>} and the set of candidate bandwidths as {<italic>h</italic><sub>1</sub>, <italic>h</italic><sub>2</sub>, ..., <italic>h</italic><sub><italic>s</italic></sub>}. The objective is to find the optimal <italic>k</italic><sub><italic>opt</italic></sub> and <italic>h</italic><sub><italic>opt</italic></sub> that yield the lowest estimated prediction error:</p>
<disp-formula id="E13"><label>(13)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo class="qopname">arg</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>C</mml:mi><mml:mi>V</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In <xref ref-type="disp-formula" rid="E13">Equation (13)</xref>, <italic>CV</italic>(<italic>k</italic><sub><italic>i</italic></sub>, <italic>h</italic><sub><italic>j</italic></sub>) represents the cross-validation error estimated over multiple random splits of the dataset into training and validation sets. Because all variables are scaled between 0 and 1, the distance between any two points will also fall within a bounded interval. This boundedness allows for the selection of h based on the maximum distance within the <italic>k</italic>-nearest neighbors for each query point, specifically for <italic>k</italic> &#x0003D; 2, 3, 5, 7. This method ensures that <italic>h</italic> is sufficiently large to encompass all neighbors in the calculation, thus being responsive to the local structure of the data and accommodating areas of varying density. The optimal (<italic>k</italic><sub><italic>opt</italic></sub>, <italic>h</italic><sub><italic>opt</italic></sub>) is found from the grid {<italic>k</italic><sub>1</sub>, <italic>k</italic><sub>2</sub>, ..., <italic>k</italic><sub><italic>r</italic></sub>} &#x000D7; {<italic>h</italic><sub>1</sub>, <italic>h</italic><sub>2</sub>, ..., <italic>h</italic><sub><italic>s</italic></sub>}.</p>
<list list-type="simple">
<list-item><p>5) <italic>Ensemble Prediction</italic></p></list-item>
</list>
<p>The ensemble&#x00027;s predictive power is harnessed by aggregating the individual KNN models&#x00027; outputs. If each model provides a prediction &#x00177;<sub><italic>b</italic></sub> for <italic>x</italic><sub>0</sub>. The final prediction is an aggregate statistic (e.g., mean) of these predictions, as shown in <xref ref-type="disp-formula" rid="E14">Equation (14)</xref>:</p>
<disp-formula id="E14"><label>(14)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x00177;</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mi>&#x00177;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The pseudocode below (<xref ref-type="table" rid="T4">Algorithm 1</xref>) concretizes the sequence of steps&#x02014;from bootstrap sampling to the ensemble prediction&#x02014;that collectively forge our proposed method.</p>
<table-wrap position="float" id="T4">
<label>Algorithm 1</label>
<caption><p>RK-KNN model for predicting responses.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-i0001.tif"/>
</table-wrap>
</sec>
<sec id="s4">
<title>4 Evaluation datasets and results</title>
<p>This section is dedicated to presenting the datasets used for benchmarking and the outcomes of the empirical evaluation conducted to assess the effectiveness of the RK-KNN regression approach. Additionally, state-of-the-art methods, including RF, ANN, and SVR, will be compared with the RK-KNN models.</p>
<sec>
<title>4.1 Datasets for benchmarking</title>
<p>For assessing the new approach alongside existing leading techniques, we utilize 15 distinct datasets. These collections of data are acquired from multiple publicly accessible platforms. An overview of each dataset is presented in <xref ref-type="table" rid="T1">Table 1</xref>, detailing the number of observations (<italic>n</italic>), the number of predictor variables (<italic>p</italic>), and the meaning of the response variable.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Datasets employed for model evaluation in RK-KNN regression with different kernels.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="left"><bold><italic>n</italic></bold></th>
<th valign="top" align="left"><bold><italic>p</italic></bold></th>
<th valign="top" align="left"><bold>Responses</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">D1</td>
<td valign="top" align="left">Student performance (Cortez, <xref ref-type="bibr" rid="B13">2014</xref>)</td>
<td valign="top" align="left">649</td>
<td valign="top" align="left">30</td>
<td valign="top" align="left">Math scores</td>
</tr> <tr>
<td valign="top" align="left">D2</td>
<td valign="top" align="left">Student performance (Cortez, <xref ref-type="bibr" rid="B13">2014</xref>)</td>
<td valign="top" align="left">649</td>
<td valign="top" align="left">30</td>
<td valign="top" align="left">Portuguese scores</td>
</tr> <tr>
<td valign="top" align="left">D3</td>
<td valign="top" align="left">Wisconsin Prognostic Breast Cancer (Wolberg et al., <xref ref-type="bibr" rid="B52">1995</xref>)</td>
<td valign="top" align="left">198</td>
<td valign="top" align="left">34</td>
<td valign="top" align="left">Recurrence time</td>
</tr> <tr>
<td valign="top" align="left">D4</td>
<td valign="top" align="left">Properties of poly-aromatic hydrocar-bons (PAH) (Todeschini et al., <xref ref-type="bibr" rid="B49">1995</xref>)</td>
<td valign="top" align="left">80</td>
<td valign="top" align="left">113</td>
<td valign="top" align="left">Effectiveness of the PDGFR inhibitors</td>
</tr> <tr>
<td valign="top" align="left">D5</td>
<td valign="top" align="left">Platelet derived growth factor receptor (PDGFR) (Guha and Jurs, <xref ref-type="bibr" rid="B22">2004</xref>)</td>
<td valign="top" align="left">79</td>
<td valign="top" align="left">304</td>
<td valign="top" align="left">Biological activity reported as IC50values</td>
</tr> <tr>
<td valign="top" align="left">D6</td>
<td valign="top" align="left">Triazines (Hirst et al., <xref ref-type="bibr" rid="B25">1994</xref>)</td>
<td valign="top" align="left">186</td>
<td valign="top" align="left">60</td>
<td valign="top" align="left">Inhibitory activity of triazine compounds</td>
</tr> <tr>
<td valign="top" align="left">D7</td>
<td valign="top" align="left">Phenethyl (Kubinyi, <xref ref-type="bibr" rid="B30">1993</xref>)</td>
<td valign="top" align="left">22</td>
<td valign="top" align="left">629</td>
<td valign="top" align="left">Phenethyl derivatives</td>
</tr> <tr>
<td valign="top" align="left">D8</td>
<td valign="top" align="left">Topo (Feng et al., <xref ref-type="bibr" rid="B18">2003</xref>)</td>
<td valign="top" align="left">8,885</td>
<td valign="top" align="left">267</td>
<td valign="top" align="left">Toxic effects from the compound structure</td>
</tr> <tr>
<td valign="top" align="left">D9</td>
<td valign="top" align="left">Tecator (Borggaard, <xref ref-type="bibr" rid="B9">1992</xref>; Thodberg, <xref ref-type="bibr" rid="B48">1996</xref>)</td>
<td valign="top" align="left">240</td>
<td valign="top" align="left">125</td>
<td valign="top" align="left">Fat content of a meat sample</td>
</tr> <tr>
<td valign="top" align="left">D10</td>
<td valign="top" align="left">Fric4 (Friedman, <xref ref-type="bibr" rid="B19">1999</xref>)</td>
<td valign="top" align="left">1,000</td>
<td valign="top" align="left">100</td>
<td valign="top" align="left">Artificially generated responses from the model.</td>
</tr> <tr>
<td valign="top" align="left">D11</td>
<td valign="top" align="left">HappinessRank (Helliwell et al., <xref ref-type="bibr" rid="B24">2017</xref>)</td>
<td valign="top" align="left">235</td>
<td valign="top" align="left">9</td>
<td valign="top" align="left">Happiness scores from The World Happiness Report</td>
</tr> <tr>
<td valign="top" align="left">D12</td>
<td valign="top" align="left">AutoHorse (OpenML, <xref ref-type="bibr" rid="B33">2024</xref>)</td>
<td valign="top" align="left">201</td>
<td valign="top" align="left">69</td>
<td valign="top" align="left">Price</td>
</tr> <tr>
<td valign="top" align="left">D13</td>
<td valign="top" align="left">Residential Building (Rafiei, <xref ref-type="bibr" rid="B35">2018</xref>)</td>
<td valign="top" align="left">372</td>
<td valign="top" align="left">108</td>
<td valign="top" align="left">Actual sales prices</td>
</tr> <tr>
<td valign="top" align="left">D14</td>
<td valign="top" align="left">Communities and Crime (Redmond, <xref ref-type="bibr" rid="B36">2009</xref>)</td>
<td valign="top" align="left">1,994</td>
<td valign="top" align="left">100</td>
<td valign="top" align="left">Total number of violent crimes per 100K population</td>
</tr> <tr>
<td valign="top" align="left">D15</td>
<td valign="top" align="left">Pumadyn (Alcal&#x000E1;-Fdez et al., <xref ref-type="bibr" rid="B2">2011</xref>)</td>
<td valign="top" align="left">8,192</td>
<td valign="top" align="left">32</td>
<td valign="top" align="left">Angular acceleration of one of the robot arms</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.2 Performance evaluation</title>
<p>The performance of the RK-KNN method, when compared to the standard KNN and R-KNN, across datasets D1 to D15, is summarized in <xref ref-type="table" rid="T2">Table 2</xref>. It reveals the effectiveness of the RK-KNN method in enhancing predictive accuracy. The RK-KNN method, particularly when employing specific kernels, consistently outperforms the standard KNN in terms of root mean square serror (RMSE), mean absolute error (MAE), and R-squared (R<sup>2</sup>) values.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Performance evaluation of KNN, R-KNN, and RK-KNN with various kernel functions across datasets (bold values represent the best performance).</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>RMSE</bold></th>
<th valign="top" align="left"><bold>MAE</bold></th>
<th valign="top" align="left"><bold>R<sup>2</sup></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">D1</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">4.1878</td>
<td valign="top" align="left">3.1845</td>
<td valign="top" align="left">0.2055</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">3.7979</td>
<td valign="top" align="left">2.7844</td>
<td valign="top" align="left">0.6291</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">3.7976</td>
<td valign="top" align="left">2.7844</td>
<td valign="top" align="left">0.6287</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">3.7975</td>
<td valign="top" align="left">2.7846</td>
<td valign="top" align="left">0.6288</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">3.7979</td>
<td valign="top" align="left">2.7844</td>
<td valign="top" align="left"><bold>0.6304</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left"><bold>3.7949</bold></td>
<td valign="top" align="left"><bold>2.7828</bold></td>
<td valign="top" align="left">0.6284</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">3.7971</td>
<td valign="top" align="left">2.7848</td>
<td valign="top" align="left">0.6275</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">3.7982</td>
<td valign="top" align="left">2.7856</td>
<td valign="top" align="left">0.6273</td>
</tr> <tr>
<td valign="top" align="left">D2</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">2.8320</td>
<td valign="top" align="left">2.0736</td>
<td valign="top" align="left">0.2579</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">2.5161</td>
<td valign="top" align="left">1.7633</td>
<td valign="top" align="left">0.6392</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">2.4875</td>
<td valign="top" align="left">1.7457</td>
<td valign="top" align="left">0.6434</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">2.4873</td>
<td valign="top" align="left">1.7445</td>
<td valign="top" align="left">0.6432</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">2.4878</td>
<td valign="top" align="left">1.7450</td>
<td valign="top" align="left">0.6434</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left"><bold>2.4852</bold></td>
<td valign="top" align="left"><bold>1.7429</bold></td>
<td valign="top" align="left"><bold>0.6436</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">2.4867</td>
<td valign="top" align="left">1.7441</td>
<td valign="top" align="left">0.6431</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">2.4874</td>
<td valign="top" align="left">1.7447</td>
<td valign="top" align="left">0.6429</td>
</tr> <tr>
<td valign="top" align="left">D3</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left"><bold>31.9275</bold></td>
<td valign="top" align="left"><bold>27.5843</bold></td>
<td valign="top" align="left"><bold>0.1452</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">32.0872</td>
<td valign="top" align="left">27.8724</td>
<td valign="top" align="left">0.1352</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">32.0903</td>
<td valign="top" align="left">27.8737</td>
<td valign="top" align="left">0.1348</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">32.0473</td>
<td valign="top" align="left">27.8755</td>
<td valign="top" align="left">0.1341</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">32.0335</td>
<td valign="top" align="left">27.8629</td>
<td valign="top" align="left">0.1352</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">32.0549</td>
<td valign="top" align="left">27.8799</td>
<td valign="top" align="left">0.1333</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">32.0614</td>
<td valign="top" align="left">27.8793</td>
<td valign="top" align="left">0.1330</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">32.0629</td>
<td valign="top" align="left">27.8772</td>
<td valign="top" align="left">0.1331</td>
</tr> <tr>
<td valign="top" align="left">D4</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.103200</td>
<td valign="top" align="left">0.088300</td>
<td valign="top" align="left">0.7049</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.100200</td>
<td valign="top" align="left">0.081800</td>
<td valign="top" align="left">0.7194</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.100027</td>
<td valign="top" align="left">0.082446</td>
<td valign="top" align="left">0.7202</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.099257</td>
<td valign="top" align="left">0.081848</td>
<td valign="top" align="left">0.7247</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.099807</td>
<td valign="top" align="left">0.082148</td>
<td valign="top" align="left">0.7224</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.098940</td>
<td valign="top" align="left">0.081626</td>
<td valign="top" align="left">0.7260</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left"><bold>0.098797</bold></td>
<td valign="top" align="left"><bold>0.081580</bold></td>
<td valign="top" align="left"><bold>0.7265</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">0.098812</td>
<td valign="top" align="left">0.081613</td>
<td valign="top" align="left"><bold>0.7265</bold></td>
</tr> <tr>
<td valign="top" align="left">D5</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.1811</td>
<td valign="top" align="left">0.1302</td>
<td valign="top" align="left">0.4476</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.1749</td>
<td valign="top" align="left">0.1242</td>
<td valign="top" align="left">0.4726</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.1745</td>
<td valign="top" align="left">0.1239</td>
<td valign="top" align="left">0.4752</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.1733</td>
<td valign="top" align="left">0.1228</td>
<td valign="top" align="left">0.4818</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.1749</td>
<td valign="top" align="left">0.1242</td>
<td valign="top" align="left">0.4726</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.1729</td>
<td valign="top" align="left">0.1223</td>
<td valign="top" align="left">0.4844</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">0.1723</td>
<td valign="top" align="left">0.1217</td>
<td valign="top" align="left">0.4879</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left"><bold>0.1722</bold></td>
<td valign="top" align="left"><bold>0.1216</bold></td>
<td valign="top" align="left"><bold>0.4886</bold></td>
</tr> <tr>
<td valign="top" align="left">D6</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.1391</td>
<td valign="top" align="left">0.1028</td>
<td valign="top" align="left">0.2026</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.1321</td>
<td valign="top" align="left">0.0970</td>
<td valign="top" align="left">0.2691</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.1320</td>
<td valign="top" align="left">0.0969</td>
<td valign="top" align="left">0.2696</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.1320</td>
<td valign="top" align="left">0.0969</td>
<td valign="top" align="left">0.2702</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.1321</td>
<td valign="top" align="left">0.0970</td>
<td valign="top" align="left">0.2691</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left"><bold>0.1318</bold></td>
<td valign="top" align="left"><bold>0.0967</bold></td>
<td valign="top" align="left"><bold>0.2720</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">0.1319</td>
<td valign="top" align="left">0.0968</td>
<td valign="top" align="left">0.2713</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">0.1319</td>
<td valign="top" align="left">0.0969</td>
<td valign="top" align="left">0.2705</td>
</tr> <tr>
<td valign="top" align="left">D7</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.1461</td>
<td valign="top" align="left">0.1233</td>
<td valign="top" align="left">0.8585</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.1454</td>
<td valign="top" align="left">0.1225</td>
<td valign="top" align="left">0.8563</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.1444</td>
<td valign="top" align="left">0.1223</td>
<td valign="top" align="left">0.8566</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.1406</td>
<td valign="top" align="left">0.1188</td>
<td valign="top" align="left">0.8646</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.1429</td>
<td valign="top" align="left">0.1213</td>
<td valign="top" align="left">0.8635</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.1400</td>
<td valign="top" align="left">0.1180</td>
<td valign="top" align="left">0.8651</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">0.1389</td>
<td valign="top" align="left">0.1171</td>
<td valign="top" align="left">0.8652</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left"><bold>0.1387</bold></td>
<td valign="top" align="left"><bold>0.1169</bold></td>
<td valign="top" align="left"><bold>0.8655</bold></td>
</tr> <tr>
<td valign="top" align="left">D8</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.0290436</td>
<td valign="top" align="left">0.0203419</td>
<td valign="top" align="left">0.0320561</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.0280759</td>
<td valign="top" align="left">0.0195862</td>
<td valign="top" align="left">0.0519505</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.0280758</td>
<td valign="top" align="left">0.0195861</td>
<td valign="top" align="left">0.0519602</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.0280756</td>
<td valign="top" align="left">0.0195860</td>
<td valign="top" align="left">0.0520021</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.0280759</td>
<td valign="top" align="left">0.0195862</td>
<td valign="top" align="left">0.0519789</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left"><bold>0.0280745</bold></td>
<td valign="top" align="left"><bold>0.0195852</bold></td>
<td valign="top" align="left"><bold>0.0520842</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">0.0280753</td>
<td valign="top" align="left">0.0195857</td>
<td valign="top" align="left">0.0520282</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">0.0280756</td>
<td valign="top" align="left">0.0195860</td>
<td valign="top" align="left">0.0520045</td>
</tr> <tr>
<td valign="top" align="left">D9</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">4.9607</td>
<td valign="top" align="left">3.6138</td>
<td valign="top" align="left">0.9043</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">4.8846</td>
<td valign="top" align="left">3.5413</td>
<td valign="top" align="left">0.9189</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">4.8699</td>
<td valign="top" align="left">3.5324</td>
<td valign="top" align="left">0.9192</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">4.8446</td>
<td valign="top" align="left">3.5176</td>
<td valign="top" align="left">0.9197</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">4.8839</td>
<td valign="top" align="left">3.5412</td>
<td valign="top" align="left">0.9189</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">4.8103</td>
<td valign="top" align="left"><bold>3.4896</bold></td>
<td valign="top" align="left"><bold>0.9205</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left"><bold>4.8101</bold></td>
<td valign="top" align="left">3.4962</td>
<td valign="top" align="left">0.9204</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">4.8204</td>
<td valign="top" align="left">3.5057</td>
<td valign="top" align="left">0.9202</td>
</tr> <tr>
<td valign="top" align="left">D10</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.9900</td>
<td valign="top" align="left">0.7839</td>
<td valign="top" align="left">0.0421</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.9532</td>
<td valign="top" align="left">0.7586</td>
<td valign="top" align="left">0.1469</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.9532</td>
<td valign="top" align="left">0.7586</td>
<td valign="top" align="left">0.1464</td>
</tr> <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.9525</td>
<td valign="top" align="left">0.7580</td>
<td valign="top" align="left">0.1440</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.9526</td>
<td valign="top" align="left">0.7581</td>
<td valign="top" align="left">0.1468</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.9526</td>
<td valign="top" align="left">0.7579</td>
<td valign="top" align="left">0.1436</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">0.9525</td>
<td valign="top" align="left">0.7579</td>
<td valign="top" align="left">0.1407</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left"><bold>0.9524</bold></td>
<td valign="top" align="left"><bold>0.7578</bold></td>
<td valign="top" align="left"><bold>0.1382</bold></td>
</tr> <tr>
<td valign="top" align="left">D11</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.3953</td>
<td valign="top" align="left">0.3048</td>
<td valign="top" align="left">0.8877</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.3802</td>
<td valign="top" align="left">0.2838</td>
<td valign="top" align="left">0.8999</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.3799</td>
<td valign="top" align="left">0.2837</td>
<td valign="top" align="left">0.9000</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.3760</td>
<td valign="top" align="left">0.2818</td>
<td valign="top" align="left">0.9023</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.3767</td>
<td valign="top" align="left">0.2821</td>
<td valign="top" align="left">0.9020</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.3754</td>
<td valign="top" align="left"><bold>0.2814</bold></td>
<td valign="top" align="left">0.9025</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left"><bold>0.3753</bold></td>
<td valign="top" align="left">0.2815</td>
<td valign="top" align="left"><bold>0.9026</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">0.3756</td>
<td valign="top" align="left">0.2818</td>
<td valign="top" align="left">0.9024</td>
</tr> <tr>
<td valign="top" align="left">D12</td>
<td valign="top" align="left">K-NN</td>
<td valign="top" align="left">3,051.2</td>
<td valign="top" align="left">2,105.2</td>
<td valign="top" align="left">0.8507</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">2,983.6</td>
<td valign="top" align="left">1,960.6</td>
<td valign="top" align="left">0.8810</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">2,961.1</td>
<td valign="top" align="left">1,949.2</td>
<td valign="top" align="left">0.8824</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">2,919.6</td>
<td valign="top" align="left">1,928.6</td>
<td valign="top" align="left">0.8848</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">2,983.6</td>
<td valign="top" align="left">1,960.6</td>
<td valign="top" align="left">0.8810</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">2,879.4</td>
<td valign="top" align="left">1,907.3</td>
<td valign="top" align="left">0.8871</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left"><bold>2,871.2</bold></td>
<td valign="top" align="left"><bold>1,904.3</bold></td>
<td valign="top" align="left"><bold>0.8872</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">2,877.5</td>
<td valign="top" align="left">1,907.5</td>
<td valign="top" align="left">0.8869</td>
</tr> <tr>
<td valign="top" align="left">D13</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">730.86</td>
<td valign="top" align="left">486.82</td>
<td valign="top" align="left">0.5993</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">710.57</td>
<td valign="top" align="left">475.59</td>
<td valign="top" align="left">0.6254</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">710.20</td>
<td valign="top" align="left">475.36</td>
<td valign="top" align="left">0.6255</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">710.71</td>
<td valign="top" align="left">475.52</td>
<td valign="top" align="left">0.6247</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">710.57</td>
<td valign="top" align="left">475.55</td>
<td valign="top" align="left">0.6254</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left"><bold>707.66</bold></td>
<td valign="top" align="left"><bold>473.38</bold></td>
<td valign="top" align="left"><bold>0.6272</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">710.54</td>
<td valign="top" align="left">475.21</td>
<td valign="top" align="left">0.6243</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">712.53</td>
<td valign="top" align="left">476.36</td>
<td valign="top" align="left">0.6224</td>
</tr> <tr>
<td valign="top" align="left">D14</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.143249</td>
<td valign="top" align="left">0.0968667</td>
<td valign="top" align="left">0.62008</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.141207</td>
<td valign="top" align="left">0.0969202</td>
<td valign="top" align="left">0.64373</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.141195</td>
<td valign="top" align="left">0.0969090</td>
<td valign="top" align="left">0.64377</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.141183</td>
<td valign="top" align="left">0.0968933</td>
<td valign="top" align="left">0.64380</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.141207</td>
<td valign="top" align="left">0.0969201</td>
<td valign="top" align="left">0.64373</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.141162</td>
<td valign="top" align="left">0.0968674</td>
<td valign="top" align="left">0.64385</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left"><bold>0.141160</bold></td>
<td valign="top" align="left"><bold>0.0968671</bold></td>
<td valign="top" align="left"><bold>0.64387</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left">0.141171</td>
<td valign="top" align="left">0.0968790</td>
<td valign="top" align="left">0.64385</td>
</tr> <tr>
<td valign="top" align="left">D15</td>
<td valign="top" align="left">KNN</td>
<td valign="top" align="left">0.027170</td>
<td valign="top" align="left">0.021513</td>
<td valign="top" align="left">0.189703</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">R-KNN</td>
<td valign="top" align="left">0.026824</td>
<td valign="top" align="left">0.021045</td>
<td valign="top" align="left">0.343268</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>RK-KNN with kernel:</bold></td>
<td/>
<td/>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Gaussian</td>
<td valign="top" align="left">0.026820</td>
<td valign="top" align="left">0.021042</td>
<td valign="top" align="left">0.343363</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Epanechnikov</td>
<td valign="top" align="left">0.026806</td>
<td valign="top" align="left">0.021032</td>
<td valign="top" align="left">0.343543</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Uniform</td>
<td valign="top" align="left">0.026822</td>
<td valign="top" align="left">0.021044</td>
<td valign="top" align="left">0.343268</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Triangular</td>
<td valign="top" align="left">0.026802</td>
<td valign="top" align="left">0.021029</td>
<td valign="top" align="left">0.343651</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Quartic</td>
<td valign="top" align="left">0.026790</td>
<td valign="top" align="left">0.021020</td>
<td valign="top" align="left">0.343954</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Tricube</td>
<td valign="top" align="left"><bold>0.026784</bold></td>
<td valign="top" align="left"><bold>0.021016</bold></td>
<td valign="top" align="left"><bold>0.344055</bold></td>
</tr></tbody>
</table>
</table-wrap>
<p>Overall, the triangular kernel emerges as the most effective, closely followed by the Tricube kernel. This observation is supported by instances across multiple datasets; for example, in dataset D1, the triangular kernel achieves an RMSE of 3.7949, an MAE of 2.7828, and an R<sup>2</sup> of 0.6304, surpassing the performance of classical KNN. Similarly, in dataset D6, the triangular kernel demonstrates superior results with an RMSE of 0.1318, an MAE of 0.0967, and an R<sup>2</sup> of 0.272.</p>
<p>The Gaussian and Epanechnikov kernels tend to not give the lowest RMSE or MAE but still perform notably well compared to the traditional KNN and R-KNN. The uniform kernel sometimes shows superiority compared to the other kernel functions.</p>
<p>Rankings are assigned to the methods from one to eight based on the values of RMSE, MAE, and R<sup>2</sup>. The lowest RMSE or MAE values receive a rank of one, indicating the best performance, while for R<sup>2</sup>, the highest value is awarded a rank of one. The rankings for RMSE, MAE, and R<sup>2</sup> are summarized in <xref ref-type="fig" rid="F3">Figures 3</xref>&#x02013;<xref ref-type="fig" rid="F5">5</xref>, respectively. These graphs demonstrate that the RK-KNN regression models generally achieve lower ranks, indicating better performance compared to the R-KNN and traditional KNN regression models. Specifically, for RMSE, the average ranks for RK-KNN with quartic, triangular, tricube, Epanechnikov, Gaussian, and uniform kernels are 1.93, 2.27, 3.13, 3.80, 5.20, and 5.27, respectively. In contrast, the R-KNN and traditional KNN models have average ranks of 6.40 and 7.53, respectively.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Comparative performance of RK-KNN with different kernel functions, R-KNN, and KNN regressions on multiple datasets using RMSE rankings.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Comparative performance of RK-KNN with different kernel functions, R-KNN, and KNN regressions on multiple datasets using MAE rankings.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparative performance of RK-KNN with different kernel functions, R-KNN, and KNN regressions on multiple datasets using R<sup>2</sup> rankings.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0005.tif"/>
</fig>
<p>For MAEs, RK-KNN models with triangular and quartic kernels exhibit nearly identical rankings, with average ranks of 2.33 and 2.67, respectively. The rankings for other methods are consistent with those observed for RMSE. The sequence features RK-KNN with tricube, Epanechnikov, Gaussian, and uniform kernels, followed by R-KNN and KNN, with respective average ranks of 3.27, 4.07, 5.07, 5.20, 6.00, and 7.07.</p>
<p>For R<sup>2</sup>, the RK-KNN model using the triangular kernel shows the best performance, achieving the lowest average rank of 2.60. It is followed by the RK-KNN models with quartic, tricube, Epanechnikov, uniform, and Gaussian kernels. R-KNN and KNN lag behind, with average ranks for R<sup>2</sup> being 3.13, 3.73, 4.07, 4.33, 4.67, 5.47, and 7.40, respectively.</p>
</sec>
<sec>
<title>4.3 Comparisons with state-of-the-art methods</title>
<p>The KNN models exhibiting the lowest RMSEs were benchmarked against RF, ANN, and SVR across fifteen diverse datasets, as detailed in <xref ref-type="table" rid="T3">Table 3</xref>. KNN-typed learners showed superior performance in datasets D3, D4, D5, D7, and D8, representing a third of the datasets. However, they were notably outperformed by RF in nine datasets (D1, D2, D6, D9, D10, D12, D13, D14, and D15) and by ANN and SVR in the remaining datasets. Although RK-KNN regression did not achieve the lowest RMSE in all datasets, it remains a competitive option, particularly against SVR and ANN. This is especially evident in datasets D5 and D8, which contain a high number of features, where RK-KNN was preferred over the other models.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Performance evaluation of best KNN-typed learner, RF, ANN, and SVR (bold values represent the best performance).</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>RMSE</bold></th>
<th valign="top" align="left"><bold>MAE</bold></th>
<th valign="top" align="left"><bold>R<sup>2</sup></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">D1</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">3.7949</td>
<td valign="top" align="left">2.7828</td>
<td valign="top" align="left">0.6284</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>1.9710</bold></td>
<td valign="top" align="left"><bold>1.1991</bold></td>
<td valign="top" align="left"><bold>0.8105</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">2.8195</td>
<td valign="top" align="left">2.1649</td>
<td valign="top" align="left">0.6123</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">2.9470</td>
<td valign="top" align="left">1.9306</td>
<td valign="top" align="left">0.5864</td>
</tr> <tr>
<td valign="top" align="left">D2</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">2.4852</td>
<td valign="top" align="left">1.7429</td>
<td valign="top" align="left">0.6436</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>1.2806</bold></td>
<td valign="top" align="left"><bold>0.8084</bold></td>
<td valign="top" align="left"><bold>0.8384</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">1.7284</td>
<td valign="top" align="left">1.2457</td>
<td valign="top" align="left">0.7065</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">1.7219</td>
<td valign="top" align="left">1.0574</td>
<td valign="top" align="left">0.7138</td>
</tr> <tr>
<td valign="top" align="left">D3</td>
<td valign="top" align="left"><bold>KNN</bold></td>
<td valign="top" align="left"><bold>31.9275</bold></td>
<td valign="top" align="left"><bold>27.5843</bold></td>
<td valign="top" align="left"><bold>0.1452</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">34.0136</td>
<td valign="top" align="left">28.8079</td>
<td valign="top" align="left">0.0326</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">46.5660</td>
<td valign="top" align="left">37.2240</td>
<td valign="top" align="left">&#x02212;0.8355</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">34.8618</td>
<td valign="top" align="left">28.7294</td>
<td valign="top" align="left">&#x02212;0.0121</td>
</tr> <tr>
<td valign="top" align="left">D4</td>
<td valign="top" align="left"><bold>Random Kernel KNN</bold></td>
<td valign="top" align="left"><bold>0.0987</bold></td>
<td valign="top" align="left"><bold>0.0815</bold></td>
<td valign="top" align="left"><bold>0.7265</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">0.1048</td>
<td valign="top" align="left">0.0854</td>
<td valign="top" align="left">0.6867</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">0.2081</td>
<td valign="top" align="left">0.1484</td>
<td valign="top" align="left">&#x02212;0.5068</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.1401</td>
<td valign="top" align="left">0.1057</td>
<td valign="top" align="left">0.6459</td>
</tr> <tr>
<td valign="top" align="left">D5</td>
<td valign="top" align="left"><bold>Random Kernel KNN</bold></td>
<td valign="top" align="left"><bold>0.1722</bold></td>
<td valign="top" align="left"><bold>0.1216</bold></td>
<td valign="top" align="left"><bold>0.4886</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">0.1826</td>
<td valign="top" align="left">0.1297</td>
<td valign="top" align="left">0.2731</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">4.9656</td>
<td valign="top" align="left">1.5578</td>
<td valign="top" align="left">&#x02212;1,988.9871</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.1945</td>
<td valign="top" align="left">0.4237</td>
<td valign="top" align="left">0.2521</td>
</tr> <tr>
<td valign="top" align="left">D6</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">0.1318</td>
<td valign="top" align="left">0.0967</td>
<td valign="top" align="left">0.2720</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>0.1243</bold></td>
<td valign="top" align="left"><bold>0.0912</bold></td>
<td valign="top" align="left"><bold>0.3758</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">0.2709</td>
<td valign="top" align="left">0.1862</td>
<td valign="top" align="left">&#x02212;1.9643</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.1398</td>
<td valign="top" align="left">0.1003</td>
<td valign="top" align="left">0.2095</td>
</tr> <tr>
<td valign="top" align="left">D7</td>
<td valign="top" align="left"><bold>Random Kernel KNN</bold></td>
<td valign="top" align="left"><bold>0.1387</bold></td>
<td valign="top" align="left"><bold>0.1169</bold></td>
<td valign="top" align="left"><bold>0.8655</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">0.1486</td>
<td valign="top" align="left">0.1208</td>
<td valign="top" align="left">0.3686</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">3.3553</td>
<td valign="top" align="left">2.2652</td>
<td valign="top" align="left">&#x02212;2,239.9117</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.2118</td>
<td valign="top" align="left">0.1659</td>
<td valign="top" align="left">&#x02212;0.2131</td>
</tr> <tr>
<td valign="top" align="left">D8</td>
<td valign="top" align="left"><bold>Random Kernel KNN</bold></td>
<td valign="top" align="left"><bold>0.0280</bold></td>
<td valign="top" align="left"><bold>0.0195</bold></td>
<td valign="top" align="left"><bold>0.0520</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">0.0297</td>
<td valign="top" align="left">0.0201</td>
<td valign="top" align="left">0.0526</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">0.1546</td>
<td valign="top" align="left">0.0746</td>
<td valign="top" align="left">&#x02212;29.6140</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.0376</td>
<td valign="top" align="left">0.0297</td>
<td valign="top" align="left">&#x02212;0.5931</td>
</tr> <tr>
<td valign="top" align="left">D9</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">4.8103</td>
<td valign="top" align="left">3.4896</td>
<td valign="top" align="left">0.9205</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>1.5273</bold></td>
<td valign="top" align="left"><bold>1.0882</bold></td>
<td valign="top" align="left"><bold>0.9883</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">3.8876</td>
<td valign="top" align="left">2.6226</td>
<td valign="top" align="left">0.9292</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">11.1648</td>
<td valign="top" align="left">7.9844</td>
<td valign="top" align="left">0.4088</td>
</tr> <tr>
<td valign="top" align="left">D10</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">0.9524</td>
<td valign="top" align="left">0.7578</td>
<td valign="top" align="left">0.1382</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>0.3749</bold></td>
<td valign="top" align="left"><bold>0.2891</bold></td>
<td valign="top" align="left"><bold>0.8640</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">1.0516</td>
<td valign="top" align="left">0.8608</td>
<td valign="top" align="left">&#x02212;0.0693</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.9534</td>
<td valign="top" align="left">0.7723</td>
<td valign="top" align="left">0.1211</td>
</tr> <tr>
<td valign="top" align="left">D11</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">0.3753</td>
<td valign="top" align="left">0.2815</td>
<td valign="top" align="left">0.9026</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">0.3839</td>
<td valign="top" align="left">0.2803</td>
<td valign="top" align="left">0.8846</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">0.5566</td>
<td valign="top" align="left">0.4287</td>
<td valign="top" align="left">0.7311</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Support Vector Regression</bold></td>
<td valign="top" align="left"><bold>0.3022</bold></td>
<td valign="top" align="left"><bold>0.1816</bold></td>
<td valign="top" align="left"><bold>0.9238</bold></td>
</tr> <tr>
<td valign="top" align="left">D12</td>
<td valign="top" align="left">RK-KNN with Quartic</td>
<td valign="top" align="left">2,871.2</td>
<td valign="top" align="left">1,904.3</td>
<td valign="top" align="left">0.8872</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>2,145.3</bold></td>
<td valign="top" align="left"><bold>1,470.5</bold></td>
<td valign="top" align="left"><bold>0.8921</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">12,614.1</td>
<td valign="top" align="left">11,164.4</td>
<td valign="top" align="left">&#x02212;2.2913</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">7,493.0</td>
<td valign="top" align="left">4,895.2</td>
<td valign="top" align="left">&#x02212;0.1041</td>
</tr> <tr>
<td valign="top" align="left">D13</td>
<td valign="top" align="left">RK-KNN with Triangular</td>
<td valign="top" align="left">707.66</td>
<td valign="top" align="left">473.38</td>
<td valign="top" align="left">0.6272</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>246.12</bold></td>
<td valign="top" align="left"><bold>129.67</bold></td>
<td valign="top" align="left"><bold>0.9564</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">198.85</td>
<td valign="top" align="left">116.84</td>
<td valign="top" align="left">0.9712</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">1,252.86</td>
<td valign="top" align="left">805.01</td>
<td valign="top" align="left">&#x02212;0.0971</td>
</tr> <tr>
<td valign="top" align="left">D14</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">0.141160</td>
<td valign="top" align="left">0.0968671</td>
<td valign="top" align="left">0.64387</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>0.136293</bold></td>
<td valign="top" align="left"><bold>0.0939809</bold></td>
<td valign="top" align="left"><bold>0.64433</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">0.196391</td>
<td valign="top" align="left">0.1486414</td>
<td valign="top" align="left">0.26220</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.142445</td>
<td valign="top" align="left">0.1037052</td>
<td valign="top" align="left">0.61202</td>
</tr> <tr>
<td valign="top" align="left">D15</td>
<td valign="top" align="left">Random Kernel KNN</td>
<td valign="top" align="left">0.026784</td>
<td valign="top" align="left">0.021016</td>
<td valign="top" align="left">0.34405</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left"><bold>Random Forest</bold></td>
<td valign="top" align="left"><bold>0.007750</bold></td>
<td valign="top" align="left"><bold>0.006169</bold></td>
<td valign="top" align="left"><bold>0.93407</bold></td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Artificial Neural Network</td>
<td valign="top" align="left">0.066096</td>
<td valign="top" align="left">0.052390</td>
<td valign="top" align="left">&#x02212;3.87338</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Support Vector Regression</td>
<td valign="top" align="left">0.030218</td>
<td valign="top" align="left">0.023477</td>
<td valign="top" align="left">&#x02212;0.00120</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion and discussion</title>
<p>This study validates the efficacy of integrating kernel functions with a random process, which includes both bootstrapping and feature selection, across 15 datasets. Our comprehensive evaluation, based on criteria such as RMSE, MAE, and R<sup>2</sup>, underscores the superiority of the RK-KNN approach, especially when employing quartic, triangular, and tricube kernel functions. These kernels have consistently demonstrated performance enhancements across various case studies.</p>
<p>Specifically, in dataset D10, RK-KNN regression markedly improves prediction accuracy. The RMSE distributions depicted in <xref ref-type="fig" rid="F6">Figure 6</xref> reveal that standard KNN exhibits higher RMSE compared to other methods, with all parameter configurations for RK-KNN outperforming standard KNN. However, achieving optimal performance across datasets may require a comprehensive search to identify the best parameters for the selected kernel functions. As shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, while the lowest RMSE values for RK-KNN across all kernel functions are superior to those of KNN, the medians of RMSEs for some kernels, like the Gaussian kernel, exceed the median RMSE of KNN. This variability indicates a critical need for tuning the optimal bandwidth and <italic>k</italic>-value to consistently achieve the lowest RMSE.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The RMSE distribution across all eight methods for dataset D10.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>The RMSE distribution across all eight methods for dataset D12.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-07-1402384-g0007.tif"/>
</fig>
<p>Moreover, the computational cost and scalability of the RK-KNN algorithm&#x00027;s cross-validation process are effectively managed through vectorized distance computations, which enhance calculation speed and reduce runtime. Standardization of features further contributes to this efficiency by simplifying the distance metric computation. A well-controlled grid search for parameter tuning, along with the ability to independently execute bootstrapping and feature selection steps, ensures computational tractability. Practical applications across multiple datasets have demonstrated that the cross-validation step, a critical aspect of the RK-KNN algorithm, is not prohibitively time-consuming. Therefore, the RK-KNN method is computationally efficient and well-suited for the analysis of large-scale data environments.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p></sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>PS: Conceptualization, Funding acquisition, Investigation, Methodology, Project administration, Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. KS: Data curation, Formal analysis, Resources, Software, Visualization, Writing &#x02013; review &#x00026; editing.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.</p>
</sec>
<ack><p>The authors would like to thank the reviewers for their valuable comments and suggestions.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abdalla</surname> <given-names>H. I.</given-names></name> <name><surname>Amer</surname> <given-names>A. A.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Towards highly-efficient k-nearest neighbor algorithm for big data classification,&#x0201D;</article-title> in <source>2022 5th International Conference on Networking, Information Systems and Security</source> (<publisher-loc>New York City, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alcal&#x000E1;-Fdez</surname> <given-names>J.</given-names></name> <name><surname>Fernandez</surname> <given-names>A.</given-names></name> <name><surname>Luengo</surname> <given-names>J.</given-names></name> <name><surname>Derrac</surname> <given-names>J.</given-names></name> <name><surname>Garc&#x000ED;a</surname> <given-names>S.</given-names></name> <name><surname>S&#x000E1;nchez</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>KEEL data-mining software tool: data set repository, integration of algorithms and experimental analysis framework</article-title>. <source>J. Mult.-Valued Log. Soft Comput</source>. <volume>17</volume>, <fpage>255</fpage>&#x02013;<lpage>287</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ali</surname> <given-names>A.</given-names></name> <name><surname>Hamraz</surname> <given-names>M.</given-names></name> <name><surname>Kumam</surname> <given-names>P.</given-names></name> <name><surname>Khan</surname> <given-names>D. M.</given-names></name> <name><surname>Khalil</surname> <given-names>U.</given-names></name> <name><surname>Sulaiman</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A k-nearest neighbours based ensemble via optimal model selection for regression</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>132095</fpage>&#x02013;<lpage>132105</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3010099</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altman</surname> <given-names>N. S.</given-names></name></person-group> (<year>1992</year>). <article-title>An introduction to kernel and nearest-neighbor nonparametric regression</article-title>. <source>Am. Stat</source>. <volume>46</volume>, <fpage>175</fpage>&#x02013;<lpage>185</lpage>. <pub-id pub-id-type="doi">10.1080/00031305.1992.10475879</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bay</surname> <given-names>S. D.</given-names></name></person-group> (<year>1999</year>). <article-title>Nearest neighbor classification from multiple feature subsets</article-title>. <source>Intell. Data Anal</source>. <volume>3</volume>, <fpage>191</fpage>&#x02013;<lpage>209</lpage>. <pub-id pub-id-type="doi">10.3233/IDA-1999-3304</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beitollahi</surname> <given-names>H.</given-names></name> <name><surname>Sharif</surname> <given-names>D. M.</given-names></name> <name><surname>Fazeli</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Application layer DDoS attack detection using cuckoo search algorithm-trained radial basis function</article-title>. <source>IEEE Access</source> <volume>10</volume>, <fpage>63844</fpage>&#x02013;<lpage>63854</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2022.3182818</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bermejo</surname> <given-names>S.</given-names></name> <name><surname>Cabestany</surname> <given-names>J.</given-names></name></person-group> (<year>2000</year>). <article-title>Adaptive soft k-nearest-neighbour classifiers</article-title>. <source>Pattern Recogn.</source> <volume>33</volume>, <fpage>1999</fpage>&#x02013;<lpage>2005</lpage>. <pub-id pub-id-type="doi">10.1016/S0031-3203(99)00186-7</pub-id><pub-id pub-id-type="pmid">34567277</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bian</surname> <given-names>C.</given-names></name> <name><surname>Huang</surname> <given-names>G. Q.</given-names></name></person-group> (<year>2024</year>). <article-title>Air pollution concentration fuzzy evaluation based on evidence theory and the K-nearest neighbor algorithm</article-title>. <source>Front. Environ. Sci</source>. <volume>12</volume>:<fpage>1243962</fpage>. <pub-id pub-id-type="doi">10.3389/fenvs.2024.1243962</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><collab>Borggaard C. and Thodberg, H. H..</collab></person-group> (<year>1992</year>). <article-title>Optimal minimal neural interpretation of spectra</article-title>. <source>Anal. Chem</source>. <volume>64</volume>, <fpage>545</fpage>&#x02013;<lpage>551</lpage>. <pub-id pub-id-type="doi">10.1021/ac00029a018</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L.</given-names></name></person-group> (<year>2001</year>). <article-title>Random forests</article-title>. <source>Mach. Learn</source>. <volume>45</volume>, <fpage>5</fpage>&#x02013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/A:1010933404324</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>D.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Deng</surname> <given-names>Z.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Zong</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;kNN algorithm with data-driven k value,&#x0201D;</article-title> in <source>Advanced Data Mining and Applications Lecture Notes in Computer Science</source>, 8933. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chung</surname> <given-names>Y.-W.</given-names></name> <name><surname>Khaki</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>T.</given-names></name> <name><surname>Chu</surname> <given-names>C.</given-names></name> <name><surname>Gadh</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>Ensemble machine learning-based algorithm for electric vehicle user behavior prediction</article-title>. <source>Appl. Energy</source> <volume>254</volume>, <fpage>113732</fpage>. <pub-id pub-id-type="doi">10.1016/j.apenergy.2019.113732</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cortez</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <source>Student Performance</source>. <publisher-loc>Boston</publisher-loc>: <publisher-name>UCI Machine Learning Repository</publisher-name>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>Z.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Cheng</surname> <given-names>D.</given-names></name> <name><surname>Zong</surname> <given-names>M. Zhang, S.</given-names></name></person-group> (<year>2016</year>). <article-title>Efficient kNN classification algorithm for big data</article-title>. <source>Neurocomputing</source> <volume>195</volume>, <fpage>143</fpage>&#x02013;<lpage>148</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2015.08.112</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dimopoulos</surname> <given-names>A. C.</given-names></name> <name><surname>Nikolaidou</surname> <given-names>M.</given-names></name> <name><surname>Caballero</surname> <given-names>F. F.</given-names></name> <name><surname>Engchuan</surname> <given-names>W.</given-names></name> <name><surname>Sanchez-Niubo</surname> <given-names>A.</given-names></name> <name><surname>Arndt</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Machine learning methodologies versus cardiovascular risk scores, in predicting disease risk</article-title>. <source>BMC Med Res Methodol</source>. <volume>18</volume>:<fpage>179</fpage>. <pub-id pub-id-type="doi">10.1186/s12874-018-0644-1</pub-id><pub-id pub-id-type="pmid">30594138</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>El-Kenawy</surname> <given-names>E.-S. M.</given-names></name> <name><surname>Mirjalili</surname> <given-names>S.</given-names></name> <name><surname>Ghoneim</surname> <given-names>S. M.</given-names></name> <name><surname>Eid</surname> <given-names>M. M.</given-names></name> <name><surname>El-Said</surname> <given-names>M.</given-names></name> <name><surname>Khan</surname> <given-names>Z. S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Advanced ensemble model for solar radiation forecasting using sine cosine algorithm and Newton&#x00027;s laws</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>115750</fpage>&#x02013;<lpage>115765</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3106233</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Enriquez</surname> <given-names>A. R. S.</given-names></name> <name><surname>Lima</surname> <given-names>S. L.</given-names></name> <name><surname>Saavedra</surname> <given-names>O. R.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;K-NN and mean-shift algorithm applied in fault diagnosis in power transformers by DGA,&#x0201D;</article-title> in <source>Presented at the 2019 20th International Conference on Intelligent System Application to Power Systems (ISAP</source>) (New Delhi: ISAP), 1-6.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feng</surname> <given-names>J.</given-names></name> <name><surname>Lurati</surname> <given-names>L.</given-names></name> <name><surname>Ouyang</surname> <given-names>H.</given-names></name> <name><surname>Robinson</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Yuan</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Predictive toxicology: benchmarking molecular descriptors and statistical methods</article-title>. <source>J. Chem. Inf. Comput. Sci</source>. <volume>43</volume>, <fpage>1463</fpage>&#x02013;<lpage>1470</lpage>. <pub-id pub-id-type="doi">10.1021/ci034032s</pub-id><pub-id pub-id-type="pmid">14502479</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Friedman</surname> <given-names>J. H.</given-names></name></person-group> (<year>1999</year>). <source>Greedy Function Approximation: A Gradient Boosting Machine</source>. <publisher-loc>Stanford, CA</publisher-loc>: <publisher-name>Technical Report, Department of Statistics, Stanford University</publisher-name>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garc&#x000ED;a-Pedrajas</surname> <given-names>N.</given-names></name> <name><surname>Ortiz-Boyer</surname> <given-names>D.</given-names></name></person-group> (<year>2009</year>). <article-title>Boosting k-nearest neighbor classifier by means of input space projection</article-title>. <source>Expert Syst. Appl</source>. <volume>36</volume>, <fpage>10570</fpage>&#x02013;<lpage>10582</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2009.02.065</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ghavami</surname> <given-names>S.</given-names></name> <name><surname>Alipour</surname> <given-names>Z.</given-names></name> <name><surname>Naseri</surname> <given-names>H.</given-names></name> <name><surname>Jahanbakhsh</surname> <given-names>H.</given-names></name> <name><surname>Karimi</surname> <given-names>M. M.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;A new ensemble prediction method for reclaimed asphalt pavement (RAP) mixtures containing different constituents</article-title>. <source>Buildings</source> <volume>13</volume>:<fpage>1787</fpage>. <pub-id pub-id-type="doi">10.3390/buildings13071787</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guha</surname> <given-names>R.</given-names></name> <name><surname>Jurs</surname> <given-names>P. C.</given-names></name></person-group> (<year>2004</year>). <article-title>Development of linear, ensemble, and nonlinear models for the prediction and interpretation of the biological activity of a set of PDGFR inhibitors</article-title>. <source>J. Chem. Inf. Comput. Sci</source>. <volume>44</volume>, <fpage>2179</fpage>&#x02013;<lpage>2189</lpage>. <pub-id pub-id-type="doi">10.1021/ci049849f</pub-id><pub-id pub-id-type="pmid">15554688</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name> <name><surname>Friedman</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <source>The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Series in Statistics, 2nd ed</source>. (New York, NY, USA: Springer).</citation>
</ref>
<ref id="B24">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Helliwell</surname> <given-names>J.</given-names></name> <name><surname>Layard</surname> <given-names>R.</given-names></name> <name><surname>Sachs</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <source>World Happiness Report 2017.</source> <publisher-loc>New York</publisher-loc>: <publisher-name>Sustainable Development Solutions Network</publisher-name>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://worldhappiness.report/ed/2017/">https://worldhappiness.report/ed/2017/</ext-link> (accessed January 9, 2024).</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hirst</surname> <given-names>J. D.</given-names></name> <name><surname>King</surname> <given-names>R. D.</given-names></name> <name><surname>Sternberg</surname> <given-names>M. J.</given-names></name></person-group> (<year>1994</year>). <article-title>Quantitative structure-activity relationships by neural networks and inductive logic programming. II. The inhibition of dihydrofolate reductase by triazines</article-title>. <source>J. Comput. Aided Mol. Des</source>. <volume>8</volume>, <fpage>421</fpage>&#x02013;<lpage>432</lpage>. <pub-id pub-id-type="doi">10.1007/BF00125376</pub-id><pub-id pub-id-type="pmid">7815093</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hofmann</surname> <given-names>T.</given-names></name> <name><surname>Sch&#x000F6;lkopf</surname> <given-names>B.</given-names></name> <name><surname>Smola</surname> <given-names>A. J.</given-names></name></person-group> (<year>2008</year>). <article-title>Kernel methods in machine learning</article-title>. <source>Ann. Statist</source>. <volume>36</volume>, <fpage>1171</fpage>&#x02013;<lpage>1220</lpage>. <pub-id pub-id-type="doi">10.1214/009053607000000677</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ingram</surname> <given-names>S.</given-names></name> <name><surname>Munzner</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>Dimensionality reduction for documents with nearest neighbor queries</article-title>. <source>Neurocomputing</source> <volume>150</volume>, <fpage>557</fpage>&#x02013;<lpage>569</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2014.07.073</pub-id><pub-id pub-id-type="pmid">37713261</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jafar</surname> <given-names>R.</given-names></name> <name><surname>Awad</surname> <given-names>A.</given-names></name> <name><surname>Hatem</surname> <given-names>I.</given-names></name> <name><surname>Jafar</surname> <given-names>K.</given-names></name> <name><surname>Awad</surname> <given-names>E.</given-names></name> <name><surname>Shahrour</surname> <given-names>I.</given-names></name></person-group> (<year>2023</year>). <article-title>Multiple linear regression and machine learning for predicting the drinking water quality index in Al-seine lake</article-title>. <source>Smart Cities</source> <volume>6</volume>, <fpage>2807</fpage>&#x02013;<lpage>2827</lpage>. <pub-id pub-id-type="doi">10.3390/smartcities6050126</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>Smith</surname> <given-names>P.</given-names></name> <name><surname>Pang</surname> <given-names>Q.</given-names></name></person-group> (<year>2023</year>). <article-title>Ensemble machine learning for modeling greenhouse gas emissions at different time scales from irrigated paddy fields</article-title>. <source>Field Crops Res</source>. <volume>292</volume>:<fpage>108821</fpage>. <pub-id pub-id-type="doi">10.1016/j.fcr.2023.108821</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kubinyi</surname> <given-names>H.</given-names></name></person-group> (<year>1993</year>). <source>QSAR: Hansch Analysis and Related Approaches. Methods and Principles in Medicinal Chemistry.</source> <publisher-loc>Weinheim; New York, NY</publisher-loc>: <publisher-name>Wiley-VCH, 438</publisher-name>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.-c.</given-names></name> <name><surname>Xu</surname> <given-names>S.-w.</given-names></name> <name><surname>Zhuang</surname> <given-names>J.-y.</given-names></name> <name><surname>Liu</surname> <given-names>J.-j.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.-x.</given-names></name></person-group> (<year>2023</year>). <article-title>Ensemble learning prediction of soybean yields in China based on meteorological data</article-title>. <source>J. Integr. Agric</source>. <volume>22</volume>, <fpage>1909</fpage>&#x02013;<lpage>1927</lpage>. <pub-id pub-id-type="doi">10.1016/j.jia.2023.02.011</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Harner</surname> <given-names>E. J.</given-names></name> <name><surname>Adjeroh</surname> <given-names>D. A.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Random KNN,&#x0201D;</article-title> in <source>Presented at the 2014 IEEE International Conference on Data Mining Workshop, Shenzhen, China</source> <publisher-loc>(Los Alamitos, CA; Washington, DC; Tokyo</publisher-loc>: <publisher-name>IEEE Computer Society Conference Publishing Services (CPS))</publisher-name> <fpage>629</fpage>&#x02013;<lpage>636</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="web"><person-group person-group-type="author"><collab>OpenML</collab></person-group> (<year>2024</year>). <source>dataset-autoHorse_fixed</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.openml.org/d/42224">https://www.openml.org/d/42224</ext-link> (accessed February 02, 2024).</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pramanik</surname> <given-names>P. K.D.</given-names></name> <name><surname>Mukhopadhyay</surname> <given-names>M.</given-names></name> <name><surname>Pal</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Big data classification: applications and challenges,&#x0201D;</article-title> in <source>Artificial Intelligence and IoT. Studies in Big Data</source>, eds. K. G. Manoharan, J. A. Nehru, and S. Balasubramanian (Singapore: Springer).</citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rafiei</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <source>Residential Building Data Set</source>. <publisher-loc>Boston</publisher-loc>: <publisher-name>UCI Machine Learning Repository</publisher-name>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Redmond</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <source>Communities and Crime</source>. <publisher-loc>Boston</publisher-loc>: <publisher-name>UCI Machine Learning Repository</publisher-name>.</citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rubio</surname> <given-names>G.</given-names></name> <name><surname>Guillen</surname> <given-names>A.</given-names></name> <name><surname>Pomares</surname> <given-names>H.</given-names></name> <name><surname>Rojas</surname> <given-names>I.</given-names></name> <name><surname>Paechter</surname> <given-names>B.</given-names></name> <name><surname>Glosekotter</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>&#x0201C;Parallelization of the nearest-neighbour search and the cross-validation error evaluation for the kernel weighted k-nn algorithm applied to large data sets in MATLAB,&#x0201D;</article-title> in <source>Presented at the 2009 International Conference on High Performance Computing &#x00026; Simulation</source> (<publisher-loc>New York City, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saadatfar</surname> <given-names>H.</given-names></name> <name><surname>Khosravi</surname> <given-names>S.</given-names></name> <name><surname>Joloudari</surname> <given-names>J. H.</given-names></name> <name><surname>Mosavi</surname> <given-names>A.</given-names></name> <name><surname>Shamshirband</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>A new K-nearest neighbors classifier for big data based on efficient data pruning</article-title>. <source>Mathematics</source> <volume>8</volume>:<fpage>286</fpage>. <pub-id pub-id-type="doi">10.3390/math8020286</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sch&#x000F6;lkopf</surname> <given-names>B.</given-names></name> <name><surname>Smola</surname> <given-names>A. J.</given-names></name></person-group> (<year>2001</year>). <source>Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond, 1st ed</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sharma</surname> <given-names>S.</given-names></name> <name><surname>Lakshmi</surname> <given-names>L. R.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Improved k-NN regression model using random forests for air pollution prediction,&#x0201D;</article-title> in <source>Presented at the International Conference on Smart Applications, Communications, and Networking (SmartNets)</source> (<publisher-loc>Istanbul</publisher-loc>: <publisher-name>SmartNets</publisher-name>).</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>H.</given-names></name> <name><surname>Choi</surname> <given-names>H.</given-names></name></person-group> (<year>2023</year>). <article-title>Forecasting stock market indices using the recurrent neural network based hybrid models: CNN-LSTM, GRU-CNN, and ensemble models</article-title>. <source>Appl. Sci</source>. <volume>13</volume>:<fpage>4644</fpage>. <pub-id pub-id-type="doi">10.3390/app13074644</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>J.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>F.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Qian</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name></person-group> (<year>2018</year>). <article-title>A novel regression modeling method for PMSLM structural design optimization using a distance-weighted KNN algorithm</article-title>. <source>IEEE Trans. Indust. Appl</source>. <volume>54</volume>, <fpage>4198</fpage>&#x02013;<lpage>4206</lpage>. <pub-id pub-id-type="doi">10.1109/TIA.2018.2836953</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srisuradetchai</surname> <given-names>P.</given-names></name></person-group> (<year>2023</year>). <article-title>A novel interval forecast for k-nearest neighbor time series: a case study of durian export in Thailand</article-title>. <source>IEEE Access</source>. <volume>12</volume>, <fpage>2032</fpage>&#x02013;<lpage>2044</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3348078</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srisuradetchai</surname> <given-names>P.</given-names></name> <name><surname>Panichkitkosolkul</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Using ensemble machine learning methods to forecast particulate matter (PM<sub>2.5</sub>) in Bangkok, Thailand,&#x0201D;</article-title> in <source>Multi-disciplinary Trends in Artificial Intelligence</source>, eds. O. Surinta and K. F. Yuen (Cham: Springer).</citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Srisuradetchai</surname> <given-names>P.</given-names></name> <name><surname>Panichkitkosolkul</surname> <given-names>W.</given-names></name> <name><surname>Phaphan</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Combining machine learning models with ARIMA for COVID-19 epidemic in Thailand,&#x0201D;</article-title> in <source>Proceedings of the 2023 Research, Invention, and Innovation Congress: Innovation in Electrical and Electronics (RI2C), Bangkok, Thailand</source> (<publisher-loc>New York City, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>155</fpage>&#x02013;<lpage>161</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steele</surname> <given-names>B. M.</given-names></name></person-group> (<year>2009</year>). <article-title>Exact bootstrap k-nearest neighbor learners</article-title>. <source>Mach. Learn</source>. <volume>74</volume>, <fpage>235</fpage>&#x02013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-008-5096-0</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>R.</given-names></name> <name><surname>Ottewill</surname> <given-names>J. R.</given-names></name> <name><surname>Thornhill</surname> <given-names>N. F.</given-names></name></person-group> (<year>2020</year>). <article-title>Monitoring statistics and tuning of Kernel principal component analysis with radial basis function kernels</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>198328</fpage>&#x02013;<lpage>198342</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3034550</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thodberg</surname> <given-names>H. H.</given-names></name></person-group> (<year>1996</year>). <article-title>A review of Bayesian neural networks with an application to near infrared spectroscopy</article-title>. <source>IEEE Trans. Neural Networks</source> <volume>7</volume>, <fpage>56</fpage>&#x02013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.1109/72.478392</pub-id><pub-id pub-id-type="pmid">18255558</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Todeschini</surname> <given-names>R. Gramatica, P.</given-names></name> <name><surname>Pravenzani</surname> <given-names>R.</given-names></name> <name><surname>Marengo</surname> <given-names>E.</given-names></name></person-group> (<year>1995</year>). <article-title>Weighted holistic invariant molecular descriptors. Part 2. Theory development and applications on modeling physicochemical properties of polyaromatic hydrocarbons</article-title>. <source>Chemometrics Intell. Lab. Syst</source>. <volume>27</volume>, <fpage>221</fpage>&#x02013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1016/0169-7439(94)00025-E</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tsybakov</surname> <given-names>A. B.</given-names></name></person-group> (<year>2009</year>). <source>Introduction to Nonparametric Estimation. 1st ed</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ukey</surname> <given-names>N.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>Y. Zhang, W.</given-names></name></person-group> (<year>2023</year>). <article-title>Survey on exact kNN queries over high-dimensional data space</article-title>. <source>Sensors</source> <volume>23</volume>:<fpage>629</fpage>. <pub-id pub-id-type="doi">10.3390/s23020629</pub-id><pub-id pub-id-type="pmid">36679422</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wolberg</surname> <given-names>W.</given-names></name> <name><surname>Mangasarian</surname> <given-names>O.</given-names></name> <name><surname>Street</surname> <given-names>N.</given-names></name> <name><surname>Street</surname> <given-names>W.</given-names></name></person-group> (<year>1995</year>). <source>Breast Cancer Wisconsin (Diagnostic).</source> <publisher-loc>Boston</publisher-loc>: <publisher-name>UCI Machine Learning Repository</publisher-name>.</citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wong</surname> <given-names>T.-T.</given-names></name> <name><surname>Yang</surname> <given-names>N.-Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Dependency analysis of accuracy estimates in k-fold cross validation</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <volume>29</volume>, <fpage>2417</fpage>&#x02013;<lpage>2427</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2017.2740926</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wong</surname> <given-names>T.-T.</given-names></name> <name><surname>Yeh</surname> <given-names>P.-Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Reliable accuracy estimates from k-fold cross validation</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <volume>32</volume>, <fpage>1586</fpage>&#x02013;<lpage>1594</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2019.2912815</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Jiang</surname> <given-names>B.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Multiple kernel k-means clustering by selecting representative kernels</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>32</volume>, <fpage>4983</fpage>&#x02013;<lpage>4996</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2020.3026532</pub-id><pub-id pub-id-type="pmid">33017298</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>G.</given-names></name> <name><surname>Cao</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;A Modified K-NN algorithm for holter waveform classification based on kernel function,&#x0201D;</article-title> in <source>2008 Fifth International Conference on Fuzzy Systems and Knowledge Discovery, Jinan, China</source> (<publisher-loc>Los Alamitos, CA; Washington, DC; Tokyo</publisher-loc>: <publisher-name>IEEE Computer Society Conference Publishing Services (CPS</publisher-name>)), <fpage>343</fpage>&#x02013;<lpage>346</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>
