<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Chem.</journal-id>
<journal-title>Frontiers in Chemistry</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Chem.</abbrev-journal-title>
<issn pub-type="epub">2296-2646</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1603948</article-id>
<article-id pub-id-type="doi">10.3389/fchem.2025.1603948</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Chemistry</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Hydrogen-centric machine learning approach for analyzing properties of tricyclic anti-depressant drugs</article-title>
<alt-title alt-title-type="left-running-head">Kour and Ravi Sankar</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fchem.2025.1603948">10.3389/fchem.2025.1603948</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Kour</surname>
<given-names>Simran</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/3078109/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Ravi Sankar</surname>
<given-names>J.</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/3019247/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/Writing - review &#x26; editing/"/>
</contrib>
</contrib-group>
<aff>
<institution>Department of Mathematics</institution>, <institution>School of Advanced Sciences</institution>, <institution>Vellore Institute of Technology</institution>, <addr-line>Vellore</addr-line>, <addr-line>Tamil Nadu</addr-line>, <country>India</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2204496/overview">Dejan Milenkovi&#x107;</ext-link>, University of Kragujevac, Serbia</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1727260/overview">Rahul Pinjari</ext-link>, Swami Ramanand Teerth Marathwada University, India</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1893267/overview">Dapeng Wang</ext-link>, Chinese Academy of Sciences (CAS), China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: J. Ravi Sankar, <email>ravisankar.j@vit.ac.in</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>03</day>
<month>06</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>13</volume>
<elocation-id>1603948</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>04</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>15</day>
<month>05</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Kour and Ravi Sankar.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Kour and Ravi Sankar</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Tricyclic anti-depressant (TCA) drugs are widely used to treat depression, but traditional methods for evaluating their physicochemical properties can be time-consuming and costly. This study examines how topological indices can help to predict the properties of TCA drugs, with a special focus on the role of the hydrogen representation.</p>
</sec>
<sec>
<title>Methods</title>
<p>Two molecular configurations were analyzed: one with only explicit hydrogen and the other including all hydrogen atoms. To assess predictive performance, linear regression (LR) and support vector regression (SVR) models were employed.</p>
</sec>
<sec>
<title>Results</title>
<p>The results showed that adding all hydrogen atoms showed strong correlations, especially for polarizability, molar refractivity, and molar volume. Among the models employed, SVR provided more accurate results. Additionally, hydrogen representation had a stronger impact on SVR&#x0027;s predictions.</p>
</sec>
<sec>
<title>Discussion</title>
<p>These findings highlight the potential of using machine learning techniques in quantitative structure-property relationship (QSPR) models for more efficient and reliable predictions of drug properties.</p>
</sec>
</abstract>
<kwd-group>
<kwd>tricyclic anti-depressant drugs</kwd>
<kwd>topological indices</kwd>
<kwd>QSPR</kwd>
<kwd>linear regression</kwd>
<kwd>support vector regression</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Theoretical and Computational Chemistry</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Mental health disorders are a group of psychiatric conditions that can severely impact an individual&#x2019;s ability to function in everyday environment, resulting in difficulties with daily activities, social connections, and behavioral stability (<xref ref-type="bibr" rid="B14">Ejima et al., 2024</xref>). Conditions such as anxiety, addiction, depression, and bipolar disorder are common, with depression being a particularly pressing public health concern that demands effective treatment options (<xref ref-type="bibr" rid="B23">Kessler et al., 2007</xref>). TCAs rank among the most commonly prescribed medications for depression, with over 25 million prescriptions written annually in the United States. However, despite their effectiveness, TCAs are frequently linked in overdose incidents, with studies showing that they contribute to nearly 25% of overdose-related hospital admissions at a major medical center (<xref ref-type="bibr" rid="B27">Marshall and Forker, 1982</xref>; <xref ref-type="bibr" rid="B37">Vandel et al., 1997</xref>). According to the 2023 NSDUH Report, 22.8% of adults (58.7 million) experienced any mental illness (AMI) in the past year, and 4.5 million adolescents reported a major depressive episode, with 20% also experiencing substance use disorders. Suicide remains a major worry, with 5.0% of adults having serious thoughts about it, 1.4% making plans, and 0.6% attempting suicide (<xref ref-type="bibr" rid="B20">Health and Services, 2023</xref>). These concerning statistics highlight the critical importance of prioritising mental health rehabilitation and preservation. While laboratory-based drug development has played a key role in advancing treatments for mental health, it is often resource intensive (<xref ref-type="bibr" rid="B22">Insel et al., 2013</xref>). Computational modeling and predictive techniques offer promising alternatives that are both cost-effective and resourceful. These approaches not only enhance traditional drug discovery but also provide accessible and effective therapy for those with neuropsychiatric disorders.</p>
<p>The process of drug design and discovery is a complex, time-consuming, and costly. To optimize this process, researchers have increasingly turned to predictive modeling techniques, particularly in resource-limited scenarios or during medical emergencies. One such approach is QSPR modeling, which predict a drug&#x2019;s physicochemical properties based on its molecular structure and descriptors, commonly referred to as topological indices. Chemical graph theory applies the principles of graph theory to chemistry by representing molecules as graphs, where vertices correspond to atoms and edges to chemical bonds (<xref ref-type="bibr" rid="B35">Thapar et al., 2022</xref>). Topological indices, numerical descriptors derived from these graphs, capture critical structural information (<xref ref-type="bibr" rid="B18">Gutman and Polansky, 2012</xref>; <xref ref-type="bibr" rid="B17">Gutman, 2006</xref>). These indices serve as essential tools in QSPR modeling, as they establish mathematical relationships between molecular structures and biological or physicochemical properties, particularly in pharmaceutical research (<xref ref-type="bibr" rid="B1">Abubakar et al., 2024a</xref>). One of the earliest and most well-known topological indices is the Wiener Index, introduced in 1947, originally designed to predict the physical properties of paraffin compounds (<xref ref-type="bibr" rid="B39">Wiener, 1947</xref>). In recent years, topological indices have gained widespread popularity in QSPR studies, offering a cost-effective and time-saving alternative to experimental methods. By enabling researchers to predict key drug properties, identify influential structural features, and optimize drug candidates, these indices play a crucial role in accelerating drug development and reducing reliance on expensive laboratory experiments (<xref ref-type="bibr" rid="B28">Parveen et al., 2022</xref>; <xref ref-type="bibr" rid="B43">Zaman et al., 2023</xref>).</p>
<p>The use of topological indices in pharmaceutical research has significantly increased in recent years, particularly in QSPR studies. But, most QSPR modeling studies primarily focused on classical graph-based topological indices and simple regression models to establish relationships between topological indices and the physical properties of compounds. A topological index that exhibits a strong linear correlation with a physical property is regarded as an effective descriptor for predicting that property (<xref ref-type="bibr" rid="B44">Zaman et al., 2024</xref>; <xref ref-type="bibr" rid="B13">Das et al., 2024</xref>; <xref ref-type="bibr" rid="B4">Arockiaraj et al., 2024</xref>; <xref ref-type="bibr" rid="B21">Huang et al., 2024</xref>; <xref ref-type="bibr" rid="B19">Hasani and Ghods, 2024</xref>). However, when the relationship between topological indices and physical properties is non-linear, more advanced approaches like machine learning, are employed to capture complex patterns and improve predictive accuracy (<xref ref-type="bibr" rid="B16">Fern&#xe1;ndez-Blanco et al., 2013</xref>; <xref ref-type="bibr" rid="B26">Madugula et al., 2021</xref>; <xref ref-type="bibr" rid="B2">Abubakar et al., 2024b</xref>). <xref ref-type="bibr" rid="B41">Zabidi et al. (2021)</xref> applied machine learning to predict HOMO and LUMO, minimizing the need for computationally expensive DFT calculations. Degree-based topological indices were employed in QSPR analysis to establish correlations with these properties and identified Linear Regression with Moment Balaban Indices as the most accurate model. <xref ref-type="bibr" rid="B12">Costa et al. (2020)</xref> proposed explored a novel method for QSAR and QSPR modeling through Molecular Graph Theory, emphasizing molecular fragment contributions. By combining Molecular Graph Theory, SMILES notation, and connection table data, they established an efficient method for fragment identification. Machine learning techniques produced accurate predictive models, and the study introduced Charming QSAR and QSPR, a Python tool designed for property estimation in chemical compounds. <xref ref-type="bibr" rid="B1">Abubakar et al. (2024a)</xref> analyzed neighborhood degree-based topological indices for QSPR modeling of anti-tuberculosis drugs, employing Support Vector Regression (SVR) and comparing it to linear regression. The results demonstrated that SVR as a better predictive tool, enhancing the understanding of the non-linear relationship. Author A and others applied QSPR modeling with neighborhood sum degree topological indices to predict antibacterial drug properties. SVR outperformed linear regression, benefiting from feature selection and hyper-parameter tuning. <xref ref-type="bibr" rid="B27">Marshall and Forker (1982)</xref> investigated an ensemble learning approach for the analysis of mental disorder drugs. Using neighborhood degree-based indices derived from SMILES notations, the study identified optimal indices for predicting key physicochemical properties. Their findings showed the role of ensemble learning in better prediction accuracy, particularly for small datasets.</p>
<p>Additionally, it is important to note that none of the cited studies considered hydrogen atoms in their topological representations, which may neglect to important contribution to molecular properties. Furthermore, the all prior work relied on degree-based topological indices, which capture only local atomic environments. In our earlier study, the predictive power of topological indices for drug properties was explored using regression models along with distance-based indices. However, the influence of hydrogen configuration was not considered (<xref ref-type="bibr" rid="B24">Kour et al., 2024</xref>). In contrast, our present study demonstrated a novel comparison of explicit hydrogen and all hydrogen structures, using distance based indices that effectively capture molecular branching and spatial arrangement. Benchmarking LR and SVR, we observed that SVR provided superior accuracy for non-linear relationships, while LR performed well in strongly linear cases. This novel approach refines QSAR modeling, demonstrating how molecular representation influences predictive accuracy and optimizing regression techniques. The primary objective of this work is to understand the impact of hydrogen configuration on the prediction of six physicochemical properties using two regression techniques. This work aims to evaluate how different molecular representations impact prediction accuracy across multiple properties.</p>
<p>The major contributions in this study are:<list list-type="simple">
<list-item>
<p>
<inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> A comparative assessment of SVR and LR in handling linear and non-linear relationships.</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> A detailed evaluation of six physicochemical properties using both molecular representations.</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> The demonstration of how distance-based topological indices, combined with SVR, can enhance property prediction and serve as practical tool to accelerate early stage drug discovery.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Methodology and data collection</title>
<sec id="s2-1">
<title>2.1 Drugs analysis</title>
<p>This study focuses on fifteen TCA drugs which have different molecular structure and clinical importance. <xref ref-type="table" rid="T1">Table 1</xref> list their chemical structures and therapeutic uses, highlighting their role in treating depression and anxiety.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>TCA drugs with their chemical structure and therapeutic uses.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Drugs</th>
<th align="left">Abbreviation</th>
<th align="left">Chemical structures</th>
<th align="left">Therapeutic uses</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Alprazolam</td>
<td align="left">ALP</td>
<td align="left">C<sub>17</sub> H<sub>13</sub> Cl N<sub>4</sub>
</td>
<td align="left">Used to treat generalized anxiety disorder, panic disorder, and off-label for insomnia, premenstrual syndrome, and depression in adults</td>
</tr>
<tr>
<td align="left">Amitriptyline</td>
<td align="left">AMT</td>
<td align="left">C<sub>20</sub> H<sub>23</sub>&#xa0;N</td>
<td align="left">Used for major depressive disorder, neuro-pathic pain, chronic tension-type headache, migraine prophylaxis in adults, and nocturnal enuresis in children aged <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:mn>6</mml:mn>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> when other treatments fail</td>
</tr>
<tr>
<td align="left">Amoxapine</td>
<td align="left">AMX</td>
<td align="left">C<sub>17</sub> H<sub>16</sub> Cl N<sub>3</sub> O</td>
<td align="left">For relieving depression symptoms in neurotic, reactive, endogenous, and psychotic depression, as well as depression associated with anxiety or agitation</td>
</tr>
<tr>
<td align="left">Buspirone</td>
<td align="left">BSP</td>
<td align="left">C<sub>21</sub> H<sub>31</sub> N<sub>5</sub> O<sub>2</sub>
</td>
<td align="left">Used to manage anxiety disorders or provide short-term relief from anxiety symptoms</td>
</tr>
<tr>
<td align="left">Clomipramine</td>
<td align="left">CLM</td>
<td align="left">C<sub>19</sub> H<sub>23</sub> Cl N<sub>2</sub>
</td>
<td align="left">Used for obsessive-compulsive disorder, related conditions, and off-label for depression, chronic pain, narcolepsy, and autism</td>
</tr>
<tr>
<td align="left">Desipramine</td>
<td align="left">DSP</td>
<td align="left">C<sub>18</sub> H<sub>22</sub> N<sub>2</sub>
</td>
<td align="left">Relieves symptoms of depressive syndromes, particularly endogenous depression, and manages chronic peripheral neuropathic pain, anxiety disorders, and ADHD (second or third-line treatment)</td>
</tr>
<tr>
<td align="left">Desvenlafaxine</td>
<td align="left">DVF</td>
<td align="left">C<sub>16</sub> H<sub>25</sub>&#xa0;N O<sub>2</sub>
</td>
<td align="left">To treat major depressive disorder in adults and is also prescribed off-label for hot flashes in menopausal women</td>
</tr>
<tr>
<td align="left">Diazepam</td>
<td align="left">DZM</td>
<td align="left">C<sub>16</sub> H<sub>13</sub> Cl N<sub>2</sub> O</td>
<td align="left">Used to treat anxiety, muscle spasms, acute alcohol withdrawal, spasticity, and as an adjunct for epilepsy, with indications for short-term anxiety relief, pre-surgical sedation in adults, and specific seizure episodes in children</td>
</tr>
<tr>
<td align="left">Fluoxetine</td>
<td align="left">FLX</td>
<td align="left">C<sub>17</sub> H<sub>18</sub> F<sub>3</sub> N O</td>
<td align="left">Used for major depressive disorder, obsessive-compulsive disorder, bulimia nervosa, acute panic disorder, PMDD, and in combination with olanzapine for Bipolar I Disorder-related and treatment-resistant depression</td>
</tr>
<tr>
<td align="left">Imipramine</td>
<td align="left">IMP</td>
<td align="left">C<sub>19</sub> H<sub>24</sub> N<sub>2</sub>
</td>
<td align="left">Used to relieve depression symptoms and reduce enuresis in children 6&#x2b;, with off-label uses for panic disorders, ADHD, bulimia nervosa, bipolar depression, PTSD, and neuropathic pain</td>
</tr>
<tr>
<td align="left">Lorazepam</td>
<td align="left">LRZ</td>
<td align="left">C<sub>15</sub> H<sub>10</sub> Cl<sub>2</sub> N<sub>2</sub> O<sub>2</sub>
</td>
<td align="left">Used for anxiety relief, sedation, and status epilepticus, with off-label uses for alcohol withdrawal, muscle spasms, insomnia, panic disorder, and more</td>
</tr>
<tr>
<td align="left">Nortriptyline</td>
<td align="left">NTP</td>
<td align="left">C<sub>19</sub> H<sub>21</sub>&#xa0;N</td>
<td align="left">Used to relieve symptoms of major depressive disorder (MDD) and off-label for chronic pain, myofascial pain, neuralgia, and irritable bowel syndrome</td>
</tr>
<tr>
<td align="left">Oxazepam</td>
<td align="left">OZP</td>
<td align="left">C<sub>15</sub> H<sub>11</sub> Cl N<sub>2</sub> O<sub>2</sub>
</td>
<td align="left">Used to manage anxiety disorders, provide short-term anxiety relief, and treat alcohol withdrawal symptoms</td>
</tr>
<tr>
<td align="left">Protriptyline</td>
<td align="left">PTP</td>
<td align="left">C<sub>19</sub> H<sub>21</sub>&#xa0;N</td>
<td align="left">Used for the treatment of major depression</td>
</tr>
<tr>
<td align="left">Trimipramine</td>
<td align="left">TMP</td>
<td align="left">C<sub>20</sub> H<sub>26</sub> N<sub>2</sub>
</td>
<td align="left">Used to treat depression, including cases accompanied by anxiety, agitation, or sleep disturbances</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Two molecular representations, one including explicit hydrogen only and the other including all the hydrogen, were analyzed to understand the influence of hydrogen on the properties of drugs. <xref ref-type="fig" rid="F1">Figure 1</xref> presents an example of Fluoxetine, showing its two different configurations-one with explicit hydrogen and other with all hydrogen. The hydrogen atoms are highlighted in red color. Six physicochemical properties listed in <xref ref-type="table" rid="T2">Table 2</xref>, were obtained from <xref ref-type="bibr" rid="B30">PubChem (2025)</xref> and <xref ref-type="bibr" rid="B10">ChemSpider (2025)</xref>. These properties help us understand the thermodynamic and structural characteristics of these compounds in further computational analyses.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Two configurations of a molecule (Example: Fluoxetine).</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g001.tif"/>
</fig>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Physicochemical properties of the TCA drugs.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Drugs</th>
<th align="center">Boiling point (BP)</th>
<th align="center">Enthalpy (E)</th>
<th align="center">Flash point (FP)</th>
<th align="center">Molar refractivity (MR)</th>
<th align="center">Polarizability (P)</th>
<th align="center">Molar volume (MV)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Alprazolam</td>
<td align="center">509</td>
<td align="center">77.9</td>
<td align="center">261.6</td>
<td align="center">88.2</td>
<td align="center">35</td>
<td align="center">225.6</td>
</tr>
<tr>
<td align="center">Amitriptyline</td>
<td align="center">398.2</td>
<td align="center">64.9</td>
<td align="center">174</td>
<td align="center">91.5</td>
<td align="center">36.3</td>
<td align="center">257.8</td>
</tr>
<tr>
<td align="center">Amoxapine</td>
<td align="center">469.9</td>
<td align="center">73.2</td>
<td align="center">238</td>
<td align="center">86.8</td>
<td align="center">34.4</td>
<td align="center">228.2</td>
</tr>
<tr>
<td align="center">Buspirone</td>
<td align="center">613.9</td>
<td align="center">91.1</td>
<td align="center">325.1</td>
<td align="center">106.8</td>
<td align="center">42.4</td>
<td align="center">310.7</td>
</tr>
<tr>
<td align="center">Clomipramine</td>
<td align="center">434.2</td>
<td align="center">69</td>
<td align="center">216.4</td>
<td align="center">93.8</td>
<td align="center">37.2</td>
<td align="center">281.2</td>
</tr>
<tr>
<td align="center">Desipramine</td>
<td align="center">407.4</td>
<td align="center">65.9</td>
<td align="center">160.5</td>
<td align="center">84.2</td>
<td align="center">33.4</td>
<td align="center">254.3</td>
</tr>
<tr>
<td align="center">Desvenlafaxine</td>
<td align="center">403.8</td>
<td align="center">69.1</td>
<td align="center">193.2</td>
<td align="center">77.8</td>
<td align="center">30.9</td>
<td align="center">236.1</td>
</tr>
<tr>
<td align="center">Diazepam</td>
<td align="center">497.4</td>
<td align="center">76.5</td>
<td align="center">254.6</td>
<td align="center">80.9</td>
<td align="center">32.1</td>
<td align="center">225.9</td>
</tr>
<tr>
<td align="center">Fluoxetine</td>
<td align="center">395.1</td>
<td align="center">64.5</td>
<td align="center">192.8</td>
<td align="center">79.9</td>
<td align="center">31.7</td>
<td align="center">266.7</td>
</tr>
<tr>
<td align="center">Imipramine</td>
<td align="center">403.1</td>
<td align="center">65.4</td>
<td align="center">179.7</td>
<td align="center">88.9</td>
<td align="center">35.3</td>
<td align="center">269.2</td>
</tr>
<tr>
<td align="center">Lorazepam</td>
<td align="center">543.6</td>
<td align="center">86.5</td>
<td align="center">282.6</td>
<td align="center">81</td>
<td align="center">32.1</td>
<td align="center">211.2</td>
</tr>
<tr>
<td align="center">Nortriptyline</td>
<td align="center">403.4</td>
<td align="center">65.5</td>
<td align="center">194.9</td>
<td align="center">86.8</td>
<td align="center">34.4</td>
<td align="center">242.9</td>
</tr>
<tr>
<td align="center">Oxazepam</td>
<td align="center">516.6</td>
<td align="center">83</td>
<td align="center">266.2</td>
<td align="center">76.4</td>
<td align="center">30.3</td>
<td align="center">201.9</td>
</tr>
<tr>
<td align="center">Protriptyline</td>
<td align="center">407.7</td>
<td align="center">66</td>
<td align="center">198.3</td>
<td align="center">84.8</td>
<td align="center">33.6</td>
<td align="center">256.5</td>
</tr>
<tr>
<td align="center">Trimipramine</td>
<td align="center">411.8</td>
<td align="center">66.4</td>
<td align="center">183.3</td>
<td align="center">93.5</td>
<td align="center">37.1</td>
<td align="center">286.1</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2-2">
<title>2.2 Topological indices</title>
<p>The study explores the relationship between molecular properties and atomic arrangements using distance-based topological indices which are presented in <xref ref-type="table" rid="T3">Table 3</xref>. Similarly, calculations were performed for fifteen TCA drugs, analyzed with explicit hydrogen and with all the hydrogen. The results, presented in <xref ref-type="table" rid="T4">Tables 4</xref>, <xref ref-type="table" rid="T5">5</xref>, provide numerical values that represent structural connectivity and molecular topology. These indices were selected due to their ability to capture spatial and str curtal complexity of the molecules. They effectively encode connectivity and branching patterns that influence molecular behavior.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>TCA drugs with their chemical structure and therapeutic uses.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Topological index</th>
<th align="left">Notation</th>
<th align="left">Formula</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Wiener Index <xref ref-type="bibr" rid="B39">Wiener&#xa0;(1947)</xref>
</td>
<td align="left">W(G)</td>
<td align="left">
<inline-formula id="inf5">
<mml:math id="m5">
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">Hyper-Wiener Index <xref ref-type="bibr" rid="B31">Randi&#x107;&#xa0;(1993)</xref>
</td>
<td align="left">WW(G)</td>
<td align="left">
<inline-formula id="inf6">
<mml:math id="m6">
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">Harary Index <xref ref-type="bibr" rid="B29">Plav&#x161;i&#x107; et al.&#xa0;(1993)</xref>
</td>
<td align="left">H(G)</td>
<td align="left">
<inline-formula id="inf7">
<mml:math id="m7">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">Detour Index <xref ref-type="bibr" rid="B25">Lukovits&#xa0;(1996)</xref>
</td>
<td align="left">D(G)</td>
<td align="left">
<inline-formula id="inf8">
<mml:math id="m8">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">Detour Harary Index <xref ref-type="bibr" rid="B15">Fang et al.&#xa0;(2018)</xref>
</td>
<td align="left">DH(G)</td>
<td align="left">
<inline-formula id="inf9">
<mml:math id="m9">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Topological indices of the drugs with explicit hydrogen.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Drugs</th>
<th align="center">W</th>
<th align="center">WW</th>
<th align="center">H</th>
<th align="center">D</th>
<th align="center">DH</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Alprazolam</td>
<td align="center">926</td>
<td align="center">2770</td>
<td align="center">81.4698</td>
<td align="center">2845</td>
<td align="center">24.1903</td>
</tr>
<tr>
<td align="center">Amitriptyline</td>
<td align="center">882</td>
<td align="center">2780</td>
<td align="center">70.7087</td>
<td align="center">2581</td>
<td align="center">26.2091</td>
</tr>
<tr>
<td align="center">Amoxapine</td>
<td align="center">1075</td>
<td align="center">3398</td>
<td align="center">85.8329</td>
<td align="center">3291</td>
<td align="center">25.4936</td>
</tr>
<tr>
<td align="center">Buspirone</td>
<td align="center">2514</td>
<td align="center">13028</td>
<td align="center">102.7360</td>
<td align="center">3764</td>
<td align="center">57.3198</td>
</tr>
<tr>
<td align="center">Clomipramine</td>
<td align="center">995</td>
<td align="center">3194</td>
<td align="center">78.0429</td>
<td align="center">2861</td>
<td align="center">28.6732</td>
</tr>
<tr>
<td align="center">Desipramine</td>
<td align="center">882</td>
<td align="center">2780</td>
<td align="center">70.7087</td>
<td align="center">2581</td>
<td align="center">26.2091</td>
</tr>
<tr>
<td align="center">Desvenlafaxine</td>
<td align="center">897</td>
<td align="center">2846</td>
<td align="center">71.6540</td>
<td align="center">1329</td>
<td align="center">46.9409</td>
</tr>
<tr>
<td align="center">Diazepam</td>
<td align="center">726</td>
<td align="center">2077</td>
<td align="center">69.4905</td>
<td align="center">1901</td>
<td align="center">24.9069</td>
</tr>
<tr>
<td align="center">Fluoxetine</td>
<td align="center">1292</td>
<td align="center">4946</td>
<td align="center">77.8563</td>
<td align="center">1772</td>
<td align="center">54.1895</td>
</tr>
<tr>
<td align="center">Imipramine</td>
<td align="center">882</td>
<td align="center">2780</td>
<td align="center">70.7087</td>
<td align="center">2581</td>
<td align="center">26.2091</td>
</tr>
<tr>
<td align="center">Lorazepam</td>
<td align="center">1034</td>
<td align="center">3076</td>
<td align="center">86.5452</td>
<td align="center">2601</td>
<td align="center">33.9740</td>
</tr>
<tr>
<td align="center">Nortriptyline</td>
<td align="center">882</td>
<td align="center">2780</td>
<td align="center">70.7087</td>
<td align="center">2581</td>
<td align="center">26.2091</td>
</tr>
<tr>
<td align="center">Oxazepam</td>
<td align="center">928</td>
<td align="center">2731</td>
<td align="center">80.6476</td>
<td align="center">2362</td>
<td align="center">30.8960</td>
</tr>
<tr>
<td align="center">Protriptyline</td>
<td align="center">882</td>
<td align="center">2780</td>
<td align="center">70.7087</td>
<td align="center">2581</td>
<td align="center">26.2091</td>
</tr>
<tr>
<td align="center">Trimipramine</td>
<td align="center">979</td>
<td align="center">3081</td>
<td align="center">78.4611</td>
<td align="center">2808</td>
<td align="center">30.3429</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Topological indices of the drugs with all the hydrogen.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Drugs</th>
<th align="center">W</th>
<th align="center">WW</th>
<th align="center">H</th>
<th align="center">D</th>
<th align="center">DH</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Alprazolam</td>
<td align="center">2936</td>
<td align="center">10289</td>
<td align="center">168.8604</td>
<td align="center">7858</td>
<td align="center">68.2119</td>
</tr>
<tr>
<td align="center">Amitriptyline</td>
<td align="center">5252</td>
<td align="center">20371</td>
<td align="center">239.9619</td>
<td align="center">12117</td>
<td align="center">126.6943</td>
</tr>
<tr>
<td align="center">Amoxapine</td>
<td align="center">3564</td>
<td align="center">12697</td>
<td align="center">193.5572</td>
<td align="center">9689</td>
<td align="center">78.4991</td>
</tr>
<tr>
<td align="center">Buspirone</td>
<td align="center">12709</td>
<td align="center">69193</td>
<td align="center">366.6186</td>
<td align="center">18125</td>
<td align="center">244.3614</td>
</tr>
<tr>
<td align="center">Clomipramine</td>
<td align="center">5462</td>
<td align="center">21023</td>
<td align="center">250.8310</td>
<td align="center">12557</td>
<td align="center">134.2503</td>
</tr>
<tr>
<td align="center">Desipramine</td>
<td align="center">4538</td>
<td align="center">16736</td>
<td align="center">226.4617</td>
<td align="center">10943</td>
<td align="center">114.0553</td>
</tr>
<tr>
<td align="center">Desvenlafaxine</td>
<td align="center">4836</td>
<td align="center">17076</td>
<td align="center">248.9564</td>
<td align="center">7008</td>
<td align="center">179.4176</td>
</tr>
<tr>
<td align="center">Diazepam</td>
<td align="center">2497</td>
<td align="center">8380</td>
<td align="center">154.1603</td>
<td align="center">5741</td>
<td align="center">71.3523</td>
</tr>
<tr>
<td align="center">Fluoxetine</td>
<td align="center">4433</td>
<td align="center">17760</td>
<td align="center">198.7027</td>
<td align="center">6065</td>
<td align="center">148.3364</td>
</tr>
<tr>
<td align="center">Imipramine</td>
<td align="center">5462</td>
<td align="center">21023</td>
<td align="center">250.8310</td>
<td align="center">12557</td>
<td align="center">134.2503</td>
</tr>
<tr>
<td align="center">Lorazepam</td>
<td align="center">2142</td>
<td align="center">7027</td>
<td align="center">139.4702</td>
<td align="center">5038</td>
<td align="center">62.4229</td>
</tr>
<tr>
<td align="center">Nortriptyline</td>
<td align="center">4346</td>
<td align="center">16147</td>
<td align="center">216.0927</td>
<td align="center">10521</td>
<td align="center">106.9993</td>
</tr>
<tr>
<td align="center">Oxazepam</td>
<td align="center">2142</td>
<td align="center">7027</td>
<td align="center">139.4702</td>
<td align="center">5038</td>
<td align="center">62.4229</td>
</tr>
<tr>
<td align="center">Protriptyline</td>
<td align="center">4267</td>
<td align="center">15583</td>
<td align="center">218.2697</td>
<td align="center">10307</td>
<td align="center">111.5984</td>
</tr>
<tr>
<td align="center">Trimipramine</td>
<td align="center">6278</td>
<td align="center">24110</td>
<td align="center">279.5762</td>
<td align="center">14063</td>
<td align="center">156.9178</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Where, <inline-formula id="inf10">
<mml:math id="m10">
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> be the distance between the vertices <inline-formula id="inf11">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf12">
<mml:math id="m12">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf13">
<mml:math id="m13">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> be the length of the longest path between the vertices <inline-formula id="inf14">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf15">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The formulas for the topological indices remain the same for all the drugs, but they are calculated in two different ways according to the representation of hydrogen, one with only explicit hydrogen and other with all hydrogen. For preprocessing, all drug structure were standardized and convert into graph-based representations. These calculations are performed using Python and its libraries. RDKit is used to handle the molecular structures, NetworkX helps in creating the adjacency and distance matrices, and NumPy takes care of the numerical operations. The input for the molecular structures is provided in SMILES format&#x2014;a simple text representation of molecules. First, molecules from PubChem are converted into graphs using RDKit. For explicit hydrogen calculation, the skeletal form of the molecule with Chem. MolFromSmiles (smiles) is used. And, for all hydrogen, we add explicit hydrogen atoms with Chem. AddHs(mol) which consider all the hydrogen present in a molecule.</p>
</sec>
<sec id="s2-3">
<title>2.3 Regression model</title>
<sec id="s2-3-1">
<title>2.3.1 QSPR model</title>
<p>QSPR is a computational model which is used to predict the physical, chemical, or biological properties of molecules based on their molecular structure. QSPR models establish a mathematical relationship between molecular descriptor (topological index) and a target property (<xref ref-type="bibr" rid="B36">Todeschini and Consonni, 2009</xref>).</p>
<p>The formulation of QSPR is represented in <xref ref-type="disp-formula" rid="e1">Equation 1</xref>.<disp-formula id="e1">
<mml:math id="m16">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where, <inline-formula id="inf16">
<mml:math id="m17">
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the target property which is a dependent variable, <inline-formula id="inf17">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the topological index, and <inline-formula id="inf18">
<mml:math id="m19">
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the mathematical function.</p>
</sec>
<sec id="s2-3-2">
<title>2.3.2 Linear regression</title>
<p>Linear regression is a method used to establish the relationship between a dependent variable and an independent variable by fitting a straight line (<xref ref-type="bibr" rid="B42">Zaid, 2015</xref>).</p>
<p>The equation is represented in <xref ref-type="disp-formula" rid="e2">Equation 2</xref>.<disp-formula id="e2">
<mml:math id="m20">
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf19">
<mml:math id="m21">
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the dependent variable, <inline-formula id="inf20">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the independent variables, <inline-formula id="inf21">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the regression coefficients, <inline-formula id="inf22">
<mml:math id="m24">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the intercept term, and <inline-formula id="inf23">
<mml:math id="m25">
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the random error.</p>
<p>The following LR model in <xref ref-type="disp-formula" rid="e3">Equation 3</xref> is employed to construct a QSPR model.<disp-formula id="e3">
<mml:math id="m26">
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where, <inline-formula id="inf24">
<mml:math id="m27">
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents target property (dependent variable), <inline-formula id="inf25">
<mml:math id="m28">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents topological index (independent variable), <inline-formula id="inf26">
<mml:math id="m29">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents intercept or constant of regression, and <inline-formula id="inf27">
<mml:math id="m30">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents regression coefficient.</p>
</sec>
<sec id="s2-3-3">
<title>2.3.3 SVR theory</title>
<p>Support Vector Machines (SVM), introduced by Vapnik and others in 1995, are based on the structural risk minimization principle and statistical learning theory (<xref ref-type="bibr" rid="B38">Vapnik, 2013</xref>). SVM has been successfully applied to a wide range of classification and regression problems (<xref ref-type="bibr" rid="B8">Cai et al., 2003b</xref>; <xref ref-type="bibr" rid="B7">Cai et al., 2003a</xref>; <xref ref-type="bibr" rid="B11">Cortes and Vapnik, 1995</xref>; <xref ref-type="bibr" rid="B34">Smola and Sch&#xf6;lkopf, 2004</xref>; <xref ref-type="bibr" rid="B32">Rao and Gopalakrishna, 2009</xref>). When used for regression, they are called support vector regression. Traditionally, QSPR models have relied on LR to predict compound properties because it is simple and interpretable. However, LR struggles with non-linear data and is sensitive to unusual data points. SVR addresses these limitations by effectively capturing non-linear patterns and showing more accurate and reliable predictions. These advantages make SVR a strong tool for combining with topological indices in QSPR studies (<xref ref-type="bibr" rid="B5">Awad and Khanna, 2015</xref>; <xref ref-type="bibr" rid="B3">Ardeshir et al., 2021</xref>; <xref ref-type="bibr" rid="B6">Ba&#x15f;tanlar and &#xd6;zuysal, 2013</xref>; <xref ref-type="bibr" rid="B40">Yang et al., 2005</xref>).</p>
<p>SVR focuses on developing a predictive model between given input features and their target values. Given a training dataset <inline-formula id="inf28">
<mml:math id="m31">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, where each input <inline-formula id="inf29">
<mml:math id="m32">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents a feature vector with dimension <inline-formula id="inf30">
<mml:math id="m33">
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf31">
<mml:math id="m34">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the corresponding target value, the goal is to determine a function <inline-formula id="inf32">
<mml:math id="m35">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> that can accurately map the approximate value of <inline-formula id="inf33">
<mml:math id="m36">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> to <inline-formula id="inf34">
<mml:math id="m37">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>SVR creates a function that is linear in a transformed feature space but can model complex, non-linear relationships in the original input space. This is done by applying a non-linear transformation <inline-formula id="inf35">
<mml:math id="m38">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> to map the input data into a higher-dimensional space where linear regression can be performed effectively.</p>
<p>The regression function is defined in <xref ref-type="disp-formula" rid="e4">Equation 4</xref>.<disp-formula id="e4">
<mml:math id="m39">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mi>&#x3b2;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>q</mml:mi>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>where:<list list-type="simple">
<list-item>
<p>
<inline-formula id="inf36">
<mml:math id="m40">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> <inline-formula id="inf37">
<mml:math id="m41">
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the weight vector that defines the orientation of the regression hyperplane in the feature space,</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf38">
<mml:math id="m42">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> <inline-formula id="inf39">
<mml:math id="m43">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denotes a feature mapping function that projects the input <inline-formula id="inf40">
<mml:math id="m44">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> into a higher-dimensional space,</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf41">
<mml:math id="m45">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> <inline-formula id="inf42">
<mml:math id="m46">
<mml:mrow>
<mml:mi>q</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> serves as a bias term, shifting the hyperplane&#x2019;s position accordingly.</p>
</list-item>
</list>
</p>
<p>The major goal of the SVR model is to find the weight vector <inline-formula id="inf43">
<mml:math id="m47">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and bias <inline-formula id="inf44">
<mml:math id="m48">
<mml:mrow>
<mml:mi>q</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> that minimize a combination of two components: a regularization term, which controls model complexity, and a loss function, which measures the prediction error. The SVR optimization minimizes the objective function <xref ref-type="disp-formula" rid="e5">Equation 5</xref> subject to constraints <xref ref-type="disp-formula" rid="e6">Equations 6</xref>-<xref ref-type="disp-formula" rid="e8">8</xref>.<disp-formula id="e5">
<mml:math id="m49">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>q</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mi>W</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:munderover>
</mml:mstyle>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>subject to:<disp-formula id="e6">
<mml:math id="m50">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22a4;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mi>&#x3b2;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>q</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
<disp-formula id="e7">
<mml:math id="m51">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22a4;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mi>&#x3b2;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>q</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m52">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>where:<list list-type="simple">
<list-item>
<p>
<inline-formula id="inf45">
<mml:math id="m53">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> The term <inline-formula id="inf46">
<mml:math id="m54">
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mi>W</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> serves as a regularization factor, aiding in managing the model&#x2019;s complexity.,</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf47">
<mml:math id="m55">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> The parameter <inline-formula id="inf48">
<mml:math id="m56">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> functions as a regularization parameter, regulating the balance between model complexity and the allowance for deviations beyond <inline-formula id="inf49">
<mml:math id="m57">
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>,</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf50">
<mml:math id="m58">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> The value <inline-formula id="inf51">
<mml:math id="m59">
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> defines the epsilon-insensitive zone (epsilon-tube) where errors within this range are not penalized in the loss function,</p>
</list-item>
<list-item>
<p>
<inline-formula id="inf52">
<mml:math id="m60">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> The slack variables <inline-formula id="inf53">
<mml:math id="m61">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf54">
<mml:math id="m62">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> quantify the extent to which training samples fall outside the epsilon-tube, allowing the model to handle data points that do not fit perfectly within the margin.</p>
</list-item>
</list>
</p>
<p>In this study, radical basis function (RBF) has been implemented. The RBF kernel is a kernel that maps data to a higher dimensional space and is defined in <xref ref-type="disp-formula" rid="e9">Equation 9</xref>.<disp-formula id="e9">
<mml:math id="m63">
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where, <inline-formula id="inf55">
<mml:math id="m64">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is a parameter, which is equal to <inline-formula id="inf56">
<mml:math id="m65">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> (<inline-formula id="inf57">
<mml:math id="m66">
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the free parameter).</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.4 Performance evaluation</title>
<sec id="s2-4-1">
<title>2.4.1 Coefficient of determination <inline-formula id="inf58">
<mml:math id="m67">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>
</title>
<p>In regression analysis, the most commonly used statistic to assess model performance is the coefficient of determination <inline-formula id="inf59">
<mml:math id="m68">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. It indicates how much of the variation in the response variable is explained by the model. The value of <inline-formula id="inf60">
<mml:math id="m69">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> ranges from 0 to 1, where a higher <inline-formula id="inf61">
<mml:math id="m70">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> signifies a better model fit (<xref ref-type="bibr" rid="B9">Cameron and Windmeijer, 1997</xref>).</p>
<p>The formula for calculating R-squared is defined in <xref ref-type="disp-formula" rid="e10">Equation 10</xref>.<disp-formula id="e10">
<mml:math id="m71">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>where, <inline-formula id="inf62">
<mml:math id="m72">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the actual value of the dependent variable, <inline-formula id="inf63">
<mml:math id="m73">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the predicted value from the regression models, and <inline-formula id="inf64">
<mml:math id="m74">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the mean of actual values.</p>
</sec>
<sec id="s2-4-2">
<title>2.4.2 Root mean squared error</title>
<p>The Root Mean Squared Error (RMSE) in the dataset is determined by taking the square root of the mean squared differences between the observed values and predicted values (<xref ref-type="bibr" rid="B5">Awad and Khanna, 2015</xref>; <xref ref-type="bibr" rid="B33">Sharma, 2005</xref>), given in <xref ref-type="disp-formula" rid="e11">Equation 11</xref>.<disp-formula id="e11">
<mml:math id="m75">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>M</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
</mml:mstyle>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>where, <inline-formula id="inf65">
<mml:math id="m76">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the number of the observations, <inline-formula id="inf66">
<mml:math id="m77">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the actual value, and <inline-formula id="inf67">
<mml:math id="m78">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is the predicted value.</p>
</sec>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and discussion</title>
<p>In this study, a QSPR analysis of fifteen TCA drugs has been performed, to understand their physicochemical properties, which play an important role in examining their efficacy, stability, and thermodynamic behavior. The focus is on the correlations between molecular topological indices and physicochemical properties, and the impact of hydrogen atoms. SVR and LR models with explicit hydrogen and all-hydrogen were used to predict the properties based on topological indices and compared their performances to find the more accurate model.</p>
<sec id="s3-1">
<title>3.1 Heatmap analysis</title>
<p>The heatmap analysis, as shown in <xref ref-type="fig" rid="F2">Figures 2</xref>, <xref ref-type="fig" rid="F3">3</xref>, compares the correlation between topological indices and physicochemical properties under two different molecular representations: explicit hydrogen and all hydrogen. The color intensity shows the strength of these relationships, darker colors mean a stronger correlation, while lighter colors mean a weaker one.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Correlation heatmap for TCA drugs with explicit hydrogen.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Correlation heatmap for TCA drugs with all the hydrogen.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g003.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F2">Figure 2</xref>, represents the dataset with explicit hydrogen only, shows a varied correlation pattern. The Harary Index has a strong correlation with Boiling Point (0.8180) and Flash Point (0.8021), but its correlation with Molar Refractivity (0.5291), Polarizability (0.5287), and Molar Volume (0.2315) is weaker. The Wiener and Hyper-Wiener indices show moderate relationships, especially with Molar Refractivity (0.6867) and Polarizability (0.6881). The Detour Harary Index has mostly weak correlations, as indicated by the lighter shades in the heatmap.</p>
<p>In contrast, <xref ref-type="fig" rid="F3">Figure 3</xref>, where all hydrogen atoms are included, the correlation pattern appears more consistent. The Detour Index showed strong correlations for Molar Refractivity (0.9250), Polarizability (0.9253), and Molar Volume (0.8657). The Harary Index, which had strong correlations in <xref ref-type="fig" rid="F1">Figure 1</xref>, now has much weaker correlations to Boiling Point (&#x2212;0.0166) and Flash Point (&#x2212;0.0793). The Detour Harary Index, which mostly has weak correlations, performs better with Molar Volume (0.8426) in this dataset.</p>
</sec>
<sec id="s3-2">
<title>3.2 SVR hyper-parameter tuning</title>
<p>The predictive model was developed using the SVR with the RBF kernel. The model was trained in Python using the scikit-learn library. The dataset was split into 80% training and 20% testing for better accuracy and validation. Hyper-parameter tuning was done to find the best values of the epsilon <inline-formula id="inf68">
<mml:math id="m79">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and cost (C) parameters, with epsilon ranging from 0.1 to 0.5 and C values set at 10, 50, 100, 500. The gamma parameter was adjusted to either &#x201c;scale&#x201d; or &#x201c;auto&#x201d; based on the requirement to achieve optimal results. To make the model more effective, 5-fold cross-validation was employed, where multiple SVR models were trained with different parameter settings. The best SVR model was trained using the optimal parameters and evaluated on the test dataset. The hyper-parameter tuning process was done separately for each physicochemical property, testing five different topological indices. The best index for each property was chosen based on the highest test <inline-formula id="inf69">
<mml:math id="m80">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> value. This tuning process helped identify the best SVR models, leading to more accurate predictions and a stronger QSPR analysis. The final results, displayed in <xref ref-type="table" rid="T6">Table 6</xref>, show the best hyper-parameter values.</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Hyper-parameter tuning.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="center">Property</th>
<th colspan="3" align="center">With explicit hydrogen</th>
<th colspan="3" align="center">With all hydrogen</th>
</tr>
<tr>
<th align="center">C</th>
<th align="center">
<inline-formula id="inf70">
<mml:math id="m81">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3f5;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>
</th>
<th align="center">Gamma</th>
<th align="center">C</th>
<th align="center">
<inline-formula id="inf71">
<mml:math id="m82">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3f5;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>
</th>
<th align="center">Gamma</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Boiling Point</td>
<td align="center">50</td>
<td align="center">0.1</td>
<td align="center">Auto</td>
<td align="center">500</td>
<td align="center">0.5</td>
<td align="center">Scale</td>
</tr>
<tr>
<td align="center">Enthalpy</td>
<td align="center">10</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
<td align="center">100</td>
<td align="center">0.1</td>
<td align="center">Scale</td>
</tr>
<tr>
<td align="center">Flash Point</td>
<td align="center">50</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
<td align="center">500</td>
<td align="center">0.2</td>
<td align="center">Scale</td>
</tr>
<tr>
<td align="center">Molar Refractivity</td>
<td align="center">10</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
<td align="center">10</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
</tr>
<tr>
<td align="center">Polarizability</td>
<td align="center">10</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
<td align="center">500</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
</tr>
<tr>
<td align="center">Molar Volume</td>
<td align="center">50</td>
<td align="center">0.5</td>
<td align="center">Auto</td>
<td align="center">50</td>
<td align="center">0.1</td>
<td align="center">Scale</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>3.3 Performance comparison: LR vs. SVR</title>
<p>In <xref ref-type="table" rid="T7">Tables 7</xref>, <xref ref-type="table" rid="T8">8</xref>, <inline-formula id="inf72">
<mml:math id="m83">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and RMSE values for LR and SVR models are presented for comparison. A higher <inline-formula id="inf73">
<mml:math id="m84">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> indicates better accuracy and lower RMSE indicates few errors. <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref> use bar graphs to visually compare these results. The findings suggest that SVR generally performs better, especially in all hydrogen model. However, LR showed better results for molar refractivity and polarizability, where it achieved much higher <inline-formula id="inf74">
<mml:math id="m85">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> values despite SVR having slightly lower RMSEs but poor <inline-formula id="inf75">
<mml:math id="m86">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> scores.</p>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>Comparison for configuration with explicit hydrogen only.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Model</th>
<th align="center">Property</th>
<th align="center">Best TI</th>
<th align="center">
<inline-formula id="inf76">
<mml:math id="m87">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>2</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>
</th>
<th align="center">RMSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="6" align="center">SVR</td>
<td align="center">Boiling Point</td>
<td align="center">Harary Index</td>
<td align="center">0.9681</td>
<td align="center">8.8991</td>
</tr>
<tr>
<td align="center">Enthalpy</td>
<td align="center">Harary Index</td>
<td align="center">0.9853</td>
<td align="center">0.7121</td>
</tr>
<tr>
<td align="center">Flash Point</td>
<td align="center">Harary Index</td>
<td align="center">0.9441</td>
<td align="center">8.4131</td>
</tr>
<tr>
<td align="center">Molar Refractivity</td>
<td align="center">Detour Index</td>
<td align="center">0.1</td>
<td align="center">2.9608</td>
</tr>
<tr>
<td align="center">Polarizability</td>
<td align="center">Detour Index</td>
<td align="center">0.1</td>
<td align="center">0.8141</td>
</tr>
<tr>
<td align="center">Molar Volume</td>
<td align="center">Harary Index</td>
<td align="center">0.6308</td>
<td align="center">10.8922</td>
</tr>
<tr>
<td rowspan="6" align="center">LR</td>
<td align="center">Boiling Point</td>
<td align="center">Harary Index</td>
<td align="center">0.6692</td>
<td align="center">37.3973</td>
</tr>
<tr>
<td align="center">Enthalpy</td>
<td align="center">Harary Index</td>
<td align="center">0.6223</td>
<td align="center">5.1795</td>
</tr>
<tr>
<td align="center">Flash Point</td>
<td align="center">Harary Index</td>
<td align="center">0.6433</td>
<td align="center">27.3783</td>
</tr>
<tr>
<td align="center">Molar Refractivity</td>
<td align="center">Detour Index</td>
<td align="center">0.6303</td>
<td align="center">4.5457</td>
</tr>
<tr>
<td align="center">Polarizability</td>
<td align="center">Detour Index</td>
<td align="center">0.6253</td>
<td align="center">1.8190</td>
</tr>
<tr>
<td align="center">Molar Volume</td>
<td align="center">Hyper-Wiener Index</td>
<td align="center">0.3662</td>
<td align="center">22.9301</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T8" position="float">
<label>TABLE 8</label>
<caption>
<p>Comparison for configuration with all hydrogen.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Model</th>
<th align="center">Property</th>
<th align="center">Best TI</th>
<th align="center">
<inline-formula id="inf77">
<mml:math id="m88">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>2</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>
</th>
<th align="center">RMSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="6" align="center">SVR</td>
<td align="center">Boiling Point</td>
<td align="center">Detour Harary Index</td>
<td align="center">0.9861</td>
<td align="center">5.8774</td>
</tr>
<tr>
<td align="center">Enthalpy</td>
<td align="center">Detour Harary Index</td>
<td align="center">0.9754</td>
<td align="center">0.9197</td>
</tr>
<tr>
<td align="center">Flash Point</td>
<td align="center">Detour Harary Index</td>
<td align="center">0.9469</td>
<td align="center">8.1953</td>
</tr>
<tr>
<td align="center">Molar Refractivity</td>
<td align="center">Harary Index</td>
<td align="center">0.1</td>
<td align="center">2.7451</td>
</tr>
<tr>
<td align="center">Polarizability</td>
<td align="center">Harary Index</td>
<td align="center">0.1</td>
<td align="center">0.8444</td>
</tr>
<tr>
<td align="center">Molar Volume</td>
<td align="center">Detour Harary Index</td>
<td align="center">0.8659</td>
<td align="center">6.5639</td>
</tr>
<tr>
<td rowspan="6" align="center">LR</td>
<td align="center">Boiling Point</td>
<td align="center">Hyper-Wiener Index</td>
<td align="center">0.1432</td>
<td align="center">60.1869</td>
</tr>
<tr>
<td align="center">Enthalpy</td>
<td align="center">Hyper-Wiener Index</td>
<td align="center">0.0938</td>
<td align="center">8.0226</td>
</tr>
<tr>
<td align="center">Flash Point</td>
<td align="center">Hyper-Wiener Index</td>
<td align="center">0.1019</td>
<td align="center">43.4411</td>
</tr>
<tr>
<td align="center">Molar Refractivity</td>
<td align="center">Detour Index</td>
<td align="center">0.8557</td>
<td align="center">2.8399</td>
</tr>
<tr>
<td align="center">Polarizability</td>
<td align="center">Detour Index</td>
<td align="center">0.8562</td>
<td align="center">1.1269</td>
</tr>
<tr>
<td align="center">Molar Volume</td>
<td align="center">Harary Index</td>
<td align="center">0.8219</td>
<td align="center">12.1539</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Comparison of <inline-formula id="inf78">
<mml:math id="m89">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> values.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g004.tif"/>
</fig>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Comparison of RMSE values.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g005.tif"/>
</fig>
</sec>
<sec id="s3-4">
<title>3.4 Comparison of actual vs. predicted values</title>
<p>It is observed from the earlier results that SVR is better than LR in term of prediction. In this section, comparison of actual values of drug properties, along with the predicted values from the SVR model and LR model, using both the explicit hydrogen and all-hydrogen are presented. Overall, the predicted values follow the actual trends closely, demonstrating the strong predictive capability of SVR for most cases.</p>
<p>
<inline-formula id="inf79">
<mml:math id="m90">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> For boiling point (<xref ref-type="fig" rid="F6">Figure 6</xref>), the all-hydrogen model with SVR closely matches actual values, especially for Amoxapine, Buspirone, Desipramine, Desvenlafaxine, and Diazepam, where the explicit hydrogen model shows large errors. For Clomipramine and Oxazepam, both explicit hydrogen and all hydrogen models with SVR perform equally. In the case of Nortriptyline, both SVR and LR work well but only with the explicit hydrogen model.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Actual vs. predicted values for boiling point.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g006.tif"/>
</fig>
<p>
<inline-formula id="inf80">
<mml:math id="m91">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> For enthalpy (<xref ref-type="fig" rid="F7">Figure 7</xref>), SVR with both explicit hydrogen and all hydrogen predicts accurately for most drugs, except Alprazolam, Clomipramine, Fluxoetine, and Lorazepam. However, for Buspirone, SVR with all hydrogen and LR with explicit hydrogen both models work similarly.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Actual vs. predicted values for enthalpy.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g007.tif"/>
</fig>
<p>
<inline-formula id="inf81">
<mml:math id="m92">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> For flash point (<xref ref-type="fig" rid="F8">Figure 8</xref>), the all-hydrogen model with SVR performs best for most drugs, but the explicit hydrogen model is just as effective for Amoxapine and Oxazepam. The explicit hydrogen model with LR also performed well for few drugs like Buspirone, Desvenlafaxine and Nortriptyline.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Actual vs. predicted values for flash point.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g008.tif"/>
</fig>
<p>
<inline-formula id="inf82">
<mml:math id="m93">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> For molar refractivity (<xref ref-type="fig" rid="F9">Figure 9</xref>), SVR with both explicit hydrogen and all hydrogen model performed well for most of the drugs, except Alprazolam, Buspirone and Imipramine. However, LR with all-hydrogen also showed good accuracy for Amitriptyline, Amoxapine, Fluoxetine, Oxazepam, while for Nortriptyline, LR worked well with both models.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Actual vs. predicted values for molar refractivity.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g009.tif"/>
</fig>
<p>
<inline-formula id="inf83">
<mml:math id="m94">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> For polarizability (<xref ref-type="fig" rid="F10">Figure 10</xref>), the all-hydrogen model provides good accuracy with SVR and LR, while the explicit hydrogen did not perform well with most of the drugs. However, for Amoxapine and Nortriptyline, the explicit hydrogen model still works well.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Actual vs. predicted values for polarizability.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g010.tif"/>
</fig>
<p>
<inline-formula id="inf84">
<mml:math id="m95">
<mml:mrow>
<mml:mo>&#x2022;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> For molar volume (<xref ref-type="fig" rid="F11">Figure 11</xref>), SVR with both explicit and all-hydrogen models performs well, but the explicit hydrogen model is more accurate. LR with the all-hydrogen model also shows good performance for a few drugs.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Actual vs. predicted values for molar volume.</p>
</caption>
<graphic xlink:href="fchem-13-1603948-g011.tif"/>
</fig>
<p>These results show that molecular representation plays a key role in prediction accuracy. Both models perform well, but SVR with all-hydrogen model consistently provides more precise and accurate predictions, making it a better option for predicting drug properties. LR shows moderate performance, with occasional improvements when paired with explicit hydrogen models. This study underscores how including hydrogen in molecular structures enhances prediction accuracy.</p>
<p>It is evident that the SVR model generally outperformed the LR model, achieving higher <inline-formula id="inf85">
<mml:math id="m96">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and RMSE lower values, especially in capturing non-linear relationships through its use of kernel functions. This allows SVR to capture complex patterns in understanding the properties of the drugs. However, for certain properties, the LR model performs comparably or even better, suggesting linear correlations in those specific cases. SVR works well with small datasets, reducing the chance of overfitting while still giving strong predictions. Overall, because most of the data is non-linear and the drug structures vary, SVR is the better choice for predicting the physicochemical properties of the TCA drugs. However, LR can still be useful for properties that show a more linear relationship.</p>
<p>The findings of this study carry significant practical implications and offer a clear path for improving predictive modeling in drug discovery. It demonstrated the value of using machine learning techniques with chemically informative indices. Accurate prediction of properties is essential in the early stages of pharmaceutical development, where reliable estimation can guide compound selection. Such a framework can significantly reduce the need of time and cost experiment. Moreover, the ability to predict multiple molecular properties with high accuracy supports faster decision making in structure-activity relationship analysis. Overall, the proposed methodology offers a data-driven tool for better drug discovery pipeline by streamlining the evaluation of molecular drugs based on their predicted properties.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s4">
<title>4 Conclusion</title>
<p>While SVR outperformed LR in most cases, LR also demonstrated strong performance for certain properties like molar refractivity and polarizability, making it valuable for understanding linear relationships between topological indices and specific molecular properties. These findings highlight the advantages of combining machine learning with topological indices for better drug property predictions and can guide future research and development of anti-depressant compounds. This study aims to predict the physicochemical properties of TCA drugs using a QSPR model that combines distance-based topological indices, SVR, and a traditional LR model. Two molecular configurations were analyzed: one with explicit hydrogen only and the other including all hydrogen. The results showed that including all hydrogen atoms led to stronger correlations, especially for properties like polarizability, molar refractivity, and molar volume. SVR outperformed LR in most of the cases, showing higher <inline-formula id="inf86">
<mml:math id="m97">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> values and lower RMSE. This highlights that SVR is better at making predictions, notably when dealing with small-sized datasets. Hyper-parameter tuning played a key role in improving accuracy, making SVR a strong choice for predicting TCA drug properties.</p>
<p>In conclusion, adding all hydrogen atoms and using SVR has shown to be an effective approach for predicting the physicochemical properties of TCA drugs. It also helps in understanding the relationship between distance-based topological indices and molecular properties. While SVR outperformed LR in most cases, LR still worked well for some properties, such as molar refractivity and polarizability. This makes LR useful for understanding simple linear relationships between topological indices and specific molecular properties. These findings highlight the advantages of combining machine learning with topological indices for better drug property predictions and can guide future research and development of anti-depressants compounds. Future work could explore additional molecular configurations and different modeling techniques could to predictions more precise.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>All molecular structures, their properties, and their topological indices used in this study are provided as part of the main manuscript. Scripts and software for data analysis are available at the following repository: <ext-link ext-link-type="uri" xlink:href="https://github.com/simran2410/tca_data.git">https://github.com/simran2410/tca_data.git</ext-link>. All necessary files and metadata are included to ensure reproducibility of the study.</p>
</sec>
<sec sec-type="author-contributions" id="s6">
<title>Author contributions</title>
<p>SK: Conceptualization, Formal Analysis, Methodology, Software, Validation, Writing &#x2013; original draft. JRS: Conceptualization, Supervision, Visualization, Writing &#x2013; review and editing.</p>
</sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<ack>
<p>The authors would like to take this opportunity to thank the management of Vellore Institute of Technology (VIT), Vellore, Tamil Nadu, India, for providing the necessary facilities and encouragement to carry out this work.</p>
</ack>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s9">
<title>Generative AI statement</title>
<p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abubakar</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Aremu</surname>
<given-names>K. O.</given-names>
</name>
<name>
<surname>Aphane</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Amusa</surname>
<given-names>L. B.</given-names>
</name>
</person-group> (<year>2024a</year>). <article-title>A qspr analysis of physical properties of antituberculosis drugs using neighbourhood degree-based topological indices and support vector regression</article-title>. <source>Heliyon</source> <volume>10</volume>, <fpage>e28260</fpage>. <pub-id pub-id-type="doi">10.1016/j.heliyon.2024.e28260</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abubakar</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Ejima</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Sanusi</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Ibrahim</surname>
<given-names>A. H.</given-names>
</name>
<name>
<surname>Aremu</surname>
<given-names>K. O.</given-names>
</name>
</person-group> (<year>2024b</year>). <article-title>Predicting antibacterial drugs properties using graph topological indices and machine learning</article-title>. <source>IEEE Access</source> <volume>12</volume>, <fpage>181420</fpage>&#x2013;<lpage>181435</lpage>. <pub-id pub-id-type="doi">10.1109/access.2024.3503760</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ardeshir</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Sanford</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hsu</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Support vector machines and linear regression coincide with very high-dimensional features</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>34</volume>, <fpage>4907</fpage>&#x2013;<lpage>4918</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arockiaraj</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Greeni</surname>
<given-names>A. B.</given-names>
</name>
<name>
<surname>Kalaam</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Aziz</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Alharbi</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Mathematical modeling for prediction of physicochemical characteristics of cardiovascular drugs via modified reverse degree topological indices</article-title>. <source>Eur. Phys. J. E</source> <volume>47</volume>, <fpage>53</fpage>. <pub-id pub-id-type="doi">10.1140/epje/s10189-024-00446-3</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Awad</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Khanna</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2015</year>). <source>Efficient learning machines: theories, concepts, and applications for engineers and system designers</source>. <publisher-name>Springer nature</publisher-name>.</citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ba&#x15f;tanlar</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>&#xd6;zuysal</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Introduction to machine learning</article-title>,&#x201d; in <source>miRNomics: MicroRNA biology and computational analysis</source>, <fpage>105</fpage>&#x2013;<lpage>128</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>Z. L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y. Z.</given-names>
</name>
</person-group> (<year>2003a</year>). <article-title>Svm-prot: web-based support vector machine software for functional classification of a protein from its primary sequence</article-title>. <source>Nucleic acids Res.</source> <volume>31</volume>, <fpage>3692</fpage>&#x2013;<lpage>3697</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkg600</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname>
<given-names>C.-z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.-L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.-z.</given-names>
</name>
</person-group> (<year>2003b</year>). <article-title>Support vector machine classification of physical and biological datasets</article-title>. <source>Int. J. Mod. Phys. C</source> <volume>14</volume>, <fpage>575</fpage>&#x2013;<lpage>585</lpage>. <pub-id pub-id-type="doi">10.1142/s0129183103004759</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cameron</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Windmeijer</surname>
<given-names>F. A.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>An r-squared measure of goodness of fit for some common nonlinear regression models</article-title>. <source>J. Econ.</source> <volume>77</volume>, <fpage>329</fpage>&#x2013;<lpage>342</lpage>. <pub-id pub-id-type="doi">10.1016/s0304-4076(96)01818-0</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<collab>ChemSpider</collab> (<year>2025</year>). <article-title>ChemSpider database</article-title>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cortes</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Vapnik</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Support-vector networks</article-title>. <source>Mach. Learn.</source> <volume>20</volume>, <fpage>273</fpage>&#x2013;<lpage>297</lpage>. <pub-id pub-id-type="doi">10.1007/bf00994018</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Costa</surname>
<given-names>P. C.</given-names>
</name>
<name>
<surname>Evangelista</surname>
<given-names>J. S.</given-names>
</name>
<name>
<surname>Leal</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Miranda</surname>
<given-names>P. C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Chemical graph theory for property modeling in qsar and qspr&#x2014;charming qsar and qspr</article-title>. <source>Mathematics</source> <volume>9</volume>, <fpage>60</fpage>. <pub-id pub-id-type="doi">10.3390/math9010060</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Das</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Rai</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>On topological indices of molnupiravir and its qspr modelling with some other antiviral drugs to treat covid-19 patients</article-title>. <source>J. Math. Chem.</source> <volume>62</volume>, <fpage>2581</fpage>&#x2013;<lpage>2624</lpage>. <pub-id pub-id-type="doi">10.1007/s10910-023-01518-z</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ejima</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Abubakar</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pawa</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>Ibrahim</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Aremu</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Ensemble learning and graph topological indices for predicting physical properties of mental disorder drugs</article-title>. <source>Phys. Scr.</source> <volume>99</volume>, <fpage>106009</fpage>. <pub-id pub-id-type="doi">10.1088/1402-4896/ad79a4</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>W.-H.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.-B.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>F.-Y.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>Z.-M.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>Z.-J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Maximum detour-harary index for some graph classes</article-title>. <source>Symmetry</source> <volume>10</volume>, <fpage>608</fpage>. <pub-id pub-id-type="doi">10.3390/sym10110608</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fern&#xe1;ndez-Blanco</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Aguiar-Pulido</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Munteanu</surname>
<given-names>C. R.</given-names>
</name>
<name>
<surname>Dorado</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Random forest classification based on star graph topological indices for antioxidant proteins</article-title>. <source>J. Theor. Biol.</source> <volume>317</volume>, <fpage>331</fpage>&#x2013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2012.10.006</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gutman</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Chemical graph theory&#x2014;the mathematical connection</article-title>. <source>Adv. Quantum Chem.</source> <volume>51</volume>, <fpage>125</fpage>&#x2013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1016/s0065-3276(06)51003-2</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gutman</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Polansky</surname>
<given-names>O. E.</given-names>
</name>
</person-group> (<year>2012</year>). <source>Mathematical concepts in organic chemistry</source>. <publisher-name>Springer Science and Business Media</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hasani</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ghods</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Predicting the physicochemical properties of drugs for the treatment of Parkinson&#x2019;s disease using topological indices and matlab programming</article-title>. <source>Mol. Phys.</source> <volume>122</volume>, <fpage>e2270082</fpage>. <pub-id pub-id-type="doi">10.1080/00268976.2023.2270082</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<collab>Health and Services, H</collab>. (<year>2023</year>). <article-title>Samhsa releases annual national survey on drug use and health</article-title>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jahanbani</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zuo</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Investigation molecular structure of anticancer drug with topological indices</article-title>. <source>Comput. Biol. Med.</source> <volume>179</volume>, <fpage>108806</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2024.108806</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Insel</surname>
<given-names>T. R.</given-names>
</name>
<name>
<surname>Voon</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Nye</surname>
<given-names>J. S.</given-names>
</name>
<name>
<surname>Brown</surname>
<given-names>V. J.</given-names>
</name>
<name>
<surname>Altevogt</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bullmore</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Innovative solutions to novel drug development in mental health</article-title>. <source>Neurosci. and Biobehav. Rev.</source> <volume>37</volume>, <fpage>2438</fpage>&#x2013;<lpage>2444</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2013.03.022</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kessler</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Angermeyer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Anthony</surname>
<given-names>J. C.</given-names>
</name>
<name>
<surname>De Graaf</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Demyttenaere</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Gasquet</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>Lifetime prevalence and age-of-onset distributions of mental disorders in the world health organization&#x2019;s world mental health survey initiative</article-title>. <source>World psychiatry</source> <volume>6</volume>, <fpage>168</fpage>&#x2013;<lpage>176</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kour</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>J.</surname>
<given-names>R. S.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Machine learning regression models for predicting anti-cancer drug properties: insights from topological indices in qspr analysis</article-title>. <source>Contemp. Math.</source>, <fpage>6515</fpage>&#x2013;<lpage>6526</lpage>. <pub-id pub-id-type="doi">10.37256/cm.5420245826</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lukovits</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>The detour index</article-title>. <source>Croat. Chem. acta</source> <volume>69</volume>, <fpage>873</fpage>&#x2013;<lpage>882</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Madugula</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>John</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nagamani</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gaur</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Poroikov</surname>
<given-names>V. V.</given-names>
</name>
<name>
<surname>Sastry</surname>
<given-names>G. N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Molecular descriptor analysis of approved drugs using unsupervised learning for drug repurposing</article-title>. <source>Comput. Biol. Med.</source> <volume>138</volume>, <fpage>104856</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104856</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marshall</surname>
<given-names>J. B.</given-names>
</name>
<name>
<surname>Forker</surname>
<given-names>A. D.</given-names>
</name>
</person-group> (<year>1982</year>). <article-title>Cardiovascular effects of tricyclic antidepressant drugs: therapeutic usage, overdose, and management of complications</article-title>. <source>Am. heart J.</source> <volume>103</volume>, <fpage>401</fpage>&#x2013;<lpage>414</lpage>. <pub-id pub-id-type="doi">10.1016/0002-8703(82)90281-2</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Parveen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hassan Awan</surname>
<given-names>N. U.</given-names>
</name>
<name>
<surname>Mohammed</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Farooq</surname>
<given-names>F. B.</given-names>
</name>
<name>
<surname>Iqbal</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Topological indices of novel drugs used in diabetes treatment and their qspr modeling</article-title>. <source>J. Math.</source> <volume>2022</volume>, <fpage>5209329</fpage>. <pub-id pub-id-type="doi">10.1155/2022/5209329</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Plav&#x161;i&#x107;</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Nikoli&#x107;</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Trinajsti&#x107;</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Mihali&#x107;</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1993</year>). <article-title>On the harary index for the characterization of chemical graphs</article-title>. <source>J. Math. Chem.</source> <volume>12</volume>, <fpage>235</fpage>&#x2013;<lpage>250</lpage>. <pub-id pub-id-type="doi">10.1007/bf01164638</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<collab>PubChem</collab> (<year>2025</year>). <article-title>PubChem database</article-title>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Randi&#x107;</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1993</year>). <article-title>Novel molecular descriptor for structure&#x2014;property studies</article-title>. <source>Chem. Phys. Lett.</source> <volume>211</volume>, <fpage>478</fpage>&#x2013;<lpage>483</lpage>. <pub-id pub-id-type="doi">10.1016/0009-2614(93)87094-j</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rao</surname>
<given-names>B. V.</given-names>
</name>
<name>
<surname>Gopalakrishna</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Hardgrove grindability index prediction using support vector regression</article-title>. <source>Int. J. Mineral Process.</source> <volume>91</volume>, <fpage>55</fpage>&#x2013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1016/j.minpro.2008.12.003</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sharma</surname>
<given-names>A. K.</given-names>
</name>
</person-group> (<year>2005</year>). <source>Text book of correlations and regression</source>. <publisher-loc>New Delhi</publisher-loc>: <publisher-name>Discovery Publishing House</publisher-name>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Smola</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Sch&#xf6;lkopf</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>A tutorial on support vector regression</article-title>. <source>Statistics Comput.</source> <volume>14</volume>, <fpage>199</fpage>&#x2013;<lpage>222</lpage>. <pub-id pub-id-type="doi">10.1023/b:stco.0000035301.49549.88</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thapar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Eyre</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Brent</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Depression in young people</article-title>. <source>Lancet</source> <volume>400</volume>, <fpage>617</fpage>&#x2013;<lpage>631</lpage>. <pub-id pub-id-type="doi">10.1016/s0140-6736(22)01012-1</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Todeschini</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Consonni</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2009</year>). <source>Molecular descriptors for chemoinformatics: volume I: alphabetical listing/volume II: appendices, references</source>. <publisher-name>John Wiley and Sons</publisher-name>.</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vandel</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Bonin</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Leveque</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Sechter</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bizouard</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Tricyclic antidepressant-induced extrapyramidal side effects</article-title>. <source>Eur. Neuropsychopharmacol.</source> <volume>7</volume>, <fpage>207</fpage>&#x2013;<lpage>212</lpage>. <pub-id pub-id-type="doi">10.1016/s0924-977x(97)00405-7</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vapnik</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2013</year>). <source>The nature of statistical learning theory</source>. <publisher-name>Springer science and business media</publisher-name>.</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wiener</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1947</year>). <article-title>Structural determination of paraffin boiling points</article-title>. <source>J. Am. Chem. Soc.</source> <volume>69</volume>, <fpage>17</fpage>&#x2013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1021/ja01193a005</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Support vector regression based qspr for the prediction of some physicochemical properties of alkyl benzenes</article-title>. <source>J. Mol. Struct. THEOCHEM</source> <volume>719</volume>, <fpage>119</fpage>&#x2013;<lpage>127</lpage>. <pub-id pub-id-type="doi">10.1016/j.theochem.2004.10.060</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zabidi</surname>
<given-names>Z. M.</given-names>
</name>
<name>
<surname>Alias</surname>
<given-names>A. N.</given-names>
</name>
<name>
<surname>Zakaria</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Mahmud</surname>
<given-names>Z. S.</given-names>
</name>
<name>
<surname>Ali</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yaakob</surname>
<given-names>M. K.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Machine learning predictor models in the electronic properties of alkanes based on degree-topology indices</article-title>. <source>Int. J. Emerg. Technol. Adv. Eng.</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.46338/ijetae1121_01</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Zaid</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Correlation and regression analysis, the statistical, economic and social research and training centre for islamic countries (sesric)</article-title>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.oicstatcom.org/file/TEXTBOOK-CORRELATION-AND-REGRESSION-ANALYSIS-EGYPTEN.pdf">http://www.oicstatcom.org/file/TEXTBOOK-CORRELATION-AND-REGRESSION-ANALYSIS-EGYPTEN.pdf</ext-link> (Accessed August 28, 2020)</comment>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zaman</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ahmed</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Sakeena</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rasool</surname>
<given-names>K. B.</given-names>
</name>
<name>
<surname>Ashebo</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Mathematical modeling and topological graph description of dominating david derived networks based on edge partitions</article-title>. <source>Sci. Rep.</source> <volume>13</volume>, <fpage>15159</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-023-42340-6</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zaman</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yaqoob</surname>
<given-names>H. S. A.</given-names>
</name>
<name>
<surname>Ullah</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sheikh</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Qspr analysis of some novel drugs used in blood cancer treatment via degree based topological indices and regression models</article-title>. <source>Polycycl. Aromat. Compd.</source> <volume>44</volume>, <fpage>2458</fpage>&#x2013;<lpage>2474</lpage>. <pub-id pub-id-type="doi">10.1080/10406638.2023.2217990</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>