Predicting the Risk of Diabetes and Heart Disease with Machine Learning Classifiers: The Mediation Analysis
Saved in:
| Title: | Predicting the Risk of Diabetes and Heart Disease with Machine Learning Classifiers: The Mediation Analysis |
|---|---|
| Language: | English |
| Authors: | Ajay Verma (ORCID |
| Source: | Measurement: Interdisciplinary Research and Perspectives. 2025 23(3):310-327. |
| Availability: | Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals |
| Peer Reviewed: | Y |
| Page Count: | 18 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Diabetes, Heart Disorders, Risk, Prediction, Classification, Artificial Intelligence, Path Analysis, Foreign Countries |
| Geographic Terms: | India |
| DOI: | 10.1080/15366367.2024.2347811 |
| ISSN: | 1536-6367 1536-6359 |
| Abstract: | Purpose: This research employs machine learning and mediation analysis, along with path analysis, to investigate the correlations between factors such as body mass index (BMI) and the occurrence of diabetes and heart disease among the Indian population. The objective is to enhance models that are specifically designed to accommodate lifestyles, genetic differences, and healthcare obstacles. Methods: Our research combines a range of data that includes aspects such as lifestyle, physical health, and mental well-being. We use mediation and path analysis techniques to identify the factors involved in the process, while also utilizing machine learning classifiers to enhance risk assessment. In addition to considering known risks, we also investigate biomarkers. Incorporate time factors through analyses. Results: Mediation and path models analyze that diabetes and heart disease are partially mediated with their coefficients a = 7.85, b = 0.01, and c-c' = 0.10. In the path analysis model, the standardized values of exposure and outcome variables are 4.14 and 6.85, respectively, showing a significant relationship with the mediator and other covariates. In classification, the Random Forest classifier shows 99% accuracy and precession, while the Decision Tree, Extra Tree, K-Nearest, and Adaboost classifiers have an accuracy of 98%, 97%, 96%, and 95%, which shows that the machine learning classifiers are more significant for the study. Conclusion: This study contributes to the development of risk management for diabetes and heart disease in India by utilizing machine learning and mediation analysis. It examines relationships, such as BMI, to provide insights for targeted measures, thereby contributing to global discussions on health. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1477877 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFn5xSqtg-7qePG-E-2ANmGAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDKVoMMXgZYXY5hG_uQIBEICBm92b6ETa0f8VdzjyPxStnm7P1Ktacki2U8XF39QDxt3L8ca6KZnVDwmk_tvcX4f5rTfRnXTl13D0x7l_zW49Bwp3flV5xDR3CiouZyKsOgjEEL_BUxPqAwJbbKFa9v7FWX5j4p4EqufE8vI1ivCU_ra9oWXKebxJs5lUZoeMfWwmHUDQIcraBGBksYCk4UtJB_N4WNTBSv3KAWQi Text: Availability: 1 Value: <anid>AN0186774769;p6i01jul.25;2025Jul23.02:58;v2.2.500</anid> <title id="AN0186774769-1">Predicting the Risk of Diabetes and Heart Disease with Machine Learning Classifiers: The Mediation Analysis </title> <p>Purpose: This research employs machine learning and mediation analysis, along with path analysis, to investigate the correlations between factors such as body mass index (BMI) and the occurrence of diabetes and heart disease among the Indian population. The objective is to enhance models that are specifically designed to accommodate lifestyles, genetic differences, and healthcare obstacles. Methods: Our research combines a range of data that includes aspects such as lifestyle, physical health, and mental well-being. We use mediation and path analysis techniques to identify the factors involved in the process, while also utilizing machine learning classifiers to enhance risk assessment. In addition to considering known risks, we also investigate biomarkers. Incorporate time factors through analyses. Results: Mediation and path models analyze that diabetes and heart disease are partially mediated with their coefficients a = 7.85, b = 0.01, and c-c' = 0.10. In the path analysis model, the standardized values of exposure and outcome variables are 4.14 and 6.85, respectively, showing a significant relationship with the mediator and other covariates. In classification, the Random Forest classifier shows 99% accuracy and precession, while the Decision Tree, Extra Tree, K-Nearest, and Adaboost classifiers have an accuracy of 98%, 97%, 96%, and 95%, which shows that the machine learning classifiers are more significant for the study. Conclusion: This study contributes to the development of risk management for diabetes and heart disease in India by utilizing machine learning and mediation analysis. It examines relationships, such as BMI, to provide insights for targeted measures, thereby contributing to global discussions on health.</p> <p>Keywords: Machine learning; mediation model; path modeling; diabetes; BMI</p> <hd id="AN0186774769-2">Introduction</hd> <p>As the global burden of diabetes and heart disease continues to escalate, there is a growing imperative to develop accurate and personalized risk prediction models. A common phenomenon observed in various systems, be they biological, mechanical, or informational, is the presence of a mediator, denoted as m, which plays a crucial role in transmitting the relationship between the variables x and y. Such an example is attached in Figure 1.</p> <p>Graph: Figure 1. An overview of the mediation Model.</p> <p>Figure 1 depicts that the direct influence of X on the outcome is represented by</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi /&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#8242;&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> after the indirect effect has been eliminated (c would be the effect before posting the indirect effect and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi /&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#8242;&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> = the indirect effect). a and b represent the indirect path of the effect of X on the outcome through the mediator (Bauer et al., [<reflink idref="bib4" id="ref1">4</reflink>]). Using the psych library in R software (Gunzler et al., [<reflink idref="bib11" id="ref2">11</reflink>]) to perform the mediation model with their impact (Imai et al., [<reflink idref="bib14" id="ref3">14</reflink>]).</p> <p>To elaborate, we assume that two sets of equations describe the relationship between the variables. First, the mediator model links the exposure to the output of the machine learning model as follows:</p> <p>(<reflink idref="bib1" id="ref4">1</reflink>)</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;a&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/math&gt; </ephtml> </p> <p>Where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the intercept,</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;a&lt;/mi&gt;&lt;/math&gt; </ephtml> is the coefficient describing the exposure relationship, and the error term is</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8764;&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;N&lt;/mi&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> . Second, the outcome model links the exposure and the output of the machine learning model to the outcome as follows:</p> <p>(<reflink idref="bib2" id="ref5">2</reflink>)</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/math&gt; </ephtml> </p> <p>Where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the intercept,</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;/math&gt; </ephtml> is the coefficient representing the mediator-to-output relationship,</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;/math&gt; </ephtml> is the direct effect of the exposure on the outcome, and the error term</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8764;&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;N&lt;/mi&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> . Once the parameters have been estimated, we can express the total effect</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;&amp;#964;&lt;/mi&gt;&lt;/math&gt; </ephtml> as the sum of the direct and indirect effects:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;&amp;#964;&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;ab&lt;/mi&gt;&lt;/math&gt; </ephtml> . This is equivalent to the decomposition obtained in a standard univariate mediation analysis (Baron and Kenny, 1986) (Shrout &amp; Bolger, [<reflink idref="bib29" id="ref6">29</reflink>]), and one can investigate whether a significant mediation effect exists by testing</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;ab&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/math&gt; </ephtml> .</p> <p>We propose jointly fitting all model parameters, including those in the machine learning model, through a single, unified modeling approach. Combining the error terms from the equation. (<reflink idref="bib1" id="ref7">1</reflink>) and (<reflink idref="bib2" id="ref8">2</reflink>), the global loss function contribution over all observations is given by:</p> <p>(<reflink idref="bib3" id="ref9">3</reflink>)</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi mathvariant="script"&gt;L&lt;/mi&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="&amp;#8741;" close="&amp;#8741;"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="&amp;#8741;" close="&amp;#8741;"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;a&lt;/mi&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>The solution to the global loss function corresponds to the maximum likelihood estimate of the three-variable path model under normality assumptions.</p> <p>We propose an iterative algorithm to estimate the model parameters that alternates between fitting the machine learning model</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> and the three-variable mediation model. Let us assume that the parameters</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> ,</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;a&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;/math&gt; </ephtml> , and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;/math&gt; </ephtml> are known. The goal is to find an optimal solution for Equation. Let</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;/math&gt; </ephtml> and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;a&lt;/mi&gt;&lt;/math&gt; </ephtml> . Then, keeping track of only those terms that involve</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> and completing the square, Eq. (<reflink idref="bib3" id="ref10">3</reflink>) becomes:</p> <p>(<reflink idref="bib4" id="ref11">4</reflink>)</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mtable rowspacing="4pt" columnspacing="1em"&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mrow&gt;&lt;mi mathvariant="script"&gt;L&lt;/mi&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;mtd&gt;&lt;mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="&amp;#8741;" close="&amp;#8741;"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="&amp;#8741;" close="&amp;#8741;"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mrow&gt;&lt;mi mathvariant="italic" /&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;mtd&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8733;&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mrow&gt;&lt;mi mathvariant="italic" /&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;mtd&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8733;&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="&amp;#8741;" close="&amp;#8741;"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;/mtable&gt;&lt;/math&gt; </ephtml> </p> <p>In the age of information, data has become a beacon guiding us through the complex terrain of healthcare (Anderson, [<reflink idref="bib2" id="ref12">2</reflink>]). With each heartbeat, every meal, and every genetic code, a narrative unfolds – a story of our well-being etched in bits and bytes (Papacharissi, [<reflink idref="bib22" id="ref13">22</reflink>]). Amidst this deluge of health data, two relentless adversaries persist diabetes and heart disease (Freudenberg, [<reflink idref="bib9" id="ref14">9</reflink>]).</p> <p>Enter the era of machine learning, where algorithms become our trusted allies in this battle for health and longevity (Bone et al., [<reflink idref="bib6" id="ref15">6</reflink>]). Picture this: your medical history, habits, and genes converge into a digital canvas painted by the brush of predictive analytics (Jin, [<reflink idref="bib15" id="ref16">15</reflink>]). It's a canvas that reveals not only your past but also the brushstrokes of your future – a future where the risks of diabetes and heart disease are foreseen and preempted (Agus, [<reflink idref="bib1" id="ref17">1</reflink>]).</p> <p>In this paper, we introduce a machine learning-based method for identifying the mediator with the help of path analysis and structural equation modeling. Our proposed approach links the mediator and the related covariates that can affect diabetes and heart disease. Our proposed machine-learning algorithm tests the standard three-variable mediation model, and the path modeling is suitable for the primary data set used. The mediation model and machine learning algorithm are fitted alternately in our suggested method using an iterated maximization technique. The method thus offers a way to integrate behavioral outcomes, high-dimensional brain measurements, and exposure variables into a single model.</p> <p>Diabetes and heart disease are significant health concerns worldwide, and their prediction and prevention are crucial for effective healthcare management (Ponikowski et al., [<reflink idref="bib23" id="ref18">23</reflink>]). This literature review aims to explore the research on predicting Diabetes and heart disease, focusing on the mediating role of body mass index (BMI) and using machine learning techniques and structural equation modeling. Several studies have investigated predicting Diabetes and heart disease using machine learning algorithms. (Yılıdrım et al., [<reflink idref="bib37" id="ref19">37</reflink>]) utilized random forest, logistic regression, and support vector machine algorithms to predict Diabetes and heart disease. Their study demonstrated the effectiveness of these algorithms in disease prediction using real-time data analytics (Makram et al., [<reflink idref="bib17" id="ref20">17</reflink>]).</p> <p>The relationship between BMI diabetes and heart disease has been extensively studied. Xu et al. ([<reflink idref="bib36" id="ref21">36</reflink>]) conducted a prospective cohort study and found a significant interaction between heart rate and BMI on incident type 2 diabetes mellitus. This suggests that BMI plays a mediating role in the development of diabetes in individuals with different heart rates. The Framingham Heart Study has been instrumental in developing risk scores for various cardiovascular diseases.</p> <p>Schnabel et al. ([<reflink idref="bib27" id="ref22">27</reflink>]) developed a risk score for atrial fibrillation using data from the Framingham Heart Study. Their study included BMI as one of the risk factors for atrial fibrillation, highlighting the importance of BMI in predicting heart disease. In addition to BMI, other covariates have also been considered in predicting Diabetes and heart disease. Metcalf et al. ([<reflink idref="bib19" id="ref23">19</reflink>]) compared the Framingham and the United Kingdom Prospective Diabetes Study risk prediction equations for coronary heart disease in individuals with type 2 diabetes. Their study revealed that the Framingham equation underestimated the risk in this population, emphasizing the need for accurate prediction models tailored to individuals with Diabetes.</p> <p>Machine learning techniques have shown promise in predicting Diabetes and heart disease. Gaber et al. ([<reflink idref="bib10" id="ref24">10</reflink>]) developed an intelligent healthcare system using machine learning algorithms such as support vector machines, linear regression, and random forest classifiers. Their study demonstrated the potential of data-driven approaches for predicting cardiovascular disease and diabetes. Structural equation modeling (SEM) has also been utilized to understand the complex relationships between variables in predicting diabetes and heart disease. Ihnaini et al. ([<reflink idref="bib13" id="ref25">13</reflink>]) proposed an intelligent healthcare recommendation system for multidisciplinary diabetes patients using deep ensemble learning and data fusion based on SEM. Their study highlighted the importance of integrating multiple data sources and utilizing SEM for accurate prediction and recommendation. The literature review highlights the significance of predicting diabetes and heart disease, with a focus on the mediating role of BMI and the use of machine learning techniques and structural equation modeling (Wong et al., [<reflink idref="bib35" id="ref26">35</reflink>]). The studies reviewed demonstrate the effectiveness of machine learning algorithms in disease prediction and emphasize the importance of considering BMI and other covariates in inaccurate risk assessment (Morgenstern et al., [<reflink idref="bib21" id="ref27">21</reflink>]).</p> <hd id="AN0186774769-3">Feature selection and engineering</hd> <p>To maximize prediction accuracy, we conduct feature selection and engineering to identify the most influential risk factors. This step is critical for model generalization and interpretability.</p> <hd id="AN0186774769-4">Machine learning models</hd> <p>We employ a variety of state-of-the-art machine learning classifiers, including but not limited to Extra Tree, Adaboost, Decision Tree, K-Nearest Classifier, and Random Forest classifiers, to build predictive models. These classifiers harness the power of data-driven insights to forecast disease risk.</p> <hd id="AN0186774769-5">Evaluation and validation</hd> <p>Rigorous cross-validation and performance metrics are applied to assess the models' predictive capabilities. To gauge model effectiveness, we examine accuracy, precession, recall, the F-1 score, the area under the receiver operating characteristic curve (AUC-ROC), and k-fold for cross-validation.</p> <hd id="AN0186774769-6">Clinical relevance</hd> <p>Beyond predictive accuracy, we discuss the clinical relevance and practical implications of the model's outputs. How can these predictions inform clinical decision-making? What are the potential benefits for patient care and healthcare resource allocation?</p> <p>By exploring the intricate interplay of factors contributing to the risk of diabetes and heart disease, this research provides a valuable framework for proactive healthcare management. Our findings promise to facilitate early intervention strategies and guide healthcare policies to reduce the burden of these prevalent chronic diseases. Ultimately, the fusion of machine learning and healthcare promises a future where prevention takes precedence and the impact of diabetes and heart disease is mitigated for individuals and societies alike.</p> <hd id="AN0186774769-7">Substantial contribution</hd> <p></p> <ulist> <item> To investigate a novel strategy for revealing hidden patterns in medical data.</item> <p></p> <item> To predict the likelihood of developing heart disease as accurately as possible.</item> <p></p> <item> To most accurately predict the likelihood of developing Diabetes.</item> <p></p> <item> Mediation and path modeling were used to make the results more accurate and relatively precise.</item> </ulist> <hd id="AN0186774769-8">Literature Review</hd> <p>In an era where bytes speak louder than beats and algorithms whisper secrets of our health, the quest to predict, preempt, and conquer the twin titans of Diabetes and heart disease has taken a captivating turn as we embark on our expedition into the realms of Machine Learning (Cohen, [<reflink idref="bib7" id="ref28">7</reflink>]).</p> <hd id="AN0186774769-9">The genesis of predictive analytics</hd> <p>Our research begins with the advent of predictive analytics, where historical health records laid the foundation for what would become a healthcare revolution. The field pioneers harnessed the power of regression models, giving us early glimpses into the relationships between variables and the probabilities of future health outcomes (Wilson et al., [<reflink idref="bib34" id="ref29">34</reflink>]). The quest was on – to not merely treat but to predict, to anticipate the storm before it rained.</p> <hd id="AN0186774769-10">The emergence of machine learning</hd> <p>As technology advanced, so did our arsenal. Enter Machine Learning – where the used classifiers give the best Accuracy regarding the previous models. In previous studies, authors used different machine learning classifiers and comparisons with the secondary data set with reasonable Accuracy (Sun et al., [<reflink idref="bib30" id="ref30">30</reflink>]; Uddin et al., [<reflink idref="bib33" id="ref31">33</reflink>]).</p> <hd id="AN0186774769-11">The data explosion</hd> <p>In the age of wearable tech and electronic health records, data has become the currency of healthcare. The digital footprints of our lives – our heartbeats, footsteps, and dietary choices – all became fodder for predictive algorithms (Sharma, [<reflink idref="bib28" id="ref32">28</reflink>]). Here, the journey to predict diabetes and heart disease has found its true north.</p> <hd id="AN0186774769-12">The predictive powerhouse</hd> <p>Studies abound showcasing the prowess of Machine Learning classifiers in forecasting disease risks. Random Forests teased apart intricate relationships between BMI and diabetes predisposition, while Support Vector Machines revealed the heartbeat of heart disease in electrocardiogram data (Huang, [<reflink idref="bib12" id="ref33">12</reflink>]). As each classifier donned its armor of code, it illuminated the path to early intervention and prevention.</p> <hd id="AN0186774769-13">Personalized medicine beckons</hd> <p>Predictive analytics is not just about prognosis; it's about precision. The literature is rife with tales of personalized medicine where machine learning doesn't just predict disease – it tailors solutions (Price &amp; Nicholson, [<reflink idref="bib25" id="ref34">25</reflink>]). From customized treatment plans to lifestyle interventions, the promise of individualized care is becoming a reality (Johnson et al., [<reflink idref="bib16" id="ref35">16</reflink>]).</p> <p>In the pages that follow, we embark on our odyssey – a mission to predict the risks of diabetes and heart disease. Armed with data, algorithms, and a commitment to precision, we aim to contribute to the ever-evolving saga of predictive healthcare (Mohamed et al., [<reflink idref="bib20" id="ref36">20</reflink>]). In a world where bytes can save lives and classifiers are the heralds of health, our journey is not just academic – it's an exploration of the future of wellness (Topol, [<reflink idref="bib32" id="ref37">32</reflink>]).</p> <hd id="AN0186774769-14">Materials and methods</hd> <p></p> <hd id="AN0186774769-15">Primary data collection</hd> <p>The researchers collected data in Madhya Pradesh, Bhopal, using a Google form distributed through social media, hospitals [67.25%, which is equal to 154 respondents], and meditation centers [32.75%, which is equal to 75 respondents]. Figure 4 displays the 2D distribution of the chosen variables, illustrating their correlation. However, the correlation between blood group vs gender and smoking vs diabetes is weak. The districts within Bhopal were selected where the percentage of women with diabetes in the state was lower than that of men (5.1% vs. 6.7%). However, the prevalence of very high diabetes in people aged 15 to 49 was slightly higher in men (2.9%) than in women (2.1%).</p> <hd id="AN0186774769-16">Sampling method</hd> <p>For the current study, data is collected according to guidelines provided by the National Family Health Survey (NFHS) (Bansode &amp; Prasad, [<reflink idref="bib3" id="ref38">3</reflink>]). This paper survey is based on a convenient sampling design using the 2011 census of India as a sampling framework to select primary units from rural and urban areas. The survey is conducted for the research as per the national family health survey norms. 229 sample size is calculated with a 95% confidence interval, and 6.35% is the margin of error.</p> <hd id="AN0186774769-17">Questionnaire design</hd> <p>The questionnaire is entirely based on the study of diabetic patients. Our target population is people who are suffering from physical and mental health issues. For ease of respondent participation, the questionnaire is designed in bilingual, as shown in Figure 2.</p> <p>Graph: Figure 2. Data collection process.</p> <hd id="AN0186774769-18">Data coding</hd> <p>Data collection has been done by using the 2-point and 5-point Likert scales, and their detailed descriptions are shown in Table 1 and Figure 3, where the 2-point scale is for diabetes and heart disease questions and the 5-point scale is for other covariates.</p> <p>Graph: Figure 3. List of the variables.</p> <p>Graph: Figure 4. 2-D distribution of the variables.</p> <p>Table 1. Detailed description of the data set.</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;S. No&lt;/td&gt;&lt;td&gt;Dataset Categories&lt;/td&gt;&lt;td&gt;Description&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;1&lt;/td&gt;&lt;td&gt;Diabetes with Comorbidity&lt;/td&gt;&lt;td&gt;In the first data set, Diabetes is a dependent variable and the rest of the variables are independent variables as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;Asthma with Comorbidity&lt;/td&gt;&lt;td&gt;In the second data set, Asthma is a dependent variable and other variables are independent variables used as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;Stroke with Comorbidity&lt;/td&gt;&lt;td&gt;In the third data set, Stroke is a dependent variable; other variables are independent variables used as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;Kidney disease with Comorbidity&lt;/td&gt;&lt;td&gt;The fourth data set is a dependent variable; other variables are independent variables used as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;Skin cancer with Comorbidity&lt;/td&gt;&lt;td&gt;In the fifth data set, Skin Cancer is a dependent variable; other variables are independent variables used as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;Heart disease with Comorbidity&lt;/td&gt;&lt;td&gt;In the sixth data set, Heart Disease is a dependent variable; other variables are independent variables used as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;Diabetes, with other variables excluding other diseases&lt;/td&gt;&lt;td&gt;In the seventh data set, Diabetes is a dependent variable and other variables are independent variables used as covariates excluding other diseases.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;Asthma, with other variables, excluding other diseases&lt;/td&gt;&lt;td&gt;In the eighth data set, Asthma is a dependent variable; other variables are independent variables used as covariates.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;Stroke, with other variables&lt;/td&gt;&lt;td&gt;excluding another disease&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;In the ninth data set, Stroke is a dependent variable, and other variables are independent variables used as covariates excluding other diseases.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;Kidney Disease, with other variables excluding another disease&lt;/td&gt;&lt;td&gt;In the tenth data set, Kidney disease is a dependent variable and other variables are independent variables used as covariates excluding other diseases.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;Skin cancer, with other variables excluding other diseases.&lt;/td&gt;&lt;td&gt;In the eleventh data set, Skin Cancer is a dependent variable and other variables are independent variables used as covariates excluding other diseases.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;12&lt;/td&gt;&lt;td&gt;Heart Disable, with other variables excluding other diseases.&lt;/td&gt;&lt;td&gt;In the twelfth data set, Heart Disease is a dependent variable and other variables are independent variables used as covariates, excluding other diseases.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0186774769-19">Methodology</hd> <p>This proposed methodology uses the body mass index to mediate diabetes and heart disease, using the Pearson chi-square correlation to establish the relationship between the variables. Exploratory data analysis is used for feature selection and data representation to present all the variables used in the data set. We conducted experiments using various classification and ensemble algorithms to make predictions about diabetes and heart disease using a test size of 0.25 and a random state of 30. We will provide a concise overview of our research phases in the upcoming sections.</p> <hd id="AN0186774769-20">K-nearest neighbor</hd> <p>KNN, a supervised machine learning technique, is versatile, handling both classification and regression tasks. It's a "lazy" algorithm that assumes similar data points cluster together and groups new data based on similarity. Utilizing a tree-like structure to calculate distances between points, KNN identifies the nearest neighbors (K being a positive integer) within the training dataset to make predictions. Algorithm: The primary data set has been collected from the different hospitals in Bhopal, Madhya Pradesh, India.</p> <p>−choose a test dataset of attributes and rows.</p> <p>−Find the Euclidean distance with the help of the formula-</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;Eculidean&lt;/mi&gt;&lt;mtext /&gt;&lt;mi mathvariant="italic"&gt;Distance&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msqrt&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;l&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;n&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="italic"&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;j&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;l&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;l&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/msqrt&gt;&lt;/math&gt; </ephtml> </p> <p>Then, Decide a random value of</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;K&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . is the no. of nearest neighbors -Then, with the help of these minimum distances and Euclidean distance finds out the nth column of each. -Find out the same output values.</p> <hd id="AN0186774769-21">Adaboost classifier</hd> <p></p> <ulist> <item> Input: a set of training samples with labels</item> </ulist> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mfenced open="{" close=""&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;...&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> a Component Learn algorithm and the number of cycles,</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;T&lt;/mi&gt;&lt;/math&gt; </ephtml> .</p> <p></p> <ulist> <item> Initialize: the weights of training samples:</item> </ulist> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msubsup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mrow&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;N&lt;/mi&gt;&lt;/math&gt; </ephtml> , for all</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;...&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;N&lt;/mi&gt;&lt;/math&gt; </ephtml> .</p> <p></p> <p>• Do for</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;...&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;T&lt;/mi&gt;&lt;/math&gt; </ephtml> </p> <p></p> <ulist> <item> Use the Component Learn algorithm to train a component classifier</item> </ulist> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> on the weighted training samples.</p> <p></p> <ulist> <item> Calculate the training error of</item> </ulist> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8800;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> .</p> <p></p> <ulist> <item> Set weight for the component classifier</item> </ulist> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mo form="prefix"&gt;ln&lt;/mo&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> .</p> <p></p> <ulist> <item> Update the weights of the training samples:</item> </ulist> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo form="prefix"&gt;exp&lt;/mo&gt;&lt;mfenced open="{" close="}"&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> ,</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;...&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;N&lt;/mi&gt;&lt;/math&gt; </ephtml> where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is a normalization constant, and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/math&gt; </ephtml> .</p> <p>4. Output:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;f&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;sign&lt;/mi&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> .</p> <hd id="AN0186774769-22">Extra tree classifier</hd> <p>Step 1: The n subsets of the training sample are extracted using the bootstrap sampling process from the complete training sample set D. The sample size of the subsets is similar to the overall sample set for training D.</p> <p>Step 2: n decision trees are built in line with the n subsets, and results of the n classification are obtained.</p> <p>Step 3: each decision tree casts one voting unit for the most common class, which decides optimal outcomes</p> <hd id="AN0186774769-23">Decision tree classifier</hd> <p>A decision tree serves as a fundamental classification technique within supervised learning. It finds application when dealing with categorical response variables. Structured like a tree, this model delineates the classification procedure according to input features, encompassing a wide range of variable types such as graphs, text, and discrete and continuous data. Information gain</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Info&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mspace width="thinmathspace" /&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mo form="prefix"&gt;log&lt;/mo&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>Where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;P&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the probability that an arbitrary tuple in</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> belongs to class</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;C&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Inf&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;o&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;A&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;V&lt;/mi&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mspace width="1em" /&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;XInfo&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;j&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> </p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Gain&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;A&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Info&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Inf&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;o&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;A&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/math&gt; </ephtml> Gini Index The Gini index favors larger partitions. It uses a squared proportion of classes. Perfectly classified, the Gini Index would be zero. The variable split should have a low Gini index.</p> <hd id="AN0186774769-24">Random forest classifier</hd> <p>Random Forest, a powerful ensemble learning approach, excels in classification and regression tasks, offering superior accuracy compared to many other models. This method effortlessly manages large data sets. Leo Breiman (Popescu, [<reflink idref="bib24" id="ref39">24</reflink>]) pioneered it, making it a widely acclaimed ensemble learning technique. Random forest enhances the performance of decision trees by reducing variance. It achieves this by creating numerous decision trees during training and ultimately delivering the most common class or the mean prediction of individual trees for classification or regression, respectively.</p> <p>Algorithm: The first step is to select the "</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;R&lt;/mi&gt;&lt;/math&gt; </ephtml> " features from the total features "</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;m&lt;/mi&gt;&lt;/math&gt; </ephtml> " where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;R&lt;/mi&gt;&lt;mo&gt;&amp;#8810;&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;M&lt;/mi&gt;&lt;/math&gt; </ephtml> .</p> <p>- Among the "</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;R&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> " features, the node using the best-split point.</p> <p>- Split the node into sub-nodes using the best split.</p> <p>- Repeat the a to c steps until the "l" number of nodes has been reached.</p> <p>- Built forest by repeating steps a to d for "a" number of times to create "</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;n&lt;/mi&gt;&lt;/math&gt; </ephtml> " number of trees.</p> <p>The random forest finds the best split using the Gin-Index Cost Function, which is given by:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Gini&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;k&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> </p> <p>Where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the Each class and proportion of training instances. Random Forest is used here for predicting Diabetes.</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;MSE&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;&amp;#945;&lt;/mi&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>Where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;N&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the no. of the variables used in the data set</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;f&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;i&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the value returned by the model and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;i&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the actual value for data point</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;i&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> .</p> <hd id="AN0186774769-25">K-fold cross validation</hd> <p>We described that a common goal of machine learning is to find an algorithm that produces predictors</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mover&gt;&lt;mi&gt;Y&lt;/mi&gt;&lt;mo stretchy="false"&gt;&amp;#710;&lt;/mo&gt;&lt;/mover&gt;&lt;/math&gt; </ephtml> for an outcome</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;Y&lt;/mi&gt;&lt;/math&gt; </ephtml> that minimizes the MSE:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;MSE&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mfenced open="{" close="}"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mover&gt;&lt;mi&gt;Y&lt;/mi&gt;&lt;mo stretchy="false"&gt;&amp;#710;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;Y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> </p> <p>When all we have at our disposal is one dataset, we can estimate the MSE with the observed MSE like this:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;MSE&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mover&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo stretchy="false"&gt;&amp;#710;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>These two are often referred to as actual and apparent errors, respectively.</p> <p>First, fixing all the algorithm parameters is important before we start the cross-validation procedure. Although we will train the algorithm on the set of training sets, the parameters</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;&amp;#955;&lt;/mi&gt;&lt;/math&gt; </ephtml> will be the same across all training sets. We will use</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo stretchy="false"&gt;&amp;#710;&lt;/mo&gt;&lt;/mover&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;&amp;#955;&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/math&gt; </ephtml> to denote the predictors obtained when we use parameters</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;&amp;#955;&lt;/mi&gt;&lt;/math&gt; </ephtml> . So, if we are going to imitate this definition:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;MSE&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;&amp;#955;&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mfenced open="(" close=")"&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mover&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo stretchy="false"&gt;&amp;#710;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;i&lt;/mi&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;&amp;#955;&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>We want to consider data sets that can be considered independent random samples, and we want to do this several times. With</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;K&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> -fold cross-validation, we do it</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;K&lt;/mi&gt;&lt;/math&gt; </ephtml> times. In the cartoons, we are showing an example that uses</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;K&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/math&gt; </ephtml> . We will eventually end up with</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;K&lt;/mi&gt;&lt;/math&gt; </ephtml> samples, but let's start by describing how to construct the first: we pick</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;M&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;K&lt;/mi&gt;&lt;/math&gt; </ephtml> observations at random (we round if</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;M&lt;/mi&gt;&lt;/math&gt; </ephtml> is not a round number) and think of these as a random sample</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;...&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> , with</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;b&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/math&gt; </ephtml> .</p> <hd id="AN0186774769-26">Mathematical path model</hd> <p>The mathematical representation of a path model involves specifying a set of linear equations that describe the relationships between variables. In a SEM context, the model is often represented in matrix notation.</p> <p>For example, if we have variables</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;X&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Y&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> , and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Z&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> in a simple path model where</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;X&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> influences</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;Y&lt;/mi&gt;&lt;/math&gt; </ephtml> and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;Y&lt;/mi&gt;&lt;/math&gt; </ephtml> influences</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;Z&lt;/mi&gt;&lt;/math&gt; </ephtml> , the equations might look like:</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace"&gt;&lt;mtr&gt;&lt;mtd /&gt;&lt;mtd&gt;&lt;mi mathvariant="italic"&gt;Y&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;X&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;mtr&gt;&lt;mtd /&gt;&lt;mtd&gt;&lt;mi mathvariant="italic"&gt;Z&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi mathvariant="italic"&gt;Y&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;/mtable&gt;&lt;/math&gt; </ephtml> </p> <p>Where: -</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> are path coefficients. -</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#1013;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> are error terms.</p> <hd id="AN0186774769-27">Experimental findings</hd> <p></p> <hd id="AN0186774769-28">Mediation model</hd> <p>The given mediation model shows that the exposure variable X, i.e., diabetes, and the outcome variable Y, i.e., heart disease, are mediated by the mediator variable with values of a = 7.85, b = 0.01, c = 0.91, and</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi /&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#8242;&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0.81&lt;/mn&gt;&lt;/math&gt; </ephtml> . This psych library of R software has been used to test the mediation between diabetes and heart disease.</p> <p>mediate(y = Heart Disease Diabetic + (BMI) – Blood Pressure – Sex – Age + data = bf + n.iter = 500 +)</p> <p>Mediation/Moderation Analysis Call: mediate(y = Heart Disease Diabetic + (BMI) – Blood Pressure – Sex – Age, data = bf, n.iter = 500)</p> <p>The DV (Y) was heart disease. The IV (X) was Intercept* Diabetic*. The mediating variable(s) = BMI*. Variable(s) partially out were blood pressure, sex, and age.</p> <p>Total effect(c) of Intercept* on Heart Disease* = 0.91 S.E. = 0.02 <emph>t</emph> = 46.29 df = 416 with <emph>p</emph> = 3.3e-166</p> <p>Direct effect (</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi /&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#8242;&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> ) of Intercept* on heart disease: removing BMI* = 0.81 S.E. = 0.03 <emph>t</emph> = 30.38 df = 415 with <emph>p</emph> = 1.5e-107</p> <p>Indirect effect (ab) of Intercept* on Heart Disease* through BMI* = 0.11 The mean bootstrapped indirect effect is 0.11, with a standard error of 0.03 Lower CI = 0.05, upper CI = 0.</p> <p>summary(mediation psych)</p> <p>Call: mediate(y = Heart Disease Diabetic + (BMI), data = dat, n.iter = 5000)</p> <p>Direct effect estimates (traditional regression):</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi /&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#8242;&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/math&gt; </ephtml> X + M on Y</p> <p> <emph>R</emph> = 0.92</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> = 0.85; F = 1180.27 on 2 and 418 DF p-value: 1.17e-172</p> <p>Tables 2 and 3 depict that the value of the coefficient of determination is 85% (0.85), which means that 85% of the variability in the exposure variable is accounted for by the independent variable(s) included in the regression model. The remaining 15% of the variability is unaccounted for as a random error. Figure 5 depicts that the value of the direct effect is 0.10 which is statistically significant for the mediation model. Tables 4, 5 , and 6 depict that the indirect effect of BMI on diabetes and heart disease is significant because bootstrapping a lower level confidence limit of 0.16 and the upper limit of 0.17 is positive, and zero does not lie between the two limits. Hence, a significant mediation exists; however, the mediation is partial because the coefficient of the indirect effect is less than the total effect.</p> <p>Graph: Figure 5. Mediation model by using psych library.</p> <p>Table 2. Direct effect estimates (traditional regression)</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi /&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#8242;&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/math&gt; </ephtml> X + M on Y.</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;Heart Disease&lt;/td&gt;&lt;td&gt;se&lt;/td&gt;&lt;td&gt;t&lt;/td&gt;&lt;td&gt;df&lt;/td&gt;&lt;td&gt;Prob&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Intercept&lt;/td&gt;&lt;td&gt;&amp;#8722;0.30&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;5.28&lt;/td&gt;&lt;td&gt;418&lt;/td&gt;&lt;td&gt;2.03e-07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Diabetic&lt;/td&gt;&lt;td&gt;0.80&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;30.52&lt;/td&gt;&lt;td&gt;418&lt;/td&gt;&lt;td&gt;1.97e-108&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;BMI&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;td&gt;6.02&lt;/td&gt;&lt;td&gt;418&lt;/td&gt;&lt;td&gt;3.73e-09&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 3. Total effect estimates (c) (X on Y).</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;Heart Disease&lt;/td&gt;&lt;td&gt;se&lt;/td&gt;&lt;td&gt;t&lt;/td&gt;&lt;td&gt;df&lt;/td&gt;&lt;td&gt;Prob&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Intercept&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;2.36&lt;/td&gt;&lt;td&gt;419&lt;/td&gt;&lt;td&gt;1.86e-02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Diabetic&lt;/td&gt;&lt;td&gt;0.91&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;46.30&lt;/td&gt;&lt;td&gt;419&lt;/td&gt;&lt;td&gt;7.26e-167&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 4. "a" effect estimates (X on M).</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;BMI&lt;/td&gt;&lt;td&gt;se&lt;/td&gt;&lt;td&gt;t&lt;/td&gt;&lt;td&gt;df&lt;/td&gt;&lt;td&gt;Prob&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Intercept&lt;/td&gt;&lt;td&gt;24.02&lt;/td&gt;&lt;td&gt;0.28&lt;/td&gt;&lt;td&gt;85.36&lt;/td&gt;&lt;td&gt;419&lt;/td&gt;&lt;td&gt;4.79e-267&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Diabetic&lt;/td&gt;&lt;td&gt;7.92&lt;/td&gt;&lt;td&gt;0.40&lt;/td&gt;&lt;td&gt;19.74&lt;/td&gt;&lt;td&gt;419&lt;/td&gt;&lt;td&gt;8.28e-62&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 5. "b" effect estimates (M on Y controlling for X).</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;Heart Disease&lt;/td&gt;&lt;td&gt;se&lt;/td&gt;&lt;td&gt;t&lt;/td&gt;&lt;td&gt;df&lt;/td&gt;&lt;td&gt;Prob&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;BMI&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.6&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;418&lt;/td&gt;&lt;td&gt;3.73e-09&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 6. "ab" effect estimates (through all mediators).</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;Heart Disease&lt;/td&gt;&lt;td&gt;boot&lt;/td&gt;&lt;td&gt;sd&lt;/td&gt;&lt;td&gt;lower&lt;/td&gt;&lt;td&gt;upper&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;BMI&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;0.16&lt;/td&gt;&lt;td&gt;0.17&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0186774769-29">Path analysis</hd> <p>A path analysis between diabetes, BMI, and heart disease has been performed with the help of AMOS SPSS. The standardized values are shown in the path model Figure 6, which shows that the correlation between covariates [Gender, Smoking, Alcohol Drinking, Stroke, Asthma, Kidney Disease, Age, Blood Pressure] and the mediator variable BMI has a significant impact on diabetes and heart disease.</p> <p>Graph: Figure 6. Path coefficient by SEM.</p> <hd id="AN0186774769-30">Machine learning classifiers</hd> <p>In this analysis, we have used five machine-learning algorithms for the prediction of diabetes and heart disease. In the given Table 7 and Figure 7, it is depicted that the random forest classifier gives the highest accuracy of 99%, the F1-score is 98%, and the K-fold cross-validation is 98%, which shows the perfect prediction of the mentioned disease. whereas the other 4 classifiers, Extra Tree, Adaboost, Decision Tree, and K Nearest, have less accuracy and precision than the random forest classifier.</p> <p>Graph: Figure 7. Visual representation of the machine learning classifiers.</p> <p>Table 7. Classification report of machine learning algorithms.</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;Classifiers&lt;/td&gt;&lt;td&gt;Accuracy&lt;/td&gt;&lt;td&gt;Precision&lt;/td&gt;&lt;td&gt;F-1 Score&lt;/td&gt;&lt;td&gt;Cross Validation&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Random Forest Classifier&lt;/td&gt;&lt;td&gt;0.99&lt;/td&gt;&lt;td&gt;0.99&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Extra Tree Classifier&lt;/td&gt;&lt;td&gt;0.97&lt;/td&gt;&lt;td&gt;0.96&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.97&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;AdaBoost Classifier&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;0.94&lt;/td&gt;&lt;td&gt;0.97&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Decision Tree Classifier&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;0.99&lt;/td&gt;&lt;td&gt;0.97&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;K-Nearest Classifier&lt;/td&gt;&lt;td&gt;0.96&lt;/td&gt;&lt;td&gt;0.96&lt;/td&gt;&lt;td&gt;0.99&lt;/td&gt;&lt;td&gt;0.96&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>The 3-dimensional histogram shows that the random forest classifier is the most efficient in comparison to the other classifiers. The other classifiers, named Extra Tree Classifier, Adaboost Classifier, Decision Tree Classifier, and K Nearest Classifier, are less efficient for predicting diabetes and heart disease.</p> <hd id="AN0186774769-31">ROC curve</hd> <p>The AUC is a scalar value that represents the overall performance of the model. A perfect model is a random forest classifier. In Figure 8 extra tree and decision tree classifier's AUC value is 96% while Adaboost and KNN classifiers have an AUC value of 90% which is not significant for the study.</p> <p>Graph: Figure 8. ROC Plot.</p> <hd id="AN0186774769-32">Comparison of machine learning classifiers with mediation analysis models</hd> <p>The first mediation model in the study shows that body BMI partially mediates diabetes and heart disease. The second structural equations model depicts that the mediator has a significant impact on the exposure and outcome variables in the presence of the other covariates, which are mentioned in the data coding section. For the classification of the results, we have used some machine learning algorithms in which the random forest classifier has the best accuracy, whereas decision tree classifiers, extra tree classifiers, K-nearest classifiers, and Adaboost classifiers are the ones that give lower accuracy and precision than the random forest classifier. For the validation of the used classifiers, we have used the k-fold cross-validation method. After using cross-validation, we found that the used classifier is giving the ideal accuracy, precision, and f1-scores. The cross-validation method is used to reduce the biasedness of the results. By using various methods, authors found that the mediation model depicts partial mediation, the path model shows the correlation between the exposure and outcome with the other covariates, but the machine learning algorithm depicts that the prediction of diabetes and heart disease is mediated by the BMI and other covariates with highest accuracy 99%, precision 99%, F-1 score 98%, and k fold cross validation value is also 98%. The ROC curve provides a comprehensive view of a model's ability to discriminate between classes across different decision thresholds. In this study, ROC is particularly useful when the class distribution is imbalanced.</p> <hd id="AN0186774769-33">Results and interpretations</hd> <p>Our study employed several machine learning classifiers, including Random Forest, Extra Tree, Adaboost, KNN, and Decision Tree classifiers, to predict the risk of diabetes and heart disease. Table 7 and Figure 7 displays the performance metrics, including accuracy, precision, recall, F1-score, and k-fold cross-validation, for each classifier. The ROC curves and corresponding AUC values are presented in Figure 8, demonstrating the discriminative ability of the models. Feature importance analysis revealed key predictors influencing the risk of diabetes and heart disease. Mediation analysis was conducted to explore potential mediating factors in the relationship between predictors and outcomes. Significant indirect effects were observed [BMI partially mediated diabetes and heart disease]. This suggests that certain factors play a mediating role in the development of diabetes and heart disease. The path analysis model elucidated direct and indirect pathways between predictors, mediators, and outcomes. Results indicate the strength and significance of each path, providing a comprehensive understanding of the complex relationships in the risk prediction model. Our machine learning models demonstrated high performance in predicting the risk of diabetes and heart disease.</p> <hd id="AN0186774769-34">Conclusion</hd> <p>In conclusion, the integration of machine learning classifiers with mediation analysis holds significant promise for advancing our understanding of the complex relationships underlying the risk of diabetes and heart disease. This innovative approach not only facilitates the identification of key mediators but also empowers us to develop more accurate predictive models tailored to the Indian context. As our understanding of the intricate interplay between lifestyle factors, genetics, and health outcomes continues to evolve, the potential for refining these predictive models becomes increasingly apparent (Sagner et al., [<reflink idref="bib26" id="ref40">26</reflink>]). As we navigate the ever-expanding landscape of data science and healthcare, the synergistic integration of machine learning and mediation analysis emerges as a beacon of hope in our quest to preemptively manage and mitigate the burden of diabetes and heart disease (Bledsoe &amp; Scherrer, [<reflink idref="bib5" id="ref41">5</reflink>]). By continually refining our models, embracing diverse datasets, and adapting to the evolving healthcare landscape, we stand poised to usher in a new era of personalized and effective preventive strategies, ultimately contributing to improved public health outcomes in the Indian population and beyond.</p> <hd id="AN0186774769-35">Limitations and future work</hd> <p>India has 77 million diabetics, ranking second globally in the worldwide diabetes epidemic behind China. Out of them, 12.1 million are over 65, and by 2045, that number is expected to rise to 27.5 million (Mehta et al., [<reflink idref="bib18" id="ref42">18</reflink>]). Future research could explore the incorporation of genetic factors into diabetes and heart disease prediction models in the Indian population. Understanding the interplay between genetic predispositions and environmental factors could enhance the accuracy of predictions. Conducting long-term prospective studies in India would provide valuable insights into the temporal relationships between potential mediators, risk factors, and the development of diabetes and heart disease (Franklin et al., [<reflink idref="bib8" id="ref43">8</reflink>]). This longitudinal approach can enhance the precision of predictive models. It is crucial to validate mediation analysis models across diverse populations within India, considering the country's substantial heterogeneity in lifestyle, diet, and socioeconomic status (Tillmann, [<reflink idref="bib31" id="ref44">31</reflink>]). This would ensure the generalizability and robustness of the predictive models. Exploring and incorporating novel biomarkers related to metabolic health, inflammation, and oxidative stress could improve the sensitivity and specificity of mediation analysis models. This might involve integrating data from advanced omics technologies. Leveraging digital health technologies, such as wearable devices and mobile health applications, for real-time data collection could offer a more dynamic and comprehensive approach to monitoring potential mediators and refining prediction models. The limited availability of high-quality, standardized data on lifestyle factors, dietary habits, and health outcomes may pose a challenge. Future research should aim to address data gaps and ensure the reliability of input variables. Mediation analysis relies on the assumption of causal relationships between variables. Establishing causality is challenging in observational studies, and future work should explore experimental designs or advanced statistical methods to strengthen causal inference. The applicability of models developed for the Indian population to other ethnic groups may be limited. Future research should assess the external validity of the models and consider potential adaptations for different populations.</p> <hd id="AN0186774769-36">Acknowledgments</hd> <p>The author thanks all participants who responded to the questionnaire questions.</p> <hd id="AN0186774769-37">Disclosure statement</hd> <p>No potential conflict of interest was reported by the author(s).</p> <hd id="AN0186774769-38">Data availability statement</hd> <p>This study is based on the primary data set.</p> <hd id="AN0186774769-39">Ethics approval and consent to participate</hd> <p>The Ethical approval has been given by the VIT Bhopal University with the Indian Council of Medical Research Bhopal India.</p> <ref id="AN0186774769-40"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref4" type="bt">1</bibl> <bibtext> These authors contributed equally to this work.</bibtext> </blist> </ref> <ref id="AN0186774769-41"> <title> References </title> <blist> <bibtext> Agus, D. B. (2012). The end of illness. Simon and Schuster.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref5" type="bt">2</bibl> <bibtext> Anderson, R. (2015). The collection, linking and use of data in biomedical research and health care: Ethical issues.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref9" type="bt">3</bibl> <bibtext> Bansode, B., &amp; Prasad, J. B. (2022). Burden of comorbidities among diabetic patients in latur, india. Clinical Epidemiology and Global Health, 13, 100957. https://doi.org/10.1016/j.cegh.2021.100957</bibtext> </blist> <blist> <bibl id="bib4" idref="ref1" type="bt">4</bibl> <bibtext> Bauer, D. J., Preacher, K. J., &amp; Gil, K. M. (2006). Conceptualizing and testing random indirect effects and moderated mediation in multilevel models: New procedures and recommendations. Psychological Methods, 11 (2), 142. https://doi.org/10.1037/1082-989X.11.2.142</bibtext> </blist> <blist> <bibl id="bib5" idref="ref41" type="bt">5</bibl> <bibtext> Bledsoe, C. H., &amp; Scherrer, R. F. (2007). The dialectics of disruption: Paradoxes of nature and professionalism in contemporary American childbearing. In Reproductive disruptions: Gender, technology, and biopolitics in the new millennium (Vol. 11, pp. 47). https://org/QP251.R444473</bibtext> </blist> <blist> <bibl id="bib6" idref="ref15" type="bt">6</bibl> <bibtext> Bone, D., Lee, C.-C., Chaspari, T., Gibson, J., &amp; Narayanan, S. (2017). Signal processing and machine learning for mental health research and clinical applications [perspectives]. IEEE Signal Processing Magazine, 34 (5), 196 – 195. https://doi.org/10.1109/MSP.2017.2718581</bibtext> </blist> <blist> <bibl id="bib7" idref="ref28" type="bt">7</bibl> <bibtext> Cohen, E. (2022). On learning to heal: Or, what medicine doesn't know. Duke University Press.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref43" type="bt">8</bibl> <bibtext> Franklin, B. A., Brook, R., &amp; Pope, C. A., III. (2015). Air pollution and cardiovascular disease. Current Problems in Cardiology, 40 (5), 207 – 238. https://doi.org/10.1016/j.cpcardiol.2015.01.003</bibtext> </blist> <blist> <bibl id="bib9" idref="ref14" type="bt">9</bibl> <bibtext> Freudenberg, N. (2014). Lethal but legal: Corporations, consumption, and protecting public health. Oxford University Press.</bibtext> </blist> <blist> <bibtext> Gaber, T., El-Ghamry, A., &amp; Hassanien, A. E. (2022). Injection attack detection using machine learning for smart iot applications. Physical Communication, 52, 101685. https://doi.org/10.1016/j.phycom.2022.101685</bibtext> </blist> <blist> <bibtext> Gunzler, D., Chen, T., Wu, P., &amp; Zhang, H. (2013). Introduction to mediation analysis with structural equation modeling. Shanghai Archives of Psychiatry, 25 (6), 390.</bibtext> </blist> <blist> <bibtext> Huang, Y. (2023). Enhancing general language models for biomedical test retrieval via diversified prior knowledge.</bibtext> </blist> <blist> <bibtext> Ihnaini, B., Khan, M., Khan, T. A., Abbas, S., Daoud, M. S., Ahmad, M., &amp; Khan, M. A. (2021). A smart healthcare recommendation system for multidisciplinary diabetes patients with data fusion based on deep ensemble learning. Computational Intelligence and Neuroscience, 2021, 1 – 11. https://doi.org/10.1155/2021/4243700</bibtext> </blist> <blist> <bibtext> Imai, K., Keele, L., Tingley, D., &amp; Yamamoto, T. (2010). Causal mediation analysis using r. In H. Vinod (Ed.), Advances in social science research using R. Lecture Notes in Statistics (Vol. 196, pp. 129 – 154). Springer. https://doi.org/10.1007/978-1-4419-1764-5_8</bibtext> </blist> <blist> <bibtext> Jin, Y. (2016). Interactive medical record visualization based on symptom location in a 2d human body [ PhD thesis ]. Université d'Ottawa/University of Ottawa.</bibtext> </blist> <blist> <bibtext> Johnson, K. B., Wei, W.-Q., Weeraratne, D., Frisse, M. E., Misulis, K., Rhee, K., Zhao, J., &amp; Snowdon, J. L. (2021). Precision medicine, ai, and the future of personalized health care. Clinical and Translational Science, 14 (1), 86 – 93. https://doi.org/10.1111/cts.12884</bibtext> </blist> <blist> <bibtext> Makram, M., Ali, N., &amp; Mohammed, A. (2022). Machine learning approach for diagnosis of heart diseases. 2022 2nd International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), Cairo, Egypt (pp. 69 – 74). IEEE. https://doi.org/10.1109/MIUCC55081.2022.9781735</bibtext> </blist> <blist> <bibtext> Mehta, S., Kashyap, A., &amp; Das, S. (2009). Diabetes mellitus in India: The modern scourge. Medical Journal Armed Forces India, 65 (1), 50 – 54. https://doi.org/10.1016/S0377-1237(09)80056-7</bibtext> </blist> <blist> <bibtext> Metcalf, P. A., Wells, S., &amp; Jackson, R. T. (2014). Assessing 10-year coronary heart disease risk in people with type 2 diabetes mellitus: Framingham versus United Kingdom prospective diabetes study. Journal of Diabetes Mellitus, 4 (1), 12 – 18. https://doi.org/10.4236/jdm.2014.41003</bibtext> </blist> <blist> <bibtext> Mohamed, S. H. P., Thangam, M. M. N., Das, M. A., &amp; Keshamma, E. (2022). Artificial intelligence in the field of health: A new paradigm aspects: Clinical, ethical and legal. Priya Lokare.</bibtext> </blist> <blist> <bibtext> Morgenstern, J. D., Rosella, L. C., Costa, A. P., Souza, R. J., &amp; Anderson, L. N. (2021). Perspective: Big data and machine learning could help advance nutritional epidemiology. Advances in Nutrition, 12 (3), 621 – 631. https://doi.org/10.1093/advances/nmaa183</bibtext> </blist> <blist> <bibtext> Papacharissi, Z. (2018). A networked self and human augmentics, artificial intelligence, sentience. Routledge.</bibtext> </blist> <blist> <bibtext> Ponikowski, P., Anker, S. D., AlHabib, K. F., Cowie, M. R., Force, T. L., Hu, S., Jaarsma, T., Krum, H., Rastogi, V., Rohde, L. E., Samal, U. C., Shimokawa, H., Budi Siswanto, B., Sliwa, K., &amp; Filippatos, G. (2014). Heart failure: Preventing disease and death worldwide. ESC Heart Failure, 1 (1), 4 – 25. https://doi.org/10.1002/ehf2.12005</bibtext> </blist> <blist> <bibtext> Popescu, B. E. (2004). Ensemble learning for prediction. Stanford University.</bibtext> </blist> <blist> <bibtext> Price, W., &amp; Nicholson, I. (2019). Medical AI and contextual bias. Harv JL &amp; Tech, 33, 65.</bibtext> </blist> <blist> <bibtext> Sagner, M., McNeil, A., Puska, P., Auffray, C., Price, N. D., Hood, L., Lavie, C. J., Han, Z.-G., Chen, Z., Brahmachari, S. K., McEwen, B. S., Soares, M. B., Balling, R., Epel, E., &amp; Arena, R. (2017). The P4 health spectrum – a predictive, preventive, personalized and participatory continuum for promoting Healthspan. Progress in Cardiovascular Diseases, 59 (5), 506 – 521. https://doi.org/10.1016/j.pcad.2016.08.002</bibtext> </blist> <blist> <bibtext> Schnabel, R. B., Sullivan, L. M., Levy, D., Pencina, M. J., Massaro, J. M., D'Agostino, R. B., Newton-Cheh, C., Yamamoto, J. F., Magnani, J. W., Tadros, T. M., Kannel, W. B., Wang, T. J., Ellinor, P. T., Wolf, P. A., Vasan, R. S., &amp; Benjamin, E. J. (2009). Development of a risk score for atrial fibrillation (framingham heart study): A community-based cohort study. Lancet, 373 (9665), 739 – 745. https://doi.org/10.1016/S0140-6736(09)60443-8</bibtext> </blist> <blist> <bibtext> Sharma, M. (2023). Artificial intelligence and iot for smart cities. In Smart urban computing applications (pp. 155 – 189). River Publishers.</bibtext> </blist> <blist> <bibtext> Shrout, P. E., &amp; Bolger, N. (2002). Mediation in experimental and nonexperimental studies: New procedures and recommendations. Psychological Methods, 7 (4), 422. https://doi.org/10.1037/1082-989X.7.4.422</bibtext> </blist> <blist> <bibtext> Sun, L., Tang, L., Shao, G., Qiu, Q., Lan, T., &amp; Shao, J. (2020). A machine learning-based classification system for urban built-up areas using multiple classifiers and data sources. Remote Sensing, 12 (1), 91. https://doi.org/10.3390/rs12010091</bibtext> </blist> <blist> <bibtext> Tillmann, T. (2018). Psychosocial and socioeconomic factors in the development of cardiovascular disease: A study of causality, mediation, international variation, and prediction in predominantly Eastern European Settings [ PhD thesis ]. UCL (University College London).</bibtext> </blist> <blist> <bibtext> Topol, E. (2019). Deep medicine: How artificial intelligence can make healthcare human again. Hachette UK.</bibtext> </blist> <blist> <bibtext> Uddin, S., Khan, A., Hossain, M. E., &amp; Moni, M. A. (2019). Comparing different supervised machine learning algorithms for disease prediction. BMC Medical Informatics &amp; Decision Making, 19 (1), 1 – 16. https://doi.org/10.1186/s12911-019-1004-8</bibtext> </blist> <blist> <bibtext> Wilson, B. S., Tucci, D. L., Moses, D. A., Chang, E. F., Young, N. M., Zeng, F.-G., Lesica, N. A., Bur, A. M., Kavookjian, H., Mussatto, C., Penn, J., Goodwin, S., Kraft, S., Wang, G., Cohen, J. M., Ginsburg, G. S., Dawson, G., &amp; Francis, H. W. (2022). Harnessing the power of artificial intelligence in otolaryngology and the communication sciences. Journal of the Association for Research in Otolaryngology, 23 (3), 319 – 349. https://doi.org/10.1007/s10162-022-00846-2</bibtext> </blist> <blist> <bibtext> Wong, W. E. J., Chan, S. P., Yong, J. K., Tham, Y. Y. S., Lim, J. R. G., Sim, M. A., Soh, C. R., Ti, L. K., &amp; Chew, T. H. S. (2021). Assessment of acute kidney injury risk using a machine-learning guided generalized structural equation model: A cohort study. BMC Nephrology, 22 (1), 1 – 8. https://doi.org/10.1186/s12882-021-02238-9</bibtext> </blist> <blist> <bibtext> Xu, P. H., Hui, C. K., Lui, M. M., Lam, D. C., Fong, D. Y., &amp; Ip, M. S. (2019). Incident type 2 diabetes in osa and effect of cpap treatment: A retrospective clinic cohort study. Chest, 156 (4), 743 – 753. https://doi.org/10.1016/j.chest.2019.04.130</bibtext> </blist> <blist> <bibtext> Yılıdrım, E., Çalhan, A., &amp; Cicioğlu, M. (2022). Performance analysis of disease diagnostic system using iomt and real-time data analytics. Concurrency &amp; Computation: Practice &amp; Experience, 34 (13), 6916. https://doi.org/10.1002/cpe.6916</bibtext> </blist> </ref> <aug> <p>By Ajay Verma and Manisha Jain</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib11" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib14" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib29" firstref="ref6"></nolink> <nolink nlid="nl4" bibid="bib22" firstref="ref13"></nolink> <nolink nlid="nl5" bibid="bib15" firstref="ref16"></nolink> <nolink nlid="nl6" bibid="bib23" firstref="ref18"></nolink> <nolink nlid="nl7" bibid="bib37" firstref="ref19"></nolink> <nolink nlid="nl8" bibid="bib17" firstref="ref20"></nolink> <nolink nlid="nl9" bibid="bib36" firstref="ref21"></nolink> <nolink nlid="nl10" bibid="bib27" firstref="ref22"></nolink> <nolink nlid="nl11" bibid="bib19" firstref="ref23"></nolink> <nolink nlid="nl12" bibid="bib10" firstref="ref24"></nolink> <nolink nlid="nl13" bibid="bib13" firstref="ref25"></nolink> <nolink nlid="nl14" bibid="bib35" firstref="ref26"></nolink> <nolink nlid="nl15" bibid="bib21" firstref="ref27"></nolink> <nolink nlid="nl16" bibid="bib34" firstref="ref29"></nolink> <nolink nlid="nl17" bibid="bib30" firstref="ref30"></nolink> <nolink nlid="nl18" bibid="bib33" firstref="ref31"></nolink> <nolink nlid="nl19" bibid="bib28" firstref="ref32"></nolink> <nolink nlid="nl20" bibid="bib12" firstref="ref33"></nolink> <nolink nlid="nl21" bibid="bib25" firstref="ref34"></nolink> <nolink nlid="nl22" bibid="bib16" firstref="ref35"></nolink> <nolink nlid="nl23" bibid="bib20" firstref="ref36"></nolink> <nolink nlid="nl24" bibid="bib32" firstref="ref37"></nolink> <nolink nlid="nl25" bibid="bib24" firstref="ref39"></nolink> <nolink nlid="nl26" bibid="bib26" firstref="ref40"></nolink> <nolink nlid="nl27" bibid="bib18" firstref="ref42"></nolink> <nolink nlid="nl28" bibid="bib31" firstref="ref44"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1477877 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Predicting the Risk of Diabetes and Heart Disease with Machine Learning Classifiers: The Mediation Analysis – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Ajay+Verma%22">Ajay Verma</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0994-4812">0000-0002-0994-4812</externalLink>)<br /><searchLink fieldCode="AR" term="%22Manisha+Jain%22">Manisha Jain</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3961-3827">0000-0003-3961-3827</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Measurement%3A+Interdisciplinary+Research+and+Perspectives%22"><i>Measurement: Interdisciplinary Research and Perspectives</i></searchLink>. 2025 23(3):310-327. – Name: Avail Label: Availability Group: Avail Data: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 18 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Diabetes%22">Diabetes</searchLink><br /><searchLink fieldCode="DE" term="%22Heart+Disorders%22">Heart Disorders</searchLink><br /><searchLink fieldCode="DE" term="%22Risk%22">Risk</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Path+Analysis%22">Path Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22India%22">India</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1080/15366367.2024.2347811 – Name: ISSN Label: ISSN Group: ISSN Data: 1536-6367<br />1536-6359 – Name: Abstract Label: Abstract Group: Ab Data: Purpose: This research employs machine learning and mediation analysis, along with path analysis, to investigate the correlations between factors such as body mass index (BMI) and the occurrence of diabetes and heart disease among the Indian population. The objective is to enhance models that are specifically designed to accommodate lifestyles, genetic differences, and healthcare obstacles. Methods: Our research combines a range of data that includes aspects such as lifestyle, physical health, and mental well-being. We use mediation and path analysis techniques to identify the factors involved in the process, while also utilizing machine learning classifiers to enhance risk assessment. In addition to considering known risks, we also investigate biomarkers. Incorporate time factors through analyses. Results: Mediation and path models analyze that diabetes and heart disease are partially mediated with their coefficients a = 7.85, b = 0.01, and c-c' = 0.10. In the path analysis model, the standardized values of exposure and outcome variables are 4.14 and 6.85, respectively, showing a significant relationship with the mediator and other covariates. In classification, the Random Forest classifier shows 99% accuracy and precession, while the Decision Tree, Extra Tree, K-Nearest, and Adaboost classifiers have an accuracy of 98%, 97%, 96%, and 95%, which shows that the machine learning classifiers are more significant for the study. Conclusion: This study contributes to the development of risk management for diabetes and heart disease in India by utilizing machine learning and mediation analysis. It examines relationships, such as BMI, to provide insights for targeted measures, thereby contributing to global discussions on health. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1477877 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1477877 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/15366367.2024.2347811 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 18 StartPage: 310 Subjects: – SubjectFull: Diabetes Type: general – SubjectFull: Heart Disorders Type: general – SubjectFull: Risk Type: general – SubjectFull: Prediction Type: general – SubjectFull: Classification Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Path Analysis Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: India Type: general Titles: – TitleFull: Predicting the Risk of Diabetes and Heart Disease with Machine Learning Classifiers: The Mediation Analysis Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Ajay Verma – PersonEntity: Name: NameFull: Manisha Jain IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1536-6367 – Type: issn-electronic Value: 1536-6359 Numbering: – Type: volume Value: 23 – Type: issue Value: 3 Titles: – TitleFull: Measurement: Interdisciplinary Research and Perspectives Type: main |
| ResultId | 1 |