A comparative study of ordinal logistic regression and machine learning models for predicting women’s malnutrition in bangladesh: evidence from BDHS 2022. Kulsum, U., Haque, A., Barai, P., & Hossain, M., M. Journal of Health, Population and Nutrition 2026, BioMed Central, 2026.
A comparative study of ordinal logistic regression and machine learning models for predicting women’s malnutrition in bangladesh: evidence from BDHS 2022 [link]Website  doi  abstract   bibtex   
Malnutrition, including both undernutrition and overnutrition, remains a major public health concern in Bangladesh, particularly among women of reproductive age. This study aims to identify key determinants of women’s malnutrition in Bangladesh and compare the predictive performance of ordinal logistic regression and machine learning methods for predicting women’s malnutrition using data from the 2022 Bangladesh Demographic and Health Survey. This study utilized data from 8,728 ever-married women aged 15–49 years extracted from the BDHS 2022. Six ML algorithms, including Random Forest, Extreme Gradient Boosting (XGBoost), Support Vector Machine, Naïve Bayes, AdaBoost, and Multilayer Perceptron (MLP), were compared with ordinal logistic regression by evaluating their performances using accuracy, precision, recall, $$\:F_1$$ score, Cohen’s kappa, and area under the curve (AUC). Data preprocessing included SMOTE to address class imbalance, and models were assessed using stratified k-fold cross-validation. Findings of Ordinal Logistic Regression (OLR) suggest that age, division, residence, wealth index, current breastfeeding status, husband’s education, currently working, and age at first marriage are the significant predictors of women’s malnutrition. However, its predictive performance was modest, with an accuracy of 49% and macro-averaged $$\:F_1$$ score was 0.47. In contrast, ML models outperformed OLR across all evaluation metrics. Random Forest and XGBoost achieved the highest test accuracy (64%), with Random Forest attaining a macro-averaged $$\:F_1$$ score of 0.64 and achieved 66.2% accuracy (10-fold CV). Traditional models, such as OLR, are more explainable, but machine learning models demonstrate higher accuracy in classifying malnutrition. The findings can help policymakers and health professionals prioritize resources and plan targeted nutrition programs, considering the risk factors identified in this study, to lessen the burden of both undernutrition and overnutrition among women in Bangladesh.
@article{
 title = {A comparative study of ordinal logistic regression and machine learning models for predicting women’s malnutrition in bangladesh: evidence from BDHS 2022},
 type = {article},
 year = {2026},
 keywords = {Clinical Nutrition,Epidemiology,Health Promotion and Disease Prevention,Infectious Diseases,Maternal and Child Health,Public Health},
 websites = {https://link.springer.com/article/10.1186/s41043-025-01236-z},
 publisher = {BioMed Central},
 id = {6e5dbae4-1bb0-35e9-a1f8-017df4adc13f},
 created = {2026-01-31T10:01:38.228Z},
 file_attached = {false},
 profile_id = {3d6b17c2-7de8-3a82-bf4f-ddb0e3081e5f},
 last_modified = {2026-01-31T10:01:38.228Z},
 read = {false},
 starred = {false},
 authored = {true},
 confirmed = {true},
 hidden = {false},
 source_type = {JOUR},
 private_publication = {false},
 abstract = {Malnutrition, including both undernutrition and overnutrition, remains a major public health concern in Bangladesh, particularly among women of reproductive age. This study aims to identify key determinants of women’s malnutrition in Bangladesh and compare the predictive performance of ordinal logistic regression and machine learning methods for predicting women’s malnutrition using data from the 2022 Bangladesh Demographic and Health Survey. This study utilized data from 8,728 ever-married women aged 15–49 years extracted from the BDHS 2022. Six ML algorithms, including Random Forest, Extreme Gradient Boosting (XGBoost), Support Vector Machine, Naïve Bayes, AdaBoost, and Multilayer Perceptron (MLP), were compared with ordinal logistic regression by evaluating their performances using accuracy, precision, recall, $$\:F_1$$ score, Cohen’s kappa, and area under the curve (AUC). Data preprocessing included SMOTE to address class imbalance, and models were assessed using stratified k-fold cross-validation. Findings of Ordinal Logistic Regression (OLR) suggest that age, division, residence, wealth index, current breastfeeding status, husband’s education, currently working, and age at first marriage are the significant predictors of women’s malnutrition. However, its predictive performance was modest, with an accuracy of 49% and macro-averaged $$\:F_1$$ score was 0.47. In contrast, ML models outperformed OLR across all evaluation metrics. Random Forest and XGBoost achieved the highest test accuracy (64%), with Random Forest attaining a macro-averaged $$\:F_1$$ score of 0.64 and achieved 66.2% accuracy (10-fold CV). Traditional models, such as OLR, are more explainable, but machine learning models demonstrate higher accuracy in classifying malnutrition. The findings can help policymakers and health professionals prioritize resources and plan targeted nutrition programs, considering the risk factors identified in this study, to lessen the burden of both undernutrition and overnutrition among women in Bangladesh.},
 bibtype = {article},
 author = {Kulsum, Umme and Haque, Ahsanul and Barai, Pallab and Hossain, Md. Moyazzem},
 doi = {10.1186/S41043-025-01236-Z},
 journal = {Journal of Health, Population and Nutrition 2026}
}

Downloads: 0