Poster Presentation Clinical Oncology Society of Australia Annual Scientific Meeting 2026

Do machine learning models outperform logistic regression for risk stratification in resource-constrained settings? insights from a real-world cohort with implications for oncology care (145897)

Sachpreet Singh Singh 1 , Talvir Sidhu 1 , Malkiat Singh 1 , Anubha Garg 1
  1. Government Medical College and Rajindra Hospital, Patiala, Punjab, India, Gurdaspur, PUNJAB, India

Aims: Early risk stratification is central to optimising outcomes, resource allocation, and follow-up in oncology, particularly in resource-limited settings. While machine learning (ML) models are increasingly proposed, their real-world advantage over conventional approaches remains uncertain. We evaluated whether ML improves prediction of unfavourable outcomes compared with logistic regression (LR) using routinely collected programmatic data from a high-burden chronic disease cohort.

Methods: We conducted a retrospective cohort study (2020–2025) within a large public-sector program. Unfavourable outcomes included death, loss to follow-up, treatment failure, and regimen modification. Predictors comprised baseline demographic, clinical, microbiological, nutritional, treatment, and temporal variables. Among 384 patients, 174 (45.3%) had unfavourable outcomes. LR, random forest, and XGBoost models were evaluated using repeated stratified five-fold cross-validation. Performance was assessed using discrimination (AUC), calibration, and decision-curve analysis. Model interpretability was examined using SHAP values.

Results: LR achieved the highest discrimination (AUC 0.642, 95% CI 0.617–0.666), outperforming random forest (0.622) and XGBoost (0.611). Calibration was modest across all models. Decision-curve analysis demonstrated limited but consistent clinical utility, with no incremental benefit of ML over LR across clinically relevant thresholds. Key predictors included age, disease classification, nutritional status, treatment regimen, and geographic factors, with consistent importance across models.

Conclusion: In this real-world cohort, ML models did not outperform conventional LR for predicting adverse outcomes. Given its comparable performance, greater interpretability, and ease of implementation, LR remains a pragmatic choice for risk stratification in routine clinical settings. These findings have direct implications for oncology, where adoption of complex ML models should be balanced against real-world utility. Improving prediction performance will likely require integration of longitudinal and treatment-response data, particularly in resource-constrained health systems.

  1. Huang RJ, Kwon NS, Tomizawa Y, Choi AY, Hernandez-Boussard T, Hwang JH. A comparison of logistic regression against machine learning algorithms for gastric cancer risk prediction within real-world clinical data streams. JCO Clin Cancer Inform. 2022;6:e2200039. doi:10.1200/CCI.22.00039.
  2. Leonard G, South C, Balentine C, Porembka M, Mansour J, Wang S, et al. Machine learning improves prediction over logistic regression in resected colon cancer patients. J Surg Res. 2022;275:181–193. doi:10.1016/j.jss.2022.01.012.