JAMIA Open
. 2026 Jun 23;9(3):ooag079.
doi: 10.1093/jamiaopen/ooag079. eCollection 2026 Jun.
Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods: analysis from 2 international randomized controlled trials
Christian Møller Jensen 1 , Ramtin Zargari Marandi 1 , Kasper Sommerlund Moestrup 1 , Ahmad Mourad 2 3 , Alfredo J Mena Lora 4 , Brad T Sherman 5 , David M Vock 6 , Jacqueline A Nordwall 6 , Joanne M Carson 7 , Nathan Peiffer-Smadja 8 9 , Neil R Aggarwal 10 , Nicole E Naiman 11 , Samuel M Brown 12 , Thomas W Barrett 13 14 , Timothy Hatlen 15 , Victoria S Kjærgaard 1 , Weizhong Chang 5 , Matthew R Sydes 16 17 , Jens Lundgren 1 18 19 , Tomas O Jensen 1 ; STRIVE Network and ITAC and TICO Study Groups
Collaborators, Affiliations
Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort.
Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC.
Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (-15.8%), but performance remained above chance level.
Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors.
Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.
Keywords: SARS-CoV-2; biostatistics; machine learning; medical informatics; randomized controlled trial.
. 2026 Jun 23;9(3):ooag079.
doi: 10.1093/jamiaopen/ooag079. eCollection 2026 Jun.
Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods: analysis from 2 international randomized controlled trials
Christian Møller Jensen 1 , Ramtin Zargari Marandi 1 , Kasper Sommerlund Moestrup 1 , Ahmad Mourad 2 3 , Alfredo J Mena Lora 4 , Brad T Sherman 5 , David M Vock 6 , Jacqueline A Nordwall 6 , Joanne M Carson 7 , Nathan Peiffer-Smadja 8 9 , Neil R Aggarwal 10 , Nicole E Naiman 11 , Samuel M Brown 12 , Thomas W Barrett 13 14 , Timothy Hatlen 15 , Victoria S Kjærgaard 1 , Weizhong Chang 5 , Matthew R Sydes 16 17 , Jens Lundgren 1 18 19 , Tomas O Jensen 1 ; STRIVE Network and ITAC and TICO Study Groups
Collaborators, Affiliations
- PMID: 42344108
- PMCID: PMC13289610
- DOI: 10.1093/jamiaopen/ooag079
Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort.
Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC.
Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (-15.8%), but performance remained above chance level.
Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors.
Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.
Keywords: SARS-CoV-2; biostatistics; machine learning; medical informatics; randomized controlled trial.