tetano
Editor, Senior Moderator
JAMIA Open
. 2026 Jun 23;9(3)
oag079.
doi: 10.1093/jamiaopen/ooag079. eCollection 2026 Jun.
Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods: analysis from 2 international randomized controlled trials
Christian Møller Jensen[SUP] 1 [/SUP], Ramtin Zargari Marandi[SUP] 1 [/SUP], Kasper Sommerlund Moestrup[SUP] 1 [/SUP], Ahmad Mourad[SUP] 2 3 [/SUP], Alfredo J Mena Lora[SUP] 4 [/SUP], Brad T Sherman[SUP] 5 [/SUP], David M Vock[SUP] 6 [/SUP], Jacqueline A Nordwall[SUP] 6 [/SUP], Joanne M Carson[SUP] 7 [/SUP], Nathan Peiffer-Smadja[SUP] 8 9 [/SUP], Neil R Aggarwal[SUP] 10 [/SUP], Nicole E Naiman[SUP] 11 [/SUP], Samuel M Brown[SUP] 12 [/SUP], Thomas W Barrett[SUP] 13 14 [/SUP], Timothy Hatlen[SUP] 15 [/SUP], Victoria S Kjærgaard[SUP] 1 [/SUP], Weizhong Chang[SUP] 5 [/SUP], Matthew R Sydes[SUP] 16 17 [/SUP], Jens Lundgren[SUP] 1 18 19 [/SUP], Tomas O Jensen[SUP] 1 [/SUP]; STRIVE Network and ITAC and TICO Study Groups
Collaborators, Affiliations
Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort.
Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC.
Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (-15.8%), but performance remained above chance level.
Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors.
Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.
Keywords: SARS-CoV-2; biostatistics; machine learning; medical informatics; randomized controlled trial.
. 2026 Jun 23;9(3)
doi: 10.1093/jamiaopen/ooag079. eCollection 2026 Jun.
Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods: analysis from 2 international randomized controlled trials
Christian Møller Jensen[SUP] 1 [/SUP], Ramtin Zargari Marandi[SUP] 1 [/SUP], Kasper Sommerlund Moestrup[SUP] 1 [/SUP], Ahmad Mourad[SUP] 2 3 [/SUP], Alfredo J Mena Lora[SUP] 4 [/SUP], Brad T Sherman[SUP] 5 [/SUP], David M Vock[SUP] 6 [/SUP], Jacqueline A Nordwall[SUP] 6 [/SUP], Joanne M Carson[SUP] 7 [/SUP], Nathan Peiffer-Smadja[SUP] 8 9 [/SUP], Neil R Aggarwal[SUP] 10 [/SUP], Nicole E Naiman[SUP] 11 [/SUP], Samuel M Brown[SUP] 12 [/SUP], Thomas W Barrett[SUP] 13 14 [/SUP], Timothy Hatlen[SUP] 15 [/SUP], Victoria S Kjærgaard[SUP] 1 [/SUP], Weizhong Chang[SUP] 5 [/SUP], Matthew R Sydes[SUP] 16 17 [/SUP], Jens Lundgren[SUP] 1 18 19 [/SUP], Tomas O Jensen[SUP] 1 [/SUP]; STRIVE Network and ITAC and TICO Study Groups
Collaborators, Affiliations
- PMID: 42344108
- PMCID: PMC13289610
- DOI: 10.1093/jamiaopen/ooag079
Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort.
Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC.
Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (-15.8%), but performance remained above chance level.
Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors.
Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.
Keywords: SARS-CoV-2; biostatistics; machine learning; medical informatics; randomized controlled trial.