tetano
Editor, Senior Moderator
Sci Rep
. 2025 Nov 28;15(1):42712.
doi: 10.1038/s41598-025-26705-7. Large language models versus classical machine learning performance in COVID-19 mortality prediction using high-dimensional tabular data
Mohammadreza Ghaffarzadeh-Esfahani[SUP] 1 2 [/SUP], Mahdi Ghaffarzadeh-Esfahani[SUP] 2 [/SUP], Aryan Salahi-Niri[SUP] 1 [/SUP], Hossein Toreyhi[SUP] 1 [/SUP], Zahra Atf[SUP] 3 [/SUP], Amirali Mohsenzadeh-Kermani[SUP] 2 [/SUP], Mahshad Sarikhani[SUP] 4 [/SUP], Zohreh Tajabadi[SUP] 5 [/SUP], Fatemeh Shojaeian[SUP] 6 [/SUP], Mohammad Hassan Bagheri[SUP] 2 [/SUP], Aydin Feyzi[SUP] 7 [/SUP], Mohamadamin Tarighat-Payma[SUP] 4 [/SUP], Narges Gazmeh[SUP] 7 [/SUP], Fateme Heydari[SUP] 4 [/SUP], Hossein Afshar[SUP] 7 [/SUP], Amirreza Allahgholipour[SUP] 7 [/SUP], Farid Alimardani[SUP] 7 [/SUP], Ameneh Salehi[SUP] 4 [/SUP], Naghmeh Asadimanesh[SUP] 4 [/SUP], Mohammad Amin Khalafi[SUP] 4 [/SUP], Hadis Shabanipour[SUP] 7 [/SUP], Ali Moradi[SUP] 7 [/SUP], Sajjad Hossein Zadeh[SUP] 7 [/SUP], Omid Yazdani[SUP] 4 [/SUP], Romina Esbati[SUP] 4 [/SUP], Moozhan Maleki[SUP] 7 [/SUP], Danial Samiei Nasr[SUP] 4 [/SUP], Amirali Soheili[SUP] 4 [/SUP], Hossein Majlesi[SUP] 4 [/SUP], Saba Shahsavan[SUP] 4 [/SUP], Alireza Soheilipour[SUP] 4 [/SUP], Nooshin Goudarzi[SUP] 1 [/SUP], Erfan Taherifard[SUP] 8 [/SUP], Hamidreza Hatamabadi[SUP] 9 [/SUP], Jamil S Samaan[SUP] 10 [/SUP], Thomas Savage[SUP] 11 [/SUP], Ankit Sakhuja[SUP] 12 [/SUP], Ali Soroush[SUP] 12 [/SUP], Girish Nadkarni[SUP] 12 [/SUP], Ilad Alavi Darazam[SUP] 13 14 [/SUP], Mohamad Amin Pourhoseingholi[SUP] 15 16 [/SUP], Seyed Amir Ahmad Safavi-Naini[SUP] 17 18 [/SUP]
Affiliations
This study compared the performance of classical feature-based machine learning models (CMLs) and large language models (LLMs) in predicting COVID-19 mortality using high-dimensional tabular data from 9,134 patients across four hospitals. Seven CML models, including XGBoost and random forest (RF), were evaluated alongside eight LLMs, such as GPT-4 and Mistral-7b, which performed zero-shot classification on text-converted structured data. Additionally, Mistral-7b was fine-tuned using the QLoRA approach. XGBoost and RF demonstrated superior performance among CMLs, achieving F1 scores of 0.87 and 0.83 for internal and external validation, respectively. GPT-4 led the LLM category with an F1 score of 0.43, while fine-tuning Mistral-7b significantly improved its recall from 1% to 79%, yielding a stable F1 score of 0.74 during external validation. Although LLMs showed moderate performance in zero-shot classification, fine-tuning substantially enhanced their effectiveness, potentially bridging the gap with CML models. However, CMLs still outperformed LLMs in handling high-dimensional tabular data tasks. This study highlights the potential of both CMLs and fine-tuned LLMs in medical predictive modeling, while emphasizing the current superiority of CMLs for structured data analysis.
Keywords: COVID-19 mortality; Fine-tuning; Large language models; Machine learning; Structured data; Zero-shot classification.
. 2025 Nov 28;15(1):42712.
doi: 10.1038/s41598-025-26705-7. Large language models versus classical machine learning performance in COVID-19 mortality prediction using high-dimensional tabular data
Mohammadreza Ghaffarzadeh-Esfahani[SUP] 1 2 [/SUP], Mahdi Ghaffarzadeh-Esfahani[SUP] 2 [/SUP], Aryan Salahi-Niri[SUP] 1 [/SUP], Hossein Toreyhi[SUP] 1 [/SUP], Zahra Atf[SUP] 3 [/SUP], Amirali Mohsenzadeh-Kermani[SUP] 2 [/SUP], Mahshad Sarikhani[SUP] 4 [/SUP], Zohreh Tajabadi[SUP] 5 [/SUP], Fatemeh Shojaeian[SUP] 6 [/SUP], Mohammad Hassan Bagheri[SUP] 2 [/SUP], Aydin Feyzi[SUP] 7 [/SUP], Mohamadamin Tarighat-Payma[SUP] 4 [/SUP], Narges Gazmeh[SUP] 7 [/SUP], Fateme Heydari[SUP] 4 [/SUP], Hossein Afshar[SUP] 7 [/SUP], Amirreza Allahgholipour[SUP] 7 [/SUP], Farid Alimardani[SUP] 7 [/SUP], Ameneh Salehi[SUP] 4 [/SUP], Naghmeh Asadimanesh[SUP] 4 [/SUP], Mohammad Amin Khalafi[SUP] 4 [/SUP], Hadis Shabanipour[SUP] 7 [/SUP], Ali Moradi[SUP] 7 [/SUP], Sajjad Hossein Zadeh[SUP] 7 [/SUP], Omid Yazdani[SUP] 4 [/SUP], Romina Esbati[SUP] 4 [/SUP], Moozhan Maleki[SUP] 7 [/SUP], Danial Samiei Nasr[SUP] 4 [/SUP], Amirali Soheili[SUP] 4 [/SUP], Hossein Majlesi[SUP] 4 [/SUP], Saba Shahsavan[SUP] 4 [/SUP], Alireza Soheilipour[SUP] 4 [/SUP], Nooshin Goudarzi[SUP] 1 [/SUP], Erfan Taherifard[SUP] 8 [/SUP], Hamidreza Hatamabadi[SUP] 9 [/SUP], Jamil S Samaan[SUP] 10 [/SUP], Thomas Savage[SUP] 11 [/SUP], Ankit Sakhuja[SUP] 12 [/SUP], Ali Soroush[SUP] 12 [/SUP], Girish Nadkarni[SUP] 12 [/SUP], Ilad Alavi Darazam[SUP] 13 14 [/SUP], Mohamad Amin Pourhoseingholi[SUP] 15 16 [/SUP], Seyed Amir Ahmad Safavi-Naini[SUP] 17 18 [/SUP]
Affiliations
- PMID: 41315569
- PMCID: PMC12663554
- DOI: 10.1038/s41598-025-26705-7
This study compared the performance of classical feature-based machine learning models (CMLs) and large language models (LLMs) in predicting COVID-19 mortality using high-dimensional tabular data from 9,134 patients across four hospitals. Seven CML models, including XGBoost and random forest (RF), were evaluated alongside eight LLMs, such as GPT-4 and Mistral-7b, which performed zero-shot classification on text-converted structured data. Additionally, Mistral-7b was fine-tuned using the QLoRA approach. XGBoost and RF demonstrated superior performance among CMLs, achieving F1 scores of 0.87 and 0.83 for internal and external validation, respectively. GPT-4 led the LLM category with an F1 score of 0.43, while fine-tuning Mistral-7b significantly improved its recall from 1% to 79%, yielding a stable F1 score of 0.74 during external validation. Although LLMs showed moderate performance in zero-shot classification, fine-tuning substantially enhanced their effectiveness, potentially bridging the gap with CML models. However, CMLs still outperformed LLMs in handling high-dimensional tabular data tasks. This study highlights the potential of both CMLs and fine-tuned LLMs in medical predictive modeling, while emphasizing the current superiority of CMLs for structured data analysis.
Keywords: COVID-19 mortality; Fine-tuning; Large language models; Machine learning; Structured data; Zero-shot classification.