tetano
Editor, Senior Moderator
Nucleic Acids Res
. 2025 Feb 8;53(4):gkaf077.
doi: 10.1093/nar/gkaf077. Using minor variant genomes and machine learning to study the genome biology of SARS-CoV-2 over time
Xiaofeng Dong[SUP] 1 [/SUP], David A Matthews[SUP] 2 [/SUP], Giulia Gallo[SUP] 3 [/SUP], Alistair Darby[SUP] 1 [/SUP], I'ah Donovan-Banfield[SUP] 1 4 [/SUP], Hannah Goldswain[SUP] 1 [/SUP], Tracy MacGill[SUP] 5 [/SUP], Todd Myers[SUP] 5 [/SUP], Robert Orr[SUP] 5 [/SUP], Dalan Bailey[SUP] 3 [/SUP], Miles W Carroll[SUP] 4 6 7 [/SUP], Julian A Hiscox[SUP] 1 4 8 [/SUP]
Affiliations
In infected individuals, viruses are present as a population consisting of dominant and minor variant genomes. Most databases contain information on the dominant genome sequence. Since the emergence of SARS-CoV-2 in late 2019, variants have been selected that are more transmissible and capable of partial immune escape. Currently, models for projecting the evolution of SARS-CoV-2 are based on using dominant genome sequences to forecast whether a known mutation will be prevalent in the future. However, novel variants of SARS-CoV-2 (and other viruses) are driven by evolutionary pressure acting on minor variant genomes, which then become dominant and form a potential next wave of infection. In this study, sequencing data from 96 209 patients, sampled over a 3-year period, were used to analyse patterns of minor variant genomes. These data were used to develop unsupervised machine learning clusters to identify amino acids that had a greater potential for mutation than others in the Spike protein. Being able to identify amino acids that may be present in future variants would better inform the design of longer-lived medical countermeasures and allow a risk-based evaluation of viral properties, including assessment of transmissibility and immune escape, thus providing candidates with early warning signals for when a new variant of SARS-CoV-2 emerges.
. 2025 Feb 8;53(4):gkaf077.
doi: 10.1093/nar/gkaf077. Using minor variant genomes and machine learning to study the genome biology of SARS-CoV-2 over time
Xiaofeng Dong[SUP] 1 [/SUP], David A Matthews[SUP] 2 [/SUP], Giulia Gallo[SUP] 3 [/SUP], Alistair Darby[SUP] 1 [/SUP], I'ah Donovan-Banfield[SUP] 1 4 [/SUP], Hannah Goldswain[SUP] 1 [/SUP], Tracy MacGill[SUP] 5 [/SUP], Todd Myers[SUP] 5 [/SUP], Robert Orr[SUP] 5 [/SUP], Dalan Bailey[SUP] 3 [/SUP], Miles W Carroll[SUP] 4 6 7 [/SUP], Julian A Hiscox[SUP] 1 4 8 [/SUP]
Affiliations
- PMID: 39970290
- PMCID: PMC11838042
- DOI: 10.1093/nar/gkaf077
In infected individuals, viruses are present as a population consisting of dominant and minor variant genomes. Most databases contain information on the dominant genome sequence. Since the emergence of SARS-CoV-2 in late 2019, variants have been selected that are more transmissible and capable of partial immune escape. Currently, models for projecting the evolution of SARS-CoV-2 are based on using dominant genome sequences to forecast whether a known mutation will be prevalent in the future. However, novel variants of SARS-CoV-2 (and other viruses) are driven by evolutionary pressure acting on minor variant genomes, which then become dominant and form a potential next wave of infection. In this study, sequencing data from 96 209 patients, sampled over a 3-year period, were used to analyse patterns of minor variant genomes. These data were used to develop unsupervised machine learning clusters to identify amino acids that had a greater potential for mutation than others in the Spike protein. Being able to identify amino acids that may be present in future variants would better inform the design of longer-lived medical countermeasures and allow a risk-based evaluation of viral properties, including assessment of transmissibility and immune escape, thus providing candidates with early warning signals for when a new variant of SARS-CoV-2 emerges.