• FluTrackers.com Inc. does not provide medical advice. Information on this web site is collected from various internet resources, and the FluTrackers board of directors makes no warranty to the safety, efficacy, correctness or completeness of the information posted on this site by any author or poster. The information collated here is for instructional and/or discussion purposes only and is NOT intended to diagnose or treat any disease, illness, or other medical condition. Every individual reader or poster should seek advice from their personal physician/healthcare practitioner before considering or using any interventions that are discussed on this website. By continuing to access this website you agree to consult your personal physican before using any interventions posted on this website, and you agree to hold harmless FluTrackers.com Inc., the board of directors, the members, and all authors and posters for any effects from use of any medication, supplement, vitamin or other substance, device, intervention, etc. mentioned in posts on this website, or other internet venues referenced in posts on this website.
  • We are not asking for any donations. Do not donate to any entity who says they are raising funds for us.

Recovery of deleted deep sequencing data sheds more light on the early Wuhan SARS-CoV-2 epidemic -preprint

Mary Wilson

Well-known member
bioRxiv preprint
doi: https://doi.org/10.1101/2021.06.18.449051; this version posted June 22, 2021.

Jesse D. Bloom
Fred Hutchinson Cancer Research Center.
Howard Hughes Medical Institute
Seattle, WA, USA


ABSTRACT

The origin and early spread of SARS-CoV-2 remains shrouded in mystery. Here I identify a data set containing SARS-CoV-2 sequences from early in the Wuhan epidemic that has been deleted from the NIH’s Sequence Read Archive. Irecover the deleted files from the Google Cloud, and reconstruct partial sequences of 13 early epidemic viruses. Phylogeneticanalysis of these sequences in the context of carefully annotated existing data suggests that the Huanan Seafood Market sequences that are the focus of the joint WHO-China report are not fully representative of the viruses in Wuhan early in the epidemic. Instead, the progenitor of known SARS-CoV-2 sequences likely contained three mutations relative to the market viruses that made it more similar to SARS-CoV-2’s bat coronavirus relatives.

https://www.biorxiv.org/content/10.1101/2021.06.18.449051v1.full.pdf
 
China confirms unauthorised labs were told to destroy early coronavirus samples

Published: 8:04pm, 15 May, 2020
Zhuang Pinghui in Beijing
  • National health authority says this was done for biosafety reasons and to ‘prevent secondary disasters caused by unidentified pathogens’
  • US Secretary of State Mike Pompeo has said Beijing declined to provide samples and destroyed them at the start of the outbreak
https://www.scmp.com/news/china/soc...rms-unauthorised-labs-were-told-destroy-early
 
I would recommend the comment to this article by ACC (whoever that is) which gives some important background. I do not thing there is anything significant about these sequences.
 
Did a Chinese team ‘obscure’ early coronavirus sequences?


n a world starved for data to clarify the origin of COVID-19, a study claiming to have unearthed early sequences of SARS-CoV-2 that were deliberately hidden was bound to ignite a sizzling debate. Last week a preprint by evolutionary biologist Jesse Bloom of the Fred Hutchinson Cancer Research Center asserted that Chinese researchers sampled viruses from some of the first COVID-19 patients in Wuhan, China, posted the viral sequences to a National Institutes of Health (NIH) database, and then later had the genetic information removed to “obscure their existence.”

The uproar over the preprint led Senator Josh Hawley (R–MO) to demand answers from NIH on why the agency agreed to “purge” the data and to call for an investigation into the matter. Even for some scientists, Bloom's work reinforced suspicions that the Chinese government has tried to hide how the pandemic started. “This is a creative and rigorous approach to investigating the provenance of SARS-CoV-2,” says Ian Lipkin, a microbiologist at Columbia University's Mailman School of Public Health. “There may have been active suppression of epidemiological and sequence data needed to track its origin.”



But critics of Bloom's bioRxiv preprint call his detective work much ado about nothing, because the Chinese scientists later published the viral information in a different form, and the recovered sequences may add little to the origin hunt. “The idea that the group was trying to hide something is farcical,” says evolutionary biologist Andrew Rambaut at the University of Edinburgh.

Bloom says he has no bias toward a particular origin hypothesis for SARS-CoV-2, and he agrees that the 13 partial sequences he recovered don't resolve whether the virus originally jumped to humans from an unknown animal or somehow leaked from a Wuhan virology center. “I don't think this bolsters either the lab origin or zoonosis hypothesis.”

Chinese health officials on 31 December 2019 tied the Huanan Seafood Market in Wuhan to an outbreak of an “unexplained pneumonia,” but within a month it had become clear that many of the earliest COVID-19 cases had no link to the market. Bloom, who studies viral evolution, set out to investigate early cases after a controversial report on the pandemic's origin issued in March by a joint commission of Chinese and foreign researchers overseen by the World Health Organization. The report deemed it “extremely unlikely” that SARS-CoV-2 escaped from a lab; Bloom helped organize a much-discussed letter, cosigned by 17 other scientists, criticizing that conclusion and calling for further investigation.

Bloom wanted to do his own analyses of the viruses detected in the earliest cases, which led him to a list of all SARS-CoV-2 sequences submitted before 31 March 2020 to the Sequence Read Archive (SRA), an NIH database. But when he checked the SRA for one of the listed projects, he couldn't find its sequences. Googling some of the project's information, he found a study, led by Ming Wang from Wuhan University's Renmin Hospital, that had been posted as a preprint on 6 March 2020 on medRxiv and published in June of that year in Small, a journal little known to virologists. That paper lists some of the earliest COVID-19 patients in Wuhan and the specific mutations in their viruses, but doesn't give the full sequence data. Further internet sleuthing led Bloom to discover that the SRA backs up its information in Google's Cloud platform, and a search there turned up files containing some of the Wang team's earlier data submissions.

The Small paper mentions no corrections to the viral sequences that might explain why they were removed from the SRA, which led Bloom to conclude that “the trusting structures of science have been abused to obscure sequences relevant to the early spread of SARS-CoV-2 in Wuhan.” (In the wake of criticisms of the initial preprint, Bloom toned down this sentence and other accusatory language.) Bloom acknowledges SARS-CoV-2 sequences can be derived from the data in the Small paper, but he says most virologists expect to be able to download whole viral genomes from a database like the SRA.

Several authors of the Small paper did not reply to requests for comment, but NIH last week noted in a statement that it removed the SRA sequences at the request of the submitting investigator, who the agency says holds the rights to the data. Bloom added NIH's email exchange to his revised preprint. “I have submitted an updated version of this SRA data to another website,” reads a 15 June email sent to NIH from a Wuhan University researcher whose name was redacted. Yet Bloom says he cannot find the sequences in any other virology database.

Bloom asserts that because the deleted sequences lack three mutations seen in COVID-19 cases linked to the seafood market, the patient viruses Wang's team analyzed more likely represent a progenitor SARS-CoV-2. But Rambaut says the differences Bloom highlighted are too few to distinguish the “roots” of the SARS-CoV-2 family tree.

Leaving aside the meaning of the sequences Bloom found, the demonstration that researchers can potentially find “new” SARS-CoV-2 data in the cloud is an exciting advance and may prompt similar sleuthing, says genomicist Sudhir Kumar of Temple University, who has also analyzed early SARS-CoV-2 sequences. “Many people feel that there is a lot more Chinese data out there, and they don't have access to it.”

https://science.sciencemag.org/content/373/6550/15
 
Tetano - and anyone else following this story the two podcasts I linked to above end up spending the better part of two hours covering this paper, which is a shame really as it is of little interest apart from the first version had a rather unscientific and accusatory tone.
The key points are the SRA is not like GISAID as it is a dump for raw data, not full gene sequences, and often includes misreads and contaminants but is useful for bigdata analysis. In this case the dump only covered spike and some of ORF10.
It is not uncommon for data to be withdraw usually if cleaned up versions are created to avoid duplication.
Wang et al's paper was primarily not about the data but about the Nanopore technology used to generate it. Across the 10% of the genome looked at they only found 3 SNPs which varied from the consensus strain all of which are in the published paper in Small. Small may not be a well know name in the virological field but is a high impact journal if you are publishing about technologies like Nanopore sequencing it has a higher impact rating (11.5) than PNAS. None of the sequence are particularly early there were plenty of full sequences being produced by this time including those with the SNPs in question.
It is a shame to lose data from the publicly accessible databases but this is a database used by small number of big data analytics analysts of genomes for all species and this particular upload did not shed any real light on anything we did not already know.
 
Back
Top Bottom