• FluTrackers.com Inc. does not provide medical advice. Information on this web site is collected from various internet resources, and the FluTrackers board of directors makes no warranty to the safety, efficacy, correctness or completeness of the information posted on this site by any author or poster. The information collated here is for instructional and/or discussion purposes only and is NOT intended to diagnose or treat any disease, illness, or other medical condition. Every individual reader or poster should seek advice from their personal physician/healthcare practitioner before considering or using any interventions that are discussed on this website. By continuing to access this website you agree to consult your personal physican before using any interventions posted on this website, and you agree to hold harmless FluTrackers.com Inc., the board of directors, the members, and all authors and posters for any effects from use of any medication, supplement, vitamin or other substance, device, intervention, etc. mentioned in posts on this website, or other internet venues referenced in posts on this website.
  • We are not asking for any donations. Do not donate to any entity who says they are raising funds for us.

Genbank Metadata

gsgs

Registered User
I recently discovered that for many sequences they have additional
data at genbank. (link)
http://www.ncbi.nlm.nih.gov/Traces/...ENTER_NAME='TIGR'+AND+TRACE_TYPE_CODE='RT-PCR'

http://www.ncbi.nlm.nih.gov/Traces/trace.cgi?cmd=stat&f=xml_list_species&m=obtain&s=species&letter=I

When they sequence a flu-genome, they get ~200 partial subsequences
from "random" positions in the genome of length ~500-800 nucleotides.
From these subsequences the genome is assembled, they usually overlap
and every position is covered multiple times.
And for ~9000 flu-genomes (7000 with ftp) these sets of subsequences
are also available.
I took the ~600 sets from wild birds only, filtered those that for some reason
didn't work with my programs and 435 were remaining for which I calculated
the alignments, the average number of subsequences that cover a position
(green in the pic) and the probability that the finally assigned nucleotide
(I don't know how they assign that value) differs from the average
(black in the pic)

I noticed that the region in H2 was more often covered than the others,
maybe subsequences break and start preferrably in that region

======================================
 

Attachments

  • covav1.gif
    covav1.gif
    9.5 KB · Views: 0
Back
Top