• FluTrackers.com Inc. does not provide medical advice. Information on this web site is collected from various internet resources, and the FluTrackers board of directors makes no warranty to the safety, efficacy, correctness or completeness of the information posted on this site by any author or poster. The information collated here is for instructional and/or discussion purposes only and is NOT intended to diagnose or treat any disease, illness, or other medical condition. Every individual reader or poster should seek advice from their personal physician/healthcare practitioner before considering or using any interventions that are discussed on this website. By continuing to access this website you agree to consult your personal physican before using any interventions posted on this website, and you agree to hold harmless FluTrackers.com Inc., the board of directors, the members, and all authors and posters for any effects from use of any medication, supplement, vitamin or other substance, device, intervention, etc. mentioned in posts on this website, or other internet venues referenced in posts on this website.
  • We are not asking for any donations. Do not donate to any entity who says they are raising funds for us.

MIR Public Health Surveill Machine Learning to Detect Self-Reporting of Symptoms, Testing Access and Recovery Associated With COVID-19 on Twitter: A

tetano

Editor, Senior Moderator
MIR Public Health Surveill


. 2020 Jun 3.
doi: 10.2196/19509. Online ahead of print.
Machine Learning to Detect Self-Reporting of Symptoms, Testing Access and Recovery Associated With COVID-19 on Twitter: A Retrospective Big-Data Infoveillance Study


Tim Mackey[SUP] 1 2 3 4 [/SUP], Vidya Purushothaman[SUP] 5 2 [/SUP], Jiawei Li[SUP] 1 2 3 4 [/SUP], Neal Shah[SUP] 1 4 [/SUP], Matthew Nali[SUP] 1 3 [/SUP], Cortni Bardier[SUP] 6 [/SUP], Bryan Liang[SUP] 2 3 [/SUP], Mingxiang Cai[SUP] 7 3 2 [/SUP], Raphael Cuomo[SUP] 1 2 [/SUP]



Affiliations

Abstract

Background: The coronavirus (COVID-19) pandemic is a global health emergency with over 6 million cases worldwide as of the beginning of June 2020. Importantly, the pandemic is historic in scope and precedent given its emergence in an increasing digital era. Importantly, there have been concerns about the accuracy of COVID-19 case counts due to issues such as lack of access to testing and difficulty in measuring recoveries.
Objective: The aims of this study were to detect and characterize user-generated conversations that could be associated with COVID-19-related symptoms, experiences with access to testing, and mentions of disease recovery using an unsupervised machine learning approach.
Methods: Tweets were collected from the Twitter public streaming API from March 3-20, 2020, filtered for general COVID-19-related keywords and then further filtered for terms that could be related to COVID-19 symptoms as self-reported by users. Tweets were analyzed using an unsupervised machine learning approach called the biterm topic model (BTM), where groups of tweets containing the same word-related themes were separated into topic clusters that included conversations about symptoms, testing, and recovery. Tweets in these clusters were then extracted and manually annotated for content analysis and also assessed for their statistical and geographic characteristics.
Results: A total of 4,492,954 tweets were collected that contained terms that could be related to COVID-19 symptoms. After using BTM to identify relevant topic clusters and removing duplicate tweets, we identified a total of 3,465 (<1%) tweets that included user generated conversations about experiences that users associated with possible COVID-19 symptoms and other disease experiences. These tweets were grouped into five main categories including first and second-hand reports of symptoms, symptom reporting concurrent with lack of testing, discussion of recovery, confirmation of negative COVID-19 diagnosis after receiving testing, and users recalling symptoms and questioning whether they might have been previously infected with COVID-19. Co-occurrence of tweets for these themes were statistically significant for users reporting symptoms with lack of testing and with discussion of recovery. Sixty-three percent (n=1112) of tweets were located in the United States.
Conclusions: This study used unsupervised machine learning for the purposes of characterizing self-reporting of symptoms, experiences with testing, and mentions of recovery related to COVID-19. Many users reported symptoms they thought were related to COVID-19, but also were not able to get tested to confirm their concerns. In the absence of testing availability and confirmation, accurate case estimations for this period of the outbreak may never be known. Future studies should continue to explore the utility of infoveillance approaches to estimate COVID-19 disease severity.
 
Back
Top Bottom