Using natural language processing to explore Dry January posts on Twitter: A longitudinal infodemiology study (Preprint)

abstract

BACKGROUND
Dry January, a temporary alcohol abstinence campaign, encourages individuals to reflect on their relationship with alcohol by temporarily abstaining from consumption during the month of January. Though Dry January has become a global phenomenon, there has been limited investigation into Dry January participants experiences. One means through which to gain insights into individuals Dry January-related experiences is by leveraging large-scale social media data (e.g., Twitter chatter) to explore and characterize public discourse concerning Dry January.
OBJECTIVE
We sought to answer the following questions: (1) What themes are present within a corpus of tweets about Dry January, and is there consistency in the language used to discuss Dry January across multiple years of tweets (2020-2022)? (2) Do unique themes or patterns emerge in Dry January 2021 tweets after the onset of the COVID-19 pandemic? and (3) What is the association with tweet composition (i.e., sentiment and human-authored vs. bot-authored) and engagement with Dry January tweets?
METHODS
We applied natural language processing techniques to a large sample of tweets (N = 222,917) containing the term dry january or dryjanuary posted from December 15 to February 15 across three separate years of participation (2020-2022). Term frequency inverse document frequency, k-means clustering, and principal component analysis were used for data visualization to identify the optimal number of clusters per year. Once data were visualized, we ran interpretation models to afford within-year (or within-cluster) comparisons. Latent Dirichlet Allocation topic modeling was used to examine content within each cluster per given year. Valence Aware Dictionary and sEntiment [sic] Reasoner sentiment analysis was used to examine affect per cluster per year. Botometer automated account check was employed to determine average bot score per cluster per year. Lastly, to assess user engagement with Dry January content, we took the average number of likes and retweets per cluster and ran correlations with other outcome variables of interest.
RESULTS
We observed several similar topics per year (e.g., Dry January resources, Dry January health benefits, updates related to Dry January progress), suggesting relative consistency in Dry January content over time. While there was overlap in themes across multiple years of tweets, unique themes related to individuals experiences with alcohol during the midst of the COVID-19 global pandemic were detected in the corpus of tweets from 2021. Also, tweet composition was associated with engagement, including number of likes, retweets, and quote-tweets per post. Bot-dominant clusters had fewer likes, retweets, or quote tweets compared with human-authored clusters.
CONCLUSIONS
Findings underscore the utility for using large-scale social media, such as discussions on Twitter, to study drinking reduction attempts and to monitor the ongoing dynamic needs of persons contemplating, preparing for, or actively pursuing attempts to quit or cut down on their drinking.

authors

author list (cited authors)

Russell, A. M., Valdez, D., Chiang, S., Montemayor, B. N., Barry, A. E., Lin, H., & Massey, P. M.

citation count

0

complete list of authors

Russell, Alex M||Valdez, Danny||Chiang, Shawn||Montemayor, Ben N||Barry, Adam E||Lin, Hsien-Chang||Massey, Philip M

Book Title

JMIR Preprints

publication date

June 2022

Using natural language processing to explore Dry January posts on Twitter: A longitudinal infodemiology study (Preprint) Institutional Repository Document

Overview

abstract

authors

author list (cited authors)

citation count

complete list of authors

Book Title

publication date

Research

keywords

Identity

Digital Object Identifier (DOI)

Other

URL