Conferences

A record of academic conferences where my research has been accepted or presented, as well as those I have attended, in reverse-chronological order, alongside my peer-review assignments.

10th International Conference on Computational Social Science, Philadelphia, USA (IC2S2-2024)

Understanding the divergence between digital health footprints and official prevalence: A multi-behaviour study across US states

Accepted as a poster presentation (Link) (Poster)

Presented by: Andrés Gvirtz

Accepted: 18 Apr 2024 · Presented: 19 Jul 2024

Authors: Andrés Gvirtz (King’s College London, UK), Sanja Šćepanović (Nokia Bell Labs, UK), Anubhab Das (Textify AI, India), Daniele Quercia (Nokia Bell Labs, UK)

Abstract: 4 in 5 US adults seek healthcare advice online, making social media and online searches common tools to study health conditions’ prevalence, outbreaks, and evolution. Two main reasons to use online data rather than, e.g., official CDC statistics to study regional distributions of health outcomes, are that online data are cheaper to collect and have a higher temporal resolution. The underlying assumption is that online data and official statistics assess the same concept and hence should converge. Despite the numerous successes of utilizing online data, there have also been notable failures such as Google Flu Trends. Approaches are variable, sometimes focusing on active engagement, e.g., posting, other times on passive engagement, e.g. searching. We ask, is a particular type of online data more suitable for public health research than others, and is this dependent on the medical condition? We further ask, what might explain the differences between different online data, and between different online data and official data?

We propose a uniformed framework that allows us to 1) compare active and passive engagement across conditions, revealing when and why one outperforms the other, and 2) explain the biases inherent in both types of online health data.

Combing the largest online derived medical taxonomy from Reddit and Google Trend Searches for a set of 17 equivalent conditions in the two data sources, we show that there is implicit knowledge to be gained from looking at cases where official, active, and passive derived online scores diverge (i.e., higher or lower online engagement in health discussions about conditions compared to their official prevalence). Specifically, we find that: 1) online searches (active engagement) more closely resemble true prevalences than social media discussions (passive engagement) for most but not all conditions, 2) the divergence between official scores and those derived from active engagement on social media discussions is explained by state socio-demographic, personality & cultural indicators to a much higher degree than the divergence between official scores and those derived from passive engagement, i.e. online searches.

Drawing from our results, certain guidelines emerge for interpreting online health data. Firstly, while search trends seem to normally outperform discussion based scores, caution is required, and comparative approaches recommended. Divergences can illuminate potential issues, including the stigma associated with conditions, the severity of conditions, and other variables impacting how individuals actively discuss their medical conditions. This study illuminates the complexities and nuances of online health behaviors, urging a more discerning approach to interpreting such data.


9th International Conference on Computational Social Science, Copenhagen, Denmark (IC2S2-2023)

Topics of Our Dreams

Accepted as an oral presentation in parallel tracks (Link) (Tweet)

Presented by: Sanja Šćepanović

Accepted: 14 Apr 2023 · Presented: 19 Jul 2023

Authors: Anubhab Das (Heritage Institute of Technology, India), Sanja Šćepanović (Nokia Bell Labs, UK), Luca Maria Aiello (IT University of Copenhagen, Denmark)

Abstract: Dreaming is a fundamental human experience, and one that is yet not fully understood. One popular method in study of dreams is content analysis. There are over 130 scales and rating systems for dream content analysis been published. One of the most well-known among the content-analysis scales is the one by Hall and Van de Castle, in which the key elements (i.e., characters, interactions, and emotions) in dreams are mapped. Another important approach to content analysis is to look at dream topics (e.g., life events, supernatural entities, or work). While previous work studied individual, or a small number of pre-selected topics, and did so using small-scale datasets that are often non-representative, there is no research uncovering a comprehensive set of topics relevant across dreams. To bridge this research gap, we relied on two recent advancements: i) the availability of large crowd-sourced datasets of dream self-reports (i.e., r/Dreams subreddit on Reddit), and ii) the improved AI methods for natural language processing (NLP). We collected over 44K dream reports published by over 34K Reddit users during 5 years, and applied BERTopic content analysis method to automatically discover the topics in each of the dreams. This resulted in 222 topics. By then applying large language model embeddings, to measure semantic similarity among topics, we grouped them into 22 higher-level categories, creating the most comprehensive taxonomy of dream topics to date. We also used complex analysis methods to understand the co-occurrence in dreams between different topics (e.g., finding that People & Relationships and Life Events topics co-occur frequently, while Space and Indoor Locations appear rarely together). Our methodology enabled us also to characterise different types of dreams in terms of topics that are peculiar to them; and to analyse temporal trends in collective dream experiences, which we found were changed by external events, such as the COVID-19 pandemic.


Peer-Reviewing

  1. Served as a peer-reviewer for the 12th International Conference on Computational Social Science, Burlington (IC2S2-2026), reviewing four extended abstracts. (Conference Website)
  2. Served as a peer-reviewer for the 11th International Conference on Computational Social Science, Sweden (IC2S2-2025), reviewing five extended abstracts. (Conference Website)