From words to wellbeing: Mining Dutch social media language to quantify wellbeing
preprint
OA: closed
CC-BY-4.0
Abstract
Wellbeing measures, traditionally based on self-reports, could potentially be enhanced with Social Media Text Mining (SMTM), providing dynamic, real-time insights into wellbeing. This study explores the application of SMTM to assess wellbeing using Dutch social media data. Utilizing a dataset from the Netherlands Twin Register (N = 411), we analysed social media posts using both closed and open vocabulary methods. Words with negative valence and those related to negative emotions were most strongly related to wellbeing. Next, we applied machine learning models to predict survey-based wellbeing scores. In these models, modest convergence between SMTM predictions and traditional self-report wellbeing measures (r = 0.12) was found. Incorporating Facebook likes and temporal information (e.g., monthly changes in word use) did not improve prediction model performances, even though they yielded insightful associations with survey-based wellbeing measures. Overall, this study highlights the potential of integrating social media data for research and its applicability to non-English contexts.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0