An LLM and Retrieval Pipeline for Reproductive-Health Misinformation and Stigma

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

Social media is a major source of health information, but reproductive-health topics such as polycystic ovary syndrome (PCOS), endometriosis, fertility, and menstruation attract both unsupported claims and stigmatizing framing. We present an evidence-grounded pipeline that combines large language models (LLMs) with biomedical retrieval to assess such claims on Bluesky, and we extend it with a novel layer that scores the stigma framing of each claim independently of its truth. From 991 English-language posts we extracted 937 distinct health claims, grounded each in evidence retrieved from PubMed, openFDA, and authoritative clinical guidelines (ACOG, NICE, NHS, WHO), and classified veracity into five categories using only the retrieved evidence. A parallel module scored six stigma dimensions (shame, secrecy, disgust, fear, blame, empowerment). Most claims (75.2%) were not supported by biomedical evidence, yet the dominant emotional frame across the corpus was empowerment rather than overt shame. Critically, stigma increased monotonically as claims departed from the evidence, from supported to dangerous claims, a graded relationship confirmed by a Kruskal–Wallis test (p <10−7) and reflected in effect sizes that grew steadily across the veracity scale. Stigma also varied sharply by condition, being highest for infertility and endometriosis and lowest for PCOS. These findings show that misinformation and stigma co-occur in reproductive-health discourse and demonstrate that evidence-grounded LLMs can jointly surface both. A single-coder human validation served as a pilot diagnostic check: stigma scores were directionally aligned with human judgment, while five-way veracity agreement was limited, indicating that categorical veracity labels require calibration against a larger multi-coder reference set. As an exploratory single-platform study, results should be interpreted within the keyword-sampled corpus rather than as platform-wide prevalence estimates.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-23T09:30:01.253652+00:00
unpaywall
last seen: 2026-08-22T06:23:50.750314+00:00
License: CC-BY-4.0