Detecting disinformation with measures of discourse derailment using LLMs

preprint OA: closed
View at publisher

Abstract

Social media presents both opportunities and risks to society. One risk includes the potential for the spread of disinformation, but methods for automatic detecting possible disinformation still mostly rely on low-level lexical cues. We present a new approach to the automatic detection of disinformation based on measuring discourse ‘derailment’ – messages that try to force the topic of discourse away from one topic and onto another. While this may include both malicious and benign derailment, it could serve as an important signal for early-warning systems. In this study, we implement a system for automatic detection of discourse derailment which uses a Large Language Model to generate expected replies and compares them to the real replies. We test this system on a set of human annotated data to show that the system outperforms various baselines and approaches the agreement between human annotators. This suggests that Large Language Models (LLMs) can be sensitive to discourse-level information. However, we also identify evidence of several limitations, including that the automatic system relies on different cues compared to human annotators, which leads to some amount of bias. Nevertheless, our project represents a considerable step towards understanding how to use LLMs to analyse discourse and a new angle on tackling disinformation.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00