Diffusion of protest behaviour: Analysing the July 2021 civil unrest in South Africa through sentiment analysis and topic modelling

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract The jailing of former South African President Jacob Zuma in July 2021 ignited protests in KwaZulu-Natal and quickly spread to Johannesburg and Pretoria. Initially peaceful, these demonstrations escalated into widespread looting and violence, culminating in the deaths of 354 individuals. This study employs a mixed-methods approach by integrating machine learning and qualitative analysis to examine the dynamics of the unrest using Twitter data from multiple South African provinces. Tweets were manually annotated for sentiment (positive, negative, neutral), and inter-annotator agreement was measured using Fleiss' Kappa, yielding a score of 0.27, indicative of fair consensus. The most accurate sentiment classification model labelled the remaining dataset, enabling the temporal tracking of sentiment and protest diffusion. Findings underscore the role of economic inequality, political instability, and racial tensions in fuelling the unrest. Also, the lifecycle of the protests progresses from mobilisation to looting, violence, and change, which was confirmed through social media discourse. The study highlights the utility of social media analytics in complementing investigative journalism and informing state responses to urban crises. It contributes to the growing literature on collective violence and the diffusion of protest behaviour in digitally networked societies.
Full text 215,778 characters · extracted from preprint-html · click to expand
Diffusion of protest behaviour: Analysing the July 2021 civil unrest in South Africa through sentiment analysis and topic modelling | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Diffusion of protest behaviour: Analysing the July 2021 civil unrest in South Africa through sentiment analysis and topic modelling Temitope Kekere, Vukosi Marivate, Marié Hattingh This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7760005/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The jailing of former South African President Jacob Zuma in July 2021 ignited protests in KwaZulu-Natal and quickly spread to Johannesburg and Pretoria. Initially peaceful, these demonstrations escalated into widespread looting and violence, culminating in the deaths of 354 individuals. This study employs a mixed-methods approach by integrating machine learning and qualitative analysis to examine the dynamics of the unrest using Twitter data from multiple South African provinces. Tweets were manually annotated for sentiment (positive, negative, neutral), and inter-annotator agreement was measured using Fleiss' Kappa, yielding a score of 0.27, indicative of fair consensus. The most accurate sentiment classification model labelled the remaining dataset, enabling the temporal tracking of sentiment and protest diffusion. Findings underscore the role of economic inequality, political instability, and racial tensions in fuelling the unrest. Also, the lifecycle of the protests progresses from mobilisation to looting, violence, and change, which was confirmed through social media discourse. The study highlights the utility of social media analytics in complementing investigative journalism and informing state responses to urban crises. It contributes to the growing literature on collective violence and the diffusion of protest behaviour in digitally networked societies. civil unrest diffusion of protest behaviour sentiment analysis topic modelling Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Introduction The protest movement that erupted in South Africa in July 2021, following the incarceration of former President Jacob Zuma for contempt of court, represents the latest chapter in the country’s long history of civil unrest (Phungula, 2024; Desai, 2023; Naidoo et al., 2023). South Africa’s protest tradition stretches back to the apartheid era, including the anti-pass law demonstrations that culminated in the Sharpeville-Langa massacre (Dubow, 2015; Coffey, 2022; Hollister, 2023), the 1976 Soweto Uprising against Afrikaans as the language of instruction (Molteno, 1979; Ndlovu, 2006), the Marikana labour strike and massacre in 2012 (Alexander, 2013; Naicker, 2016; Bond & Mottiar, 2013), and the Fees Must Fall student movement that challenged the rising cost of higher education (Langa, 2017; Griffiths, 2019). These protests reflect deeper structural grievances and unresolved conflicts, whether rooted in state repression, racial injustice, labour exploitation, or educational inequality. In this context, protest is not merely a reactionary act but a form of political participation through which citizens articulate resistance and contest state authority within the public sphere. As earlier sociological theories have suggested, protest is a global phenomenon that is not limited to South Africa or mainly Black communities. Around the world, demonstrations have occurred in various political, economic, and cultural settings. In Spain, for example, activists in Madrid, representing a racially diverse and urban population, occupied Plaza Puerta del Sol on 15 May 2011 to oppose austerity measures imposed by the political elite. Protesters set up temporary camps to demand better welfare systems and less corruption in financial and political institutions. While traditional media provided limited coverage, the movement, mostly driven by youth, used social media to mobilize widely. This made it one of the most significant political mobilizations outside of Spain’s formal labor unions and political groups (della Porta, 2015; Feenstra, 2015; Montero et al., 1997; Kaldor et al., 2015). Similarly, the United States has seen multiple waves of civil unrest throughout its democratic history. Most recently, university students across diverse campuses have protested antisemitism and the Israeli military campaign in Gaza (Tollefson, 2024). Studying protest is crucial since it provides policymakers, researchers, and community members with ways to understand collective action, reveal its motivations, and assess its effects in different socio-political contexts. This inquiry improves our understanding of the complex grievances behind protest movements, such as economic exclusion, political marginalization, and calls for justice and recognition (Tilly, 2004; Gurr, 1968). Instead of seeing protests as spontaneous or isolated occurrences, scholarly research places them in broader historical contexts and structural power relations, which allows for deeper interpretations of their causes and effects (Gamson, 1990). Importantly, empirical research on protest helps develop effective strategies for managing civil unrest. These strategies range from reducing immediate violence to creating long-term plans that tackle systemic inequalities, thus promoting stability in cities and strengthening societal resilience (Davies, 1962; Cloward & Piven, 1977). Scholarly studies of protest movements have used various methods to capture the complex nature of collective action. Quantitative methodologies, particularly surveys and structured questionnaires, enable researchers to measure predictors of protest participation, trace patterns of mobilization, and model the spatial or social diffusion of contentious episodes (Klandermans & Oegema, 1987; Opp & Gern, 1993). These methods support the empirical testing of theories such as relative deprivation, political opportunity structures, and resource mobilisation, facilitating the analysis of protest phenomena at scale (McCarthy & Zald, 1977; McAdam, 1983). Complementing this, qualitative approaches provide the contextual richness necessary to interpret protest through theoretical lenses such as framing, political process, and new social movement theory (Snow & Benford, 2005; Tarrow, 2022; Melucci et al, 1989). These frameworks are applied through case studies, ethnographies, and discourse analyses that investigate activist methods, identity formation, and the interpretive frameworks that support collective mobilization. Machine learning is one of the key areas of artificial intelligence. It offers a strong computational framework for analysing complex datasets and gaining insights into various social issues (Bishop, 2006; Goodfellow et al., 2016). Its main feature is that systems can learn from data by identifying patterns, making predictions, and improving performance on tasks without needing to be explicitly programmed for every situation (Samuel, 1959; Mitchell, 1997). Machine learning methods can be categorised into three primary types: supervised learning, unsupervised learning, and reinforcement learning (Sutton & Barto, 2018; Murphy, 2012). Unsupervised learning finds hidden structures in unlabelled data, usually through clustering. Reinforcement learning teaches agents to make the best sequential decisions based on feedback rewards. In contrast, supervised learning is commonly used in computer science and is increasingly applied in social science research. It allows researchers to learn from labelled input-output pairs to predict or classify future observations, providing a scalable alternative to traditional qualitative and survey methods. In supervised learning, a machine learning model is trained on a dataset comprising input instances (features) paired with corresponding output labels (target variables) (Hastie et al., 2009). The model, similar to a student, learns the relationship between features and labels by repeatedly processing these input-output pairs. During training, an optimisation algorithm adjusts the model’s internal parameters to reduce a loss function, which measures the difference between the predicted and actual labels (Goodfellow et al., 2016; Shalev-Shwartz & Ben-David, 2014). For example, in sentiment classification, the model connects text features like unigrams, bigrams, or other syntactic patterns with sentiment categories such as ‘positive,’ ‘negative,’ or ‘neutral.’ The model’s effectiveness is evaluated using both the training set and a separate validation set to ensure reliability and prevent overfitting. Once validated, it is tested on new data to check its generalizability (Bishop, 2006). A well-trained model can be utilised for various classification tasks, including spam detection and disease diagnosis. This study categorises social media discussions related to civil unrest. Supervised learning supports numerous important applications in Natural Language Processing (NLP), which aims to enable computers to understand and process human language (Manning & Schütze, 1999; Jurafsky & Martin, 2009). One such task is sentiment analysis, where a model trained on a labelled dataset (e.g., annotated tweets) can determine the emotional tone of new text inputs (Pang & Lee, 2008; Liu, 2012, 2022). Besides sentiment classification, supervised models are commonly used for topic classification (assigning texts to specific thematic areas), named entity recognition (identifying individuals, organizations, or locations), and hate speech detection. These tasks provide detailed insights into social discourse, enabling systematic analysis of public opinion, ideological positions, and behaviour patterns in digital communication (Aggarwal & Zhai, 2012; Fortuna & Nunes, 2018). Mooijman et al. (2018) employed machine learning techniques on social media data to investigate the rise of violent protests, specifically the clashes between protesters, police, and counter-protesters during the Baltimore unrest. The research examined how moral beliefs impact individual actions and social media networks in protest situations. It was discovered that how people view the moral legitimacy of a protest, both personally and through social reinforcement, is a crucial factor in driving collective mobilisation. The authors noted that deeply held religious or moral beliefs, when expressed in digital social networks, can increase individuals’ likelihood of engaging in civil disobedience during moments of moral outrage. In their study of the Baltimore protests, Mooijman et al. (2018) manually labelled 4,800 tweets as either moral or non-moral, creating a training dataset for a neural network that classified an additional 18 million tweets. The aim was to trace the moral arguments present in public discussions around the arrest and subsequent death of Freddie Gray, a 25-year-old Black man who suffered a fatal spinal cord injury while in police custody and was denied medical attention. Although six police officers were charged with serious offences, including second-degree murder and manslaughter, public outrage mounted over the perceived failure of the justice system. The study found that spikes in moral rhetoric were temporally aligned with the escalation of violence and correlated with increased protest-related arrests on an hourly basis. These findings demonstrate the potential of machine learning to operationalise hypotheses in protest research, ranging from measuring public sentiment and participation to uncovering the moral underpinnings of contentious political action. The Baltimore case underscores the complex intersections of racial injustice, police brutality, and institutional distrust, as reflected and amplified through social media discourse. In the study of protest movements, machine learning enables the processing and analysis of large-scale data, whether in textual, audio, or visual formats. This paper focuses on the protest activity that erupted in July 2021 following the incarceration of former President Jacob Zuma, using Twitter data to examine the evolving discourse and sentiment dynamics. The study contributes to social media analytics and protest research by demonstrating how social media data can be systematically collected and analysed to understand protest behaviour in South Africa. It presents sentiment as a mediating variable in the development and escalation of protest, traces the temporal spread of unrest through sentiment time-series analysis, and reveals provincial differences in participation through the lens of collective behaviour theory. These findings offer practical value to stakeholders, including policymakers and crisis response teams, by providing empirical evidence for anticipating, mitigating, and preventing episodes of civil unrest, thereby supporting safer and more resilient urban governance. The structure of the paper is as follows: Section 2 reviews the literature on theories of collective behaviour; Section 3 outlines the methodology; Section 4 presents the findings; and Section 5 concludes the study with implications and recommendations. Literature Review One important theoretical framework for understanding protest behaviour is the political opportunity structure (POS) model. This model examines how institutional arrangements and political settings influence collective action. POS theory suggests that the extent to which political systems enable or limit public participation influences the rise and direction of protest movements. According to this perspective, open political systems that allow public demands to be expressed through easy and responsive institutional channels may boost protest activity. They do this by legitimising dissent and making it easier to mobilise. On the other hand, more restrictive or closed systems can also trigger protest. When citizens face institutional barriers, they may turn to non-institutional methods to express their discontent. Some scholars advocate for a hybrid interpretation, suggesting that a mix of openness and constraint can create optimal conditions for protest. Eisinger's (1973) empirical study of American cities supports this perspective, revealing that protest activity was highest in localities characterised by a hybrid political structure, where limited openness coexisted with institutional friction, intensifying citizen mobilisation. Empirical evidence suggests that aggression frequently arises as a sociological response to deprivation, and that the application of military force is perceived as a form of deprivation, thereby intensifying frustration-aggression dynamics (Gurr, 1968). Within Gurr’s theory, relative deprivation is mediated by various psychological and social variables, including perceived injustice, group identity, political efficacy, and social cohesion, which collectively influence the magnitude of civil strife. While theoretically conceptualised as linear or sublinear influences in Gurr’s model, these mediating variables can be operationalised in machine learning frameworks as independent variables that predict protest behaviour. In such models, the magnitude of strife functions as the dependent variable or output. This computational framing allows researchers to empirically test the explanatory power of these variables using large-scale social data, an approach elaborated upon in the methodological section of this study. Between 1961 and 1965, Gurr’s seminal study analysed 1,100 strife events across 104 U.S. cities, identified from local newspaper sources. These events were systematically hand-coded, and the resulting variables were subjected to factor analysis. Six indicators of relative deprivation were consolidated into a composite index and statistically correlated with civil strife predictors. Gurr proposed a curvilinear relationship between the independent variable—relative deprivation, operationalised through coercive potential, institutionalisation, facility, and legitimacy—and various forms of civil unrest, including conspiracy, internal war, and turmoil. Notably, coercive force did not follow this curvilinear trend. Conceptually, the schematic model presented by Gurr parallels the architecture of a neural network, where independent variables (relative deprivation indicators) are trained to predict outputs (dimensions of civil strife). The model learns the relationship between these variables, ultimately revealing that more significant relative deprivation corresponds to increased unrest intensity. Importantly, Gurr’s framework excludes revolution from the scope of strife, focusing instead on forms of unrest that emerge within existing political systems. The basis for protest behaviour is often explained through grievance theory. This framework suggests that collective action typically arises from ongoing, unresolved societal complaints. These complaints often stem from systemic inequalities or perceived unfairness. Issues such as division among elites, problems within institutions, or the exclusionary nature of political systems can exacerbate these grievances. According to the theory, protests occur when these grievances align with openings in the political landscape, such as leadership gaps, conflicts among elites, or a decline in trust in the regime. This situation allows previously marginalised voices to organise. Economic factors, such as high unemployment, poverty, or austerity measures, can exacerbate feelings of dissatisfaction, particularly when government structures fail to address pressing public issues. Therefore, grievance theory highlights how the interaction between structural limits and lack of political openness can trigger mass mobilisation, especially when institutional redress lacks effectiveness. Methodology This study takes a computational social science approach, using natural language processing (NLP) techniques to analyse digital trace data from social media. The methodology includes several connected parts. It starts with collecting Twitter data on the July 2021 protests in South Africa. After acquiring the data, initial exploratory analyses cleaned and organised it, which helped with sentiment classification and topic extraction. Researchers annotated the dataset to train supervised learning models for sentiment classification. They then applied unsupervised topic modelling algorithms to find the main themes in the discussions. Temporal analysis was followed to track the development and spread of the unrest over time. The insights from these NLP tasks were compared with traditional media coverage and official reports to ensure they were accurate and added depth to the interpretation. This mixed-method strategy offers a solid framework for examining the digital aspects of protest movements and understanding how online discussions relate to offline events. Data Collection During the early days of the July 2021 unrest, this study collected Twitter data using a specific set of hashtags related to the protests. Hashtags like #JulyUnrest, #Looting, #FreeJacobZuma, #LootingSA, #ShutdownSA, and #PhoenixMassacre were used to gather relevant tweets from Twitter. The data collection method took geography into account, allowing us to categorise tweets by province. We sourced tweets from all ten provinces, with a strong focus on KwaZulu-Natal (Durban), where the protests began, and Gauteng (Johannesburg and Pretoria), which later saw significant unrest. We also collected additional data from other provinces, including North West (Mahikeng), Western Cape (Cape Town), Eastern Cape (Port Elizabeth), Limpopo (Thohoyandou), Free State (Mangaung), and smaller municipalities in KwaZulu-Natal, such as Umhlathuze. A distinct dataset labelled "FreeJacobZuma" was created due to the high frequency of tweets containing that hashtag. The study generated 11 provincial and thematic subsets of data, covering the protest activity from 8 July to 22 July 2021. Table 6 summarises the dataset names, tweet volumes, and relative proportions across subsets. Table 6 Name of datasets and their proportion Datasets Number of samples Percentages (%) Cape town 1427 1.97 FreeJacobZuma 24747 34.21 Johannesburg 5448 7.53 Mahikeng 77 0.11 Mangaung 159 0.22 PhoenixMassacre 36333 50.23 Port Elizabeth 285 0.39 Pretoria 1952 2.70 Thoyandou 41 0.06 Umhlathuzi 76 0.11 Total 72332 100 Data Preprocessing Extensive preprocessing was necessary to ensure data quality and analytical relevance before the harvested tweets could be used for sentiment classification and topic modelling. In adherence to Twitter’s Developer Policy and to protect user privacy, all tweets were collected without personally identifiable information. The raw tweets often contained unstructured and noisy elements such as numerals, abbreviations, internet slang, typographical errors, special characters, hyperlinks, hashtags, and other non-informative tokens. These components were systematically removed to improve textual coherence. Following this cleaning process, the tweets were tokenised, breaking each tweet into discrete word units. Then, they were lemmatised using the Natural Language Toolkit (NLTK) library to standardise the words to their base forms. This normalisation process ensured consistency in word representation across the corpus. The clean textual data was then vectorised using the Term Frequency-Inverse Document Frequency (TF-IDF) method, which transforms the corpus into a numerical format, emphasising informative terms while down-weighting frequent but uninformative words. The resulting matrix was used as input for subsequent machine learning tasks, such as sentiment classification and topic extraction. Data Labelling Supervised machine learning needs a dataset of input-output pairs, where each input (in this case, a tweet) has a corresponding output label that defines sentiment polarity. In this study, each tweet received a single sentiment label, which can either be positive, negative, or neutral, with no overlap between categories. This means a tweet could only belong to one of the three sentiment classes. This classification setup allowed the training of a machine learning model to learn the relationship between text inputs and their corresponding sentiment outputs. The model updated its internal parameters by minimising prediction errors across the training samples. Once trained, the model could generalize and classify new tweets into the right sentiment class with fewer errors. To support this supervised training process, about 5,000 tweets (6.91% of the 72,332 tweets collected) were manually labelled by human raters, creating a high-quality dataset. This labelled subset was the basis for training and evaluating the sentiment classification models used in the study. Three independent annotators manually reviewed the tweet dataset to ensure reliable sentiment labelling and assigned each tweet a positive, negative, or neutral sentiment label. Annotators worked independently to minimise potential bias or influence from others, thereby helping to maintain objectivity in the labelling process. A detailed annotation guide was provided to standardise the labelling procedure and help annotators deal with ambiguous or complex tweets. The guide also outlined a method for handling cases where annotators disagreed. When all three annotators disagreed on a tweet's sentiment, a fourth party was involved to resolve the disagreement. This adjudicator reviewed the tweet and the labels given, making a final decision. The final output included a labelled dataset with the original tweets, individual annotations, and a consolidated master sheet that showed the adjudicated sentiment label. This process ensured consistency and validity in the labelling, which is vital for training effective supervised machine learning models. To evaluate the reliability of the sentiment labels created by the annotators, the study calculated inter-annotator agreement using Fleiss’ Kappa statistic. Previously, tweets with missing, incomplete, or misspelt labels, as well as those that did not meet the annotation guidelines, were removed from the evaluation. This filtering process dropped the total number of usable labelled tweets from 4,916 to 4,847. The Kappa score from these 4,847 tweets was 0.2737, indicating a fair level of agreement among the three annotators. Although the score indicates a reasonable level of consistency in the annotators' judgments, it also highlights significant variation in their interpretations. To address these variations, the final label for each tweet used in training the sentiment analysis model was decided by a majority vote among the three annotators. This method balances personal judgment with systematic aggregation, ensuring the labelled dataset remains solid for supervised learning. Data Modelling Following the establishment of fair inter-annotator agreement, the study used a majority voting system to finalise the sentiment labels for each tweet. For each sample, the selected label was the one agreed upon by at least two annotators. This method effectively draws on the judgment of the annotators, reducing individual bias and improving the overall reliability of the labelling process. When annotators disagreed and no clear majority emerged, an adjudicator reviewed the tweet and assigned the final label. This process ensured that all 4,916 samples could be retained for training, including those with initially disputed labels. The finalised dataset, comprised of tweet-input and sentiment-label pairs, was used to train a supervised machine learning model to classify unseen tweets as positive, negative, or neutral. Prior to training, the distribution of sentiment classes was assessed to verify class balance, thereby avoiding model bias toward majority classes. Table 7 presents the frequency distribution of sentiment labels in the dataset. Table 7 Sentiment label distribution Sentiment class Sentiment samples Sentiment proportion % Negative 3063 62.31 Positive 1099 22.36 Neutral 752 15.30 Total samples 4916 99.97 Sentiment distribution of class labels From the table above, the sentiment distribution of the annotated dataset reveals a significant class imbalance. Table 7 shows that of the 4,916 tweets, 3,063 (62.31%) were classified as expressing negative sentiment. In contrast, only 1,099 tweets (22.36%) were labelled positive, and 752 (15.30%) were neutral. This skewed distribution suggests that negative sentiment predominates the dataset, posing a challenge for training a machine learning model. When one class significantly outweighs others, models tend to be biased toward the majority class, compromising the accuracy and fairness of the classification task. To address this issue, the study employed the Synthetic Minority Over-sampling Technique (SMOTE), which generates synthetic samples for the minority classes. By rebalancing the dataset, SMOTE ensures that the model does not overlearn from the overrepresented negative class. This improves the generalisation and predictive accuracy of the sentiment classifier. To create the sentiment classification model, five popular supervised learning algorithms were chosen. The models chosen are Logistic Regression (LR), Gaussian Naïve Bayes (GaussianNB), K-Nearest Neighbours (KNN), Support Vector Machines (SVM), and Extreme Gradient Boosting (XGBoost). These models were selected for their effectiveness in multi-class text classification tasks. The preprocessed and SMOTE-rebalanced dataset was divided into training and testing subsets. The training data allowed the models to learn the connection between tweets (input features) and their corresponding sentiment classes (targets). As a multi-class classification problem, the model predicted one of three mutually exclusive sentiment labels for each tweet: positive, negative, or neutral. This experimental setup enabled an empirical comparison of classifier performance, evaluated using the metrics described in the following section. Table 8 Classifier performance of Sentiment models Model performance (F1-score, macro) Base models Train Test Logistic regression 0.80 0.82 GaussainNB 0.81 0.84 KNN 0.46 0.48 SVM 0.86 0.88 XGBoost 0.76 0.77 Table 8 presents the classification performance of the machine learning algorithms used to develop the sentiment model. Logistic regression, Gaussian Naive Bayes, and support vector machines achieved approximately 80% accuracy in classifying tweets, demonstrating strong predictive performance. XGBoost followed closely with a slightly lower accuracy of 76%. By contrast, the K-nearest neighbours (KNN) algorithm exhibited the weakest performance, with a notably high error rate and an accuracy of only 54%. These performance metrics were validated using five-fold cross-validation, a robust technique that partitions the dataset into five equal subsets. In each iteration, four subsets are used for training while the remaining subset is used for testing. This process is repeated five times, ensuring each subset is used once as a test set. The average F1-score across all folds provides a reasonable estimate of the model's generalizability and helps mitigate overfitting. This study examined a range of established topic modelling algorithms to identify hidden themes in the tweet collection, where each tweet served as a document. Commonly used models in this area include Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorisation (NMF), and Latent Semantic Analysis (LSA). The study compared these models to see how well they could extract clear and understandable topics related to protest discussions on Twitter. However, it presented detailed results only from the most effective model. The topics revealed by this model highlighted the concerns and key issues of the protesters, shedding light on the reasons behind collective action. These findings were viewed through the lens of collective behaviour, allowing for a better understanding of how social unrest develops in response to shared concerns. Results This section presents the study's results. The analysis begins with an exploratory analysis of the dataset used to study the protest movement. It outlines how well the sentiment classification model performed in assessing public sentiment, with a focus on perceptions and the emotional tone of discussions related to protests. Subsequently, the section outlines the application of topic modelling to extract dominant themes, offering insight into the protesters' key grievances and focal points. An initial dataset exploration offers preliminary insights into the July 2021 protest movement as observed on Twitter. Descriptive statistics summarise and quantify the characteristics of the data collected, highlighting structural and contextual patterns within the corpus. Common statistical descriptors such as the total number of tweets, average tweet length, standard deviation, range, interquartile range, and sentiment label distribution are typically employed to characterise the dataset. Additional descriptive features such as the frequency of hashtag usage, retweet counts, and the number of likes may enrich the analysis. However, these indicators were sparsely distributed in this dataset, and the number of unique users was omitted in compliance with Twitter’s data privacy guidelines. Table 9 Geographical distribution of dataset Datasets Number of samples Percentages (%) Province Cape town 1427 1.97 Western cape Durban 1787 2.47 KwaZulu-Natal FreeJacobZuma 24747 34.21 Not applicable Johannesburg 5448 7.53 Gauteng Mahikeng 77 0.11 North West Mangaung 159 0.22 Free State PhoenixMassacre 36333 50.23 KwaZulu-Natal Port Elizabeth 285 0.39 Eastern Cape Pretoria 1952 2.70 Gauteng Thohoyandou 41 0.06 Limpopo Umhlathuzi 76 0.11 KwaZulu-Natal Total 72332 100 Table 9 shows the geographical distribution of tweets across different South African provinces and cities. Conversations referencing Phoenix constituted the largest dataset share, accounting for approximately 50.23% of all tweets collected. This significant volume is attributed to the violence and fatalities associated with the Phoenix massacre during the unrest. The FreeJacobZuma dataset was followed by 34.21%, reflecting the widespread public support and online mobilisation advocating for the release of former President Jacob Zuma. Johannesburg contributed 7.53% of the total conversation, while Pretoria and Cape Town accounted for 2.70% and 1.97%, respectively. Smaller percentages of tweets originated from other cities that were less directly involved in the protest action. This geographic variation in tweet volume offers insight into regional engagement and participation in the protest discourse, suggesting a possible correlation between online attention and on-the-ground mobilisation. Table 10 Twitter dataset descriptive statistics Mean Standard deviation Words Characters Words Characters Datasets Before cleaning After cleaning Before cleaning After cleaning Before cleaning After cleaning Before cleaning After cleaning Cape town 26 24 163 134 15 15 89 82 Durban 24 22 152 120 14 14 83 87 FreeJacobZuma 17 14 131 78 13 13 83 73 Johannesburg 23 21 147 117 14 14 84 78 Mahikeng 25 23 159 128 12 12 77 69 Mangaung 24 22 147 116 15 15 88 81 PhoenixMassacre 22 18 154 102 14 14 87 79 Port Elizabeth 22 19 140 108 13 13 76 71 Pretoria 23 21 143 113 14 14 82 76 Thohoyandou 27 25 161 136 15 15 80 78 Umhlathuzi 17 16 117 90 13 13 82 74 Table 10 presents summary statistics that describe the key characteristics of the Twitter dataset, specifically the average and standard deviation of word and character counts before and after data cleaning. These metrics offer insight into the expressiveness and variability of tweets across geographic locations. For instance, tweets from Thohoyandou and Cape Town exhibit the highest average word counts, such as 27 and 26 words per tweet, respectively, suggesting a relatively more expressive discourse in those regions. Mahikeng follows with an average of 25 words, while tweets from Mangaung and Durban average 24 words. Other provinces generally have an average of 22 to 23 words. This reflects the overall average across all regions. Interestingly, the FreeJacobZuma dataset shows a clear difference, with an average of 17 words per tweet. This shorter length suggests that expressions of frustration or calls to action might be shorter, possibly indicating stronger emotions. The average character counts in all subsets stay within Twitter’s 280-character limit. This shows that users can usually express their thoughts within the platform’s limits. The standard deviation for word counts is between 13 and 15, while character counts range from 76 to 89. This shows a moderate variation in tweet lengths across the dataset. These patterns help us understand the structure and dynamics of protest-related discussions. The statistics for word and character counts after cleaning show little change from the values before cleaning. This suggests that the text preprocessing, which included removing stop words, punctuation, and other non-essential tokens, did not significantly alter the main content of the tweets. The retained linguistic structure suggests that the preprocessing preserved the semantic integrity of the dataset, ensuring that the machine learning models were trained on meaningful and representative textual information. This outcome affirms the efficacy of the cleaning procedure, demonstrating that it effectively filtered noise without discarding valuable content necessary for sentiment analysis and topic modelling tasks. Table 11 Sentiment distribution of tweets Number of samples Proportion (%) Negative labels 3063 62.33 Positive labels 1099 22.36 Neutral labels 752 15.30 Total 4914 100 Table 11 presents the sentiment distribution of the labelled subset of tweets, representing 6.9% (5,000 out of 72,332) of the total dataset. Within this annotated sample, negative sentiment accounted for 62.33% (3,063 tweets), positive sentiment constituted 22.36% (1,099 tweets), and neutral sentiment comprised 15.3% (752 tweets). Figure 15 displays a histogram visualising this distribution, where red bars indicate negative sentiments, green bars represent positive sentiments, and blue bars denote neutral sentiments. The histogram underscores the dataset's significant class imbalance: for every one positive tweet, there are approximately three negative tweets, and for every neutral tweet, there are roughly eight negative tweets. This imbalance poses potential challenges for sentiment classification tasks and necessitates rebalancing techniques to ensure fair model training. To train the sentiment analysis model, five widely recognised classifiers were employed: Logistic Regression, Gaussian Naive Bayes (GaussianNB), K-Nearest Neighbours (KNN), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost). These models were selected for their established effectiveness in multi-class classification tasks. Table 12 presents the performance metrics of each classifier after being trained on the labelled dataset. The models were evaluated on their ability to accurately classify tweets into three sentiment categories: positive, negative, and neutral. Table 12 Performance of the sentiment classification model Model performance (F1-score, macro) Base models Train Test Logistic regression 0.78 0.79 GaussainNB 0.81 0.85 KNN 0.42 0.46 DecisionTree 0.67 0.71 SVM 0.86 0.87 XGBoost 0.78 0.78 Table 12 presents the performance metrics of the sentiment classification models. Logistic Regression, Gaussian Naive Bayes, and Support Vector Machines each achieved approximately 80% classification accuracy, demonstrating their effectiveness in sentiment prediction. XGBoost performed slightly lower, with an accuracy of 76%. In contrast, K-Nearest Neighbours (KNN) exhibited the weakest performance, with an accuracy of 54%, indicating a high error rate and limited generalisability for this task. The results reflect F1-scores obtained through five-fold cross-validation, a technique that mitigates overfitting by dividing the dataset into five equal parts. In each iteration, four folds are used for training while the remaining fold is reserved for validation. This process is repeated across all five folds, ensuring a robust estimate of model performance across different data partitions. The sentiment analysis model developed using the best-performing classifier was employed to label the remaining 93.9% (67,915 out of 72,332) of previously unlabelled tweets. Following the preprocessing steps, including cleaning and deduplication, the unlabelled dataset was reduced to 67,326 tweets. Table 13 presents the sentiment prediction results generated by the selected model, specifically the Support Vector Machine (SVM) classifier. The SVM model classified a substantial portion of the tweets as expressing negative sentiment, reinforcing the dominant tone observed in the annotated subset. Figure 16 offers a histogram visualisation of the SVM-based sentiment distribution, clearly representing the model’s output across the negative, positive, and neutral categories. Table 13 Distribution of Sentiment Labels in SVM Model Predictions on Unlabelled Tweets Number of samples Proportion (%) Negative labels 56070 83.28 Positive labels 1449 2.15 Neutral labels 9807 14.57 Total 67326 100 As illustrated in Table 13 and Figure 16, the SVM model predicted a dominant proportion of tweets, 83.28% (56,070 out of 67,326), as negative. This aligns with the earlier annotation in Table 11, where 60.33% of the manually labelled tweets were also negative. The consistency across labelled and predicted datasets, reinforced using SMOTE for class rebalancing and five-fold cross-validation to minimise bias, suggests a strong underlying narrative of discontent within the data. It is plausible that the prevalence of negative sentiment stems largely from the Phoenix subset, which constituted 50.23% of the entire dataset (as shown in Table 11). This region witnessed a significant escalation of the July 2021 unrest, marked by racially charged violence between Black South Africans and Indian residents. Thus, the overwhelmingly negative sentiment captured by the SVM model reflects the upheaval, looting, and civil tension that dominated Twitter discourse during the period. Without prejudice to the predominantly negative sentiment identified by the classifier, the study further explored the dataset's content by examining the most frequently occurring words. A bigram histogram was generated to visualise the top word pairs, offering insight into the dominant themes and summarising the discourse captured in the dataset. The unlabelled subsets underwent a comprehensive cleaning process to ensure the analysis provided meaningful and interpretable results. This included the removal of URLs, HTML tags, emojis, emoticons, usernames, hashtags, special characters, punctuation marks, and stop words. The resulting cleaned corpus served as the basis for identifying salient terms that reveal the underlying concerns and topics discussed during the protest. The bigram histogram presented in Figure 17 visualises the top 20 word pairs from the Cape Town dataset. Notably, the name "Jacob Zuma" appeared more than 50 times, reflecting the centrality of his jailing to the protest discourse. These words show that former President Jacob Zuma is at the centre of the discussion surrounding the protest movement. Variations of the name also occurred multiple times with differing frequencies, indicating diverse contexts of mention. Other frequently occurring bigrams include “South Africa,” “taxi violence,” and “Cape Town,” each appearing more than 40 times, suggesting these were significant components of the localised protest narrative. Less frequent but still relevant bigrams, each appearing fewer than 20 times, include “inciting violence,” “state capture,” “Cyril Ramaphosa,” and “people loot,” pointing to broader socio-political concerns linked to governance, leadership, and public frustration. These words show that former President Jacob Zuma is at the centre of the discussion surrounding the protest movement. The bigram histogram from Durban, the epicentre of the July 2021 protest, presented in Figure 18, identifies “Jacob Zuma” as the most frequently occurring term, appearing over 60 times in the dataset. Other notable phrases, such as “South Africa,” “Black people,” and “people looting,” appear with lower frequency counts, each under 50. The name “Bheki Cele,” South Africa’s Minister of Police, surfaces in connection with widespread criticism of the security cluster’s inability to control the unrest, as reports indicated the police force was overwhelmed by looters. Specific locations, such as KwaMashu Shopping Centre, were mentioned in relation to the extensive damage, including looting and arson, as confirmed by local media sources. Additionally, “Sihle Zikalala,” the Premier of KwaZulu-Natal and a senior member of the ruling African National Congress (ANC), was frequently cited in public discourse surrounding the province’s response to the unrest. Figure 19 presents a bar chart based on the Free Jacob Zuma subset of the dataset. The terms “President Zuma” and “South Africa” appear exceptionally frequently, each occurring over 600 times, highlighting the centrality of Zuma’s imprisonment to the protest discourse. The widespread use of the hashtag #FreeJacobZuma underscores the public’s demand for his release and reflects his symbolic role in mobilising collective action during the unrest. As the protest spread from Durban, it began to diffuse into other urban centres within Gauteng Province. Figure 19 displays the most frequently occurring terms in the Johannesburg dataset, where “South Africa” and “Jacob Zuma” emerged as the most prominent, each appearing over 100 times. Other notable expressions, occurring at roughly half the frequency, included “inciting violence,” “people looting,” “black people,” “South Africans,” and “free Zuma.” Additional terms with moderate frequency—under 50 mentions—were “President Zuma,” “security cluster,” “Cyril Ramaphosa,” and “stop looting,” reflecting both political and public safety concerns circulating during the unrest. Figure 21 presents a histogram of the most frequently occurring bigrams in the Mahikeng dataset. The phrases “stolen goods,” “South Africa,” “South African,” and “President Zuma” appeared three times, indicating their centrality to the discourse. Other notable bigrams appearing twice include “country right,” “Zuma president,” “I’m happy,” and “stop looting.” The phrase “people hungry” reveals an underlying concern about the province's food insecurity and economic deprivation. Similarly, “let protect” suggests calls for safeguarding public infrastructure, while “anger towards” signals a directed frustration, though the object of that frustration varies across the dataset. Figure 22 presents the histogram of the most frequently occurring phrases in the Mangaung dataset. “Jacob Zuma” emerged as the most frequently used phrase, appearing six times, underscoring the centrality of his imprisonment to the discourse. The phrase “looting won’t” reflects local disapproval of looting activities, suggesting a perceived disconnect between such actions and their influence on judicial outcomes. Other frequently appearing phrases, such as “won’t impact,” “impact SA,” “SA economy,” “Free State,” “people looting,” “South Africas,” and “black people”, each appeared more than three times. These expressions highlight both geographical references and socio-political concerns. Less frequent phrases, occurring approximately twice, add further nuance to the conversation, indicating mixed public sentiment around the protest, its perceived consequences, and broader national identity. Figure 23 presents the frequency distribution of the most used phrases in the Phoenix Massacre dataset. The phrase “Black people” appeared over 3,000 times, indicating the racial dimension central to the discourse. “South Africa” was mentioned approximately 1,000 times, reflecting the national significance of the event. Phrases such as “people killed,” “Bheki Cele,” and “Cyril Ramaphosa” appeared around 500 times each, highlighting the public’s focus on the loss of life and the role of state leadership in managing the crisis. Additional frequently occurring terms, as visualised in the histogram, further underscore the intensity and specificity of public engagement surrounding the events in Phoenix. Figure 24 displays the histogram of the top 20 most frequently occurring phrases in the Port Elizabeth dataset. The overall low frequency of terms suggests minimal engagement from the Port Elizabeth community in the July 2021 protest. The most prominent phrase, “Jacob Zuma,” appeared 15 times, reflecting the centrality of the former president in the discourse. “South Africans” and “let burn” appeared five times each, while other expressions such as “South Africa,” “Free Zuma,” “former president,” and “President Jacob” occurred with varying lower frequencies. Notably, the histogram also captures discussions around “taxi violence,” references to “Zuma jail,” and the provocative phrase “must loot,” indicating the presence of isolated incitements despite the region’s overall limited participation. Figure 25 presents the frequency distribution of the top 20 bigrams in the Pretoria dataset. Only two phrases, “Jacob Zuma” and “South Africa”, appeared more than 40 times, underscoring the central role of Zuma’s imprisonment in triggering the unrest and the national context in which it unfolded. Phrases occurring more than 20 times include “people looting,” “President Zuma,” “former President,” “South Africans,” “inciting violence,” “black people,” “Zuma must,” and “Zuma arrested.” Additional expressions with lower frequency, such as “rule law,” “looting shops,” “people hungry,” “poor people,” “go loot,” and “looting burning,” further highlight key themes in the discourse. Notably, phrases like “people hungry” and “poor people” suggest that economic hardship may have contributed to looting participation, while “looting burning” reflects the intensity of unrest through the destruction of property. Figure 26 presents the variation in frequently occurring bigrams within the Thohoyandou dataset. The most prominent phrase is “mislead JZ,” which appears more than six times, indicating a potential narrative that accuses individuals or groups of misleading former President Jacob Zuma. “Former President” occurred approximately three times, reaffirming the centrality of Zuma in the discussions. Phrases such as “hong hong” were likely a local language expression, and “rights limited” reflects localised discourse and concerns about civil liberties. Mentions of “Edward Zuma” may carry sarcastic undertones aimed at the former president’s family. Expressions like “dilemma poor,” “poor people,” “people know,” “know feels,” “feels deprived,” and “deprived livelihood” articulate the socio-economic challenges facing the population, though some terms appeared misspelt. Additional phrases such as “livelihood jobs,” “job retails,” and “retail loss” point to economic consequences, including employment disruption and business losses during the protest. Finally, phrases like “looting men,” “men black,” and “black majority,” despite inconsistencies in spelling, suggest that discussions in the dataset emphasised the participation of Black men in the looting incidents, reflecting community-level observations during the July 2021 unrest. Figure 27 illustrates the low-frequency distribution of phrases in the Umhlathuzi dataset, indicating minimal discourse related to the July 2021 unrest. Most phrases appear only once or twice, suggesting limited public engagement on the topic in this region. Among the few recurring phrases are “South Africa” and “President Zuma,” each appearing twice, which may indicate a peripheral awareness of the national events. Phrases such as “white sovereignty” and “white acts” introduce a racial undertone, although their low frequency renders them statistically insignificant in this dataset. Additionally, the phrase “lorch donda” serves as a casual expression, whereas several other terms are in local languages or do not align with the broader protest narrative. The sparse and disconnected nature of the discussions suggests that the unrest did not greatly impact Umhlathuzi or that the protest movement did not spread into this province or area. The limited presence of protest-related conversations in Umhlathuzi shows how topic modelling can serve as a tool for real-time monitoring. This can help identify civil unrest early and prevent it. State actors can identify emerging hotspots of discontent by monitoring the frequency with which specific keywords and phrases co-occur. These keywords might include names of political figures, mentions of violence, or social and economic complaints. This proactive method enables timely actions, such as effective communication, community involvement, or deploying resources to reduce unrest. For instance, the early presence of terms like “President Zuma” or “white sovereignty” in otherwise unaffected regions could be precursors to ideological mobilisation. Therefore, integrating topic modelling into digital surveillance frameworks provides a scalable and non-intrusive mechanism for understanding the geographical and thematic evolution of protest movements, thereby enhancing national crisis preparedness and response strategies. This study adopts a semi-supervised learning approach to extend sentiment annotation beyond the gold-standard human-labelled subset. After training and evaluating the models, five widely used classifiers were selected. These selected classifiers were Logistic Regression, Gaussian Naive Bayes, K-Nearest Neighbours (KNN), Support Vector Machines (SVM), and XGBoost. A comparative performance analysis was conducted using accuracy and F1-scores derived from five-fold cross-validation. The results revealed SVM as the most effective model for multi-class sentiment classification in this context, exhibiting the highest generalisation performance on training and unseen data. Leveraging this outcome, the trained SVM model was used to classify the remaining unlabelled 93.9% of the protest-related tweets. This semi-supervised pipeline, in which a relatively small, annotated dataset guides the labelling of a much larger corpus, exemplifies a scalable strategy for sentiment monitoring when manual annotation is prohibitively time-consuming or expensive. Comparative analysis ensures that only the most reliable classifier is deployed for large-scale sentiment estimation, thereby reducing the risks of misclassification that could compromise subsequent analysis of protest dynamics. A probability threshold was introduced during sentiment prediction to refine the interpretability of sentiment trends further and prevent false positives in identifying potential unrest. Specifically, only sentiment classifications with confidence scores exceeding 0.8 were considered in downstream analyses. This thresholding strategy filters uncertain predictions, ensuring that only high-confidence sentiment outputs influence protest diffusion tracking. This prevents overreaction to sporadic or weakly expressed sentiments that may not reflect the public's mood. Moreover, this method offers a robust early warning system by correlating temporally aggregated sentiment shifts with topic modelling outputs. Intervention strategies can be triggered when consistently high levels of negative sentiment coincide with discourse around grievance-related themes, such as economic hardship, state repression, or racial tension. Thus, combining semi-supervised classification, comparative model selection, and probabilistic thresholding provides a replicable framework for real-time sentiment monitoring in crisis contexts. Non-negative Matrix Factorisation (NMF) was employed in this study as the topic modelling algorithm to extract latent themes within tweets related to the July 2021 protest movement. The model was applied to tweets originating from KwaZulu-Natal, the epicentre of the unrest, and tweets from provinces where the protest later diffused, specifically Johannesburg and Pretoria (see Tables 9, 10, and 11, respectively). NMF works by factorising the document-term matrix into two non-negative matrices: one representing topics as rows and the other representing documents as columns. In this framework, each tweet is viewed as a composite of multiple topics, where each topic is represented by a distribution of terms. The dimensionality reduction inherent in NMF transforms tweets into sparse, non-negative vectors, making it especially suitable for modelling social media content, which is typically characterised by short text length and thematic fragmentation. Table 14 Thematic analysis of KwaZulu-Natal tweets SN KwaZulu-Natal Topics Interpretation 1 Looting, stop, looters, stores, shops, busy, burning, continues, said, mall. This topic centres around the looting and destruction of shops and malls during the unrest. It describes on the ongoing chaos, with mentions of looting, burning, and the government’s call to stop the looters. 2 Zuma, free, released, arrested, release, dudu, support Supporters of Zuma, including his daughter Dudu, continue to call for his release, saying he should no longer remain arrested, while many think the movement to free him will only grow stronger. 3 loot, didn’t, left, food, love, want, responsibly, looters, excuse, money Many looters didn’t leave behind any food or money, using their love for material things as an excuse, while others argue they acted out of desperation and not irresponsibility. 4 people, black, indians, white, longer, phoenix, killed, saying, stores, thing Tensions have escalated in Phoenix, with reports of violence between black and Indian communities, where people are no longer just saying, but acting, leading to stores being destroyed and lives being lost. 5 sandf, deployed, kzn, need, come, deploying, really, deployment, soldiers, good The deployment of SANDF soldiers in KZN has been seen as necessary, with many saying that the presence of the military is really needed to restore order, and that the deployment is a good move. 6 violence, chose, taxi, wena, inciting, kzn, acts, abantu, phoenix, cause Taxi drivers in KZN have been accused of inciting violence, particularly in Phoenix, where acts of brutality against abantu (people) have caused widespread fear and chaos 7 police, security, looters, minister, army, cluster, private, area, mall, kwamashu The police, together with private security forces, have been deployed to areas like KwaMashu malls, with the Minister urging the army and security cluster to protect the malls from looters. 8 jacob, president, cyril, zumas, ramaphosa, free, release, arrest, protests, kzn The arrest of former president Jacob Zuma has led to widespread protests in KZN, with calls for President Cyril Ramaphosa to release him, intensifying the crisis. 9 don’t, unrest, saps, know, time, protest, think, im, want, country Many citizens don’t know how long the unrest will last, but they think it's time for the SAPS to take control, as protests continue, and people express their frustration with the state of the country 10 south, africa, durban, happening, right, protesting, drive, kwazulunatal, lotus, africans In Durban and across South Africa, protesting is happening right now, with Africans driving the unrest in areas like KwaZulu-Natal and Lotus Table 15 Thematic analysis of Johannesburg tweets SN Johannesburg Topics Interpretation 1 looting, stop, police, looters, shops, say, busy, poor, gauteng, criminality Looting in Gauteng has escalated, with the police struggling to stop looters targeting shops, as criminality thrives amidst the busy and poor conditions. 2 zuma, free, jacob, jail, prison, release, president, ramaphosa, longer, he’s Debates continue over Jacob Zuma's imprisonment, with some calling for his release, while President Ramaphosa is pressured to take a stance 3 loot, come, want, na, responsibly, wan, money, didn’t, cup, looters Looters, driven by desperation, want money and resources, though they didn’t consider the consequences of their actions, ignoring calls to act responsibly 4 violence, inciting, taxi, stop, public, arrested, account, incitement, report, peace Violence continues as some incite further unrest, particularly among taxi drivers, leading to arrests and reports calling for peace 5 people, black, hungry, steal, white, stealing, unemployed, government, going, jobs Economic inequality is evident as hungry and unemployed black people resort to stealing, while the government struggles to create jobs and prevent further unrest. 6 sandf, deployed, deploy, members, saps, need, police, cops, government, situation The government has deployed SANDF and SAPS members to address the escalating situation, emphasizing the need for a stronger police presence 7 south, africa, unrest, africans, african, pray, peace, world, jacob, week As unrest continues in South Africa, Africans across the world pray for peace, reflecting on a turbulent week centred around Jacob Zuma 8 security, don’t, country, president, amp, like, know, anc, think, time Security concerns are mounting in the country, with many questioning the ANC's leadership and the president’s ability to manage the situation in time 9 protest, peaceful, free, want, criminality, say, level, join, there’s, yes Calls for peaceful protest are growing, as people want to express their desire for change without descending into criminality, urging others to join the movement. 10 mall, glen, jabulani, soweto, maponya, protea, shoprite, ridge, durban, joburg Malls like Maponya, Jabulani, and Shoprite in Soweto, Durban, and Joburg have been severely impacted by the unrest, with Glen and Protea Ridge among those affected Table 16 Thematic analysis of Pretoria tweets SN Pretoria Topics Interpretation 1 zuma, jacob, jail, prison, anymore, free, arrested, think, I’m, released Jacob Zuma is in jail, and people are debating whether he should remain in prison or be released, with many thinking his arrest is unjust. 2 looting, stop, started, shops, police, south, burning, busy, they’re, say Looting has erupted across South Africa, with shops burning and the police struggling to stop the unrest that has already started. 3 loot, didn’t, na, left, need, right, hope, burn, gon, vho Looters didn’t leave much behind, driven by a desperate need, hoping they can get what they need before everything burns. 4 violence, inciting, stop, role, taxi, arrested, wena, incite, started, need Violence has escalated, with individuals inciting others, particularly within the taxi industry, leading to arrests as authorities attempt to stop the chaos. 5 people, hungry, black, poor, going, busy, like, died, arrested, killing The unrest is deeply rooted in social inequality, with hungry and poor black people feeling like they have no other option, leading to deaths and arrests 6 don’t, steal, know, care, play, things, think, want, understand, let There’s a growing sense of apathy, with many not caring about the consequences of stealing and feeling like people just don’t understand or care anymore. 7 president, law, like, africa, south, jacob, ramaphosa, zumas, release, want The president is under pressure to uphold the law, with many in South Africa wanting Ramaphosa to release Jacob Zuma 8 protest, stealing, going, understand, jail, sa, tht, today, unemployment, funded Economic desperation has led to protests, with people resorting to stealing and feeling like they’re heading for jail if things don’t improve in South Africa 9 security, minister, state, need, police, tech, smarti, systems, that’s, guard The need for enhanced security is clear, with the Minister emphasizing the importance of police and advanced tech systems to guard against further unrest 10 sandf, saps, deployed, kzn, amp, I’m, deployment, police, deploy, ba The SANDF and SAPS have been deployed in KZN, with observers noting how the deployment of these forces will impact the situation The diffusion of collective violence refers to how civil unrest and protest activity propagate in waves, facilitated by communication networks, protest opportunities, and underlying socio-economic inequalities that mobilise public demand for change. This study's findings indicate that social media played a central role in the spread of protest during the July 2021 unrest in South Africa. With an estimated internet penetration rate of 72% and widespread smartphone use, platforms like Twitter became crucial for sharing information and gaining support. The news of the Constitutional Court’s decision to imprison former President Jacob Zuma spread quickly online, drawing public sympathy and leading to protests that started in KwaZulu-Natal and spread to almost all provinces. This mobilisation was fuelled by political instability within the ruling party and public frustration over the slow judicial processes related to the State Capture investigation. For many economically disadvantaged citizens, still dealing with inequalities from the apartheid era, Zuma’s arrest represented larger issues. What began as a solidarity march at Zuma’s Nkandla homestead, where supporters tried to prevent his arrest, quickly turned into widespread looting and property damage in KwaZulu-Natal, before moving to Johannesburg and Pretoria over two weeks. The government's efforts to restore order led to 354 deaths and more than 8,263 arrests. It is important to mention that attempts to stop the violence were ineffective until Zuma was released, and tensions increased in Phoenix, where clashes between Indian South Africans and Black residents resulted in about 38 deaths. The South African Human Rights Commission (SAHRC) conducted an inquiry into the July 2021 unrest through a National Hearing Panel, which reviewed 54 oral testimonies and 120 written submissions. The Commission's findings suggest that the unrest was staged by a group of resourceful primary actors who organised secondary participants, many of whom participated in theft and looting at commercial centres. The report noted that the timing of former President Jacob Zuma’s imprisonment and the protests was coincidental, with no direct evidence linking the two events. Although socio-economic hardship was not identified as the primary catalyst for the unrest, the report acknowledged that systemic poverty, inequality, and a widespread lack of trust in the government's capacity to improve living conditions created fertile ground for secondary actors to vent their frustrations through participation in the unrest. Among its recommendations, the report emphasised the need to develop responsive social media monitoring and intervention mechanisms to facilitate rapid responses to emerging national security threats. Conclusion How a government responds to public participation in political actions, such as protests, reflects the strength and maturity of its democratic institutions. Responsive governance supports dissent and enables government officials to receive direct feedback on how policies impact society. The diverse interest in protest dynamics, encompassing fields such as sociology, social psychology, political science, economics, and computer science, emphasises the importance of protest studies in understanding, preventing, and managing civil unrest. These insights are crucial for developing effective strategies to restore and maintain urban safety during periods of social disruption. The analysis of protests related to the imprisonment of former President Jacob Zuma reveals deeper issues of political mismanagement within South Africa’s ruling party. This mismanagement created opportunities for collective mobilisation. By using machine learning techniques, such as sentiment analysis, topic modelling, and data visualisation, the study found that the public generally opposed Zuma’s arrest. The strong presence of negative sentiment in the data shows the lasting legitimacy and support he had among parts of the population, which fuelled protests from KwaZulu-Natal to Pretoria. The quick spread of information through social media changed a peaceful march into widespread looting, violence, and destruction. Bigrams taken from various provincial datasets reveal deep political tensions within the African National Congress (ANC), echoing historical rivalries like those with the Pan Africanist Congress (PAC) during apartheid. These patterns also uncover underlying economic grievances within marginalised and impoverished communities. The study thus highlights the dual role of social media as a tool for mobilisation and a means for spreading unrest, following the path of the protests from KwaZulu-Natal to the urban centres of Johannesburg and Pretoria. Several interconnected factors contributed to the July 2021 unrest in South Africa, stemming from political power struggles, ongoing judicial inquiries, and deep-seated socio-economic issues. The study identifies the structure of the protest's mobilisation and explains its trajectory using collective behaviour theory. The protests initially served as a platform for political expression and civic dissent, but they turned into widespread chaos marked by racial tensions, property destruction, and loss of life. Topic modelling of Twitter discussions revealed economic grievances as key themes, showing that the unrest went beyond loyalty to Jacob Zuma and reflected broader frustrations among marginalised communities. What started as a peaceful mobilization in support of Zuma quickly turned into looting and vandalism across several provinces, continuing for two weeks. The clashes in the Phoenix community highlighted long-standing racial divisions, worsening the unrest. These findings line up with the South African Human Rights Commission’s investigation, which found no direct link between Zuma’s imprisonment and the violence. However, socio-economic inequality and a lack of public trust in government drove participation. Based on these observations, the life cycle of the July 2021 protest can be understood in four stages: protest, looting, destruction, and change. “Protest” includes the initial civic mobilization; “looting” refers to the opportunistic exploitation of the unrest; “destruction” describes the violence and fatalities that followed; and “change” indicates the shifts in public awareness or government response triggered by the collective action. The study shows the benefits of combining machine learning with social science research by using computational methods to analyse large amounts of text and gain insights into complex social issues. Through sentiment analysis and topic modellingg, the research demonstrates how monitoring public perception on social media can provide an essential feedback mechanism for policymakers, businesses, and government officials. These insights can inform policy formation, judicial reviews, and crisis management decisions, ultimately contributing to the creation of safer and more responsive urban environments. The visualizations produced in the study improve understanding by revealing underlying patterns and relationships within the discussions. However, the study has limitations. Its focus on Twitter data from certain South African provinces during the July 2021 unrest may limit the applicability of the findings and broader causal connections. Future research could benefit from incorporating data from multiple social media platforms and expanding the geographical scope to provide a more comprehensive view of protest dynamics. Declarations Conflict of Interest Statement On behalf of all authors, the corresponding author states that there is no conflict of interest. References Aggarwal, C. C., & Zhai, C. (2012). A survey of text clustering algorithms. Mining text data, 77-128. Alexander, P. (2013). Marikana, turning point in South African history. Review of African Political Economy, 40(138), 605-619. Bishop, C. M., & Nasrabadi, N. M. (2006). Pattern recognition and machine learning (Vol. 4, No. 4, p. 738). New York: springer. Bond, P., & Mottiar, S. (2013). Movements, protests and a massacre in South Africa. Journal of Contemporary African Studies, 31(2), 283-302. Cloward, R. A., & Piven, F. F. (1977). The acquiescence of social work. Society, 14(2), 55-63. Coffey, R. (2022). The Sharpeville Massacre, 1960: African Activism and the Press. In The British Press, Public Opinion and the End of Empire in Africa: The'Wind of Change', 1957-60 (pp. 169-212). Cham: Springer International Publishing. Daniel Jurafsky and James H. Martin. 2025. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models, 3rd edition. Online manuscript released January 12, 2025. https://web.stanford.edu/~jurafsky/slp3. Davies, J. C. (1962). Toward a theory of revolution. American sociological review, 5-19. Della Porta, D. (2015). Social movements in times of austerity: Bringing capitalism back into protest analysis. John Wiley & Sons. Desai, A. (2023). Geographies of racial capitalism: the 2021 July riots in South Africa. Ethnic and Racial Studies, 46(16), 3542-3561. Dubow, S. (2015). Were there political alternatives in the wake of the Sharpeville-Langa violence in South Africa, 1960?. The Journal of African History, 56(1), 119-142. Eisinger, P. K. (1973). The conditions of protest behavior in American cities. American political science review, 67(1), 11-28. Feenstra, R. A. (2015). Activist and citizen political repertoire in Spain: A reflection based on civil society theory and different logics of political participation. Journal of civil society, 11(3), 242-258. Fortuna, P., & Nunes, S. (2018). A survey on automatic detection of hate speech in text. Acm Computing Surveys (Csur), 51(4), 1-30. Gamson, W. A. (1991, March). Commitment and agency in social movements. In Sociological forum (Vol. 6, No. 1, pp. 27-50). New York: Kluwer Academic Publishers-Plenum Publishers. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. Book in preparation for MIT Press. URL¡ http://www. deeplearningbook. org, 1. Griffiths, D. (2019). # FeesMustFall and the decolonised university in South Africa: Tensions and opportunities in a globalising world. International Journal of Educational Research, 94, 143-149. Gurr, T. (1968). A causal model of civil strife: a comparative analysis using new indices1. American political science review, 62(4), 1104-1124. Hastie, T., Tibshirani, R., Friedman, J., Hastie, T., Tibshirani, R., & Friedman, J. (2009). Overview of supervised learning. The elements of statistical learning: Data mining, inference, and prediction , 9-41. Hollister, R. W. (2023). The Sharpeville Massacre, Violence, and the Struggles of the African National Congress, 1960-1990. Armstrong Undergraduate Journal of History, 13(1), 62-75. Jurafsky, D., & Martin, J. H. (2009). Speech and Language Processing. Kaldor, M., Selchow, S., & Murray-Leach, T. (Eds.). (2015). Subterranean politics in Europe. Springer. Klandermans, B., & Oegema, D. (1987). Potentials, networks, motivations, and barriers: Steps towards participation in social movements. American sociological review, 519-531. Langa, M. (2017). Researching the# FeesMustFall movement. Hashtag: An analysis of the# FeesMustFall movement at South African universities, 6-12. Liu, B. (2012). Sentiment analysis: A fascinating problem. In Sentiment analysis and opinion mining (pp. 1-8). Cham: Springer International Publishing. Liu, B. (2022). Sentiment analysis and opinion mining. Springer Nature. McAdam, D. (1983). Tactical innovation and the pace of insurgency. American sociological review, 735-754. McCarthy, J. D., & Zald, M. N. (1977). Resource mobilization and social movements: A partial theory. American journal of sociology, 82(6), 1212-1241. Melucci, A., Keane, J., & Mier, P. (1989). Nomads of the present: Social movements and individual needs in contemporary society. Mitchell, T. M. (1997). Does Machine Learning Really Work?. AI Magazine, 18(3), 11-20. Molteno, F. (1979). The uprising of 16th June: A review of the literature on events in South Africa 1976. Social Dynamics, 5(1), 54-89. Montero, J. R., Gunther, R., & Torcal, M. (1997). Democracy in Spain: Legitimacy, discontent, and disaffection. Studies in comparative international development, 32, 124-160. Mooijman, M., Hoover, J., Lin, Y., Ji, H., & Dehghani, M. (2018). Moralization in social networks and the emergence of violence during protests. Nature human behaviour, 2(6), 389-396. Murphy, K. P. (2012). Machine learning: a probabilistic perspective. MIT press. Naidoo, K., Lewis, S., Essop, H., Koch, G. G., Khoza, T. E., Phahlamohlaka, N. M., & Badriparsad, N. R. (2023). July 2021 civil unrest: South African diagnostic radiography students’ experiences. Health SA Gesondheid, 28(1). Ndlovu, S. M. (2006). The soweto uprising. The road to democracy in South Africa, 2, 1970-1980. Opp, K. D., & Gern, C. (1993). Dissident groups, personal networks, and spontaneous cooperation: The East German revolution of 1989. American sociological review, 659-680. Pang, B., & Lee, L. (2008). Opinion mining and sentiment analysis. Foundations and Trends® in information retrieval, 2(1–2), 1-135. Phungula, N. (2024). Understanding the dynamics of South Africa’s July 2021 social unrest. Journal of Nation-Building and Policy Studies , 8 (1), 71. Samuel, A. L. (1959). Some studies in machine learning using the game of checkers. IBM Journal of research and development, 3(3), 210-229. Shalev-Shwartz, S., & Ben-David, S. (2014). Understanding machine learning: From theory to algorithms. Cambridge university press. Snow, D. A., & Benford, R. D. (2005). Clarifying the relationship between framing and ideology. Frames of protest: Social movements and the framing perspective, 205, 209. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction . MIT Press. Tarrow, S. (2022). Power in movement . Cambridge university press. Tilly, C. (2004). Social boundary mechanisms. Philosophy of the social sciences, 34(2), 211-236. Tollefson, J. (2024). Protests over Israel-Hamas war have torn US universities apart: what's next?. Nature. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7760005","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":523403833,"identity":"4c21d16c-dd5d-4884-9c69-4caedb60693e","order_by":0,"name":"Temitope Kekere","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCklEQVRIie2RsUrEQBBA5zi4NBO2nSr5BUPaoL+yIbCprk9xYEBICheu9TMObAVHFlLl/uAKQUgnrNhYCJqEKw5JTkuLfcVus4837AA4HP8UHk+E5XAFAKvzz3FQ+ESJ8S8KnCqp/k258m5TY4tDAGjMxdtDkj+SyixsEghrnq7gfsfcdjH4VZbedWqtSTUEjYKFltMKrXf8VJm0FBgbZNMreQWL0vRzzijh66hcl0K8m0/+ypHyGzsoK/E8U/FHRYKvl1n/eRJJMQ0K0kyl7Qdr2y6qsIkjzVmk206RbBQSTVe8en9vi+IQCsxe6IMvQ69WsbWbJAi305Uj/GMZ8rixs4rD4XA4ZvkG+bFgF/skofoAAAAASUVORK5CYII=","orcid":"","institution":"University of Pretoria","correspondingAuthor":true,"prefix":"","firstName":"Temitope","middleName":"","lastName":"Kekere","suffix":""},{"id":523403834,"identity":"56bc63c6-ee2a-420d-bed0-5aebf28888e9","order_by":1,"name":"Vukosi Marivate","email":"","orcid":"","institution":"University of Pretoria","correspondingAuthor":false,"prefix":"","firstName":"Vukosi","middleName":"","lastName":"Marivate","suffix":""},{"id":523403835,"identity":"ba9f2890-e23a-4a5b-a553-f05cba21fb1f","order_by":2,"name":"Marié Hattingh","email":"","orcid":"","institution":"University of Pretoria","correspondingAuthor":false,"prefix":"","firstName":"Marié","middleName":"","lastName":"Hattingh","suffix":""}],"badges":[],"createdAt":"2025-10-01 13:38:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7760005/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7760005/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":92702222,"identity":"364a5325-3cad-4acf-af61-82d80a512248","added_by":"auto","created_at":"2025-10-03 08:47:50","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":683692,"visible":true,"origin":"","legend":"","description":"","filename":"protestdiffusion.docx","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/96b2527521ca456db1ae8d16.docx"},{"id":92702062,"identity":"801447ea-4d1e-42e8-9891-65039a21b92d","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5775,"visible":true,"origin":"","legend":"","description":"","filename":"2a40e22d730c42cdbd4a76e3e2d67710.json","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/2eaebd4e19f76f70c3cb8e73.json"},{"id":92702058,"identity":"cc44cd69-0a33-4749-b908-da81edd4128c","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":65286,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 19 Histogram of sentiment distribution\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/046d7e8baec4770cf5e33b7a.jpg"},{"id":92702221,"identity":"c872eb23-2215-435a-920b-0e9bc834ad8b","added_by":"auto","created_at":"2025-10-03 08:47:50","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":65467,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 20 Predicted sentiment class distribution\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/1668ec01db606c242e8ce6bd.jpg"},{"id":92702059,"identity":"eacbf651-22a9-4e2d-9c38-4f6f89800019","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":91893,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 21 Top 20 Bigrams for Cape Town dataset\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/f67feee8740cb45d0dc0b3ce.jpg"},{"id":92702061,"identity":"086c0255-39d8-4ba3-885a-4c48ff6934d4","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":91215,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 22 Frequency Distribution of Top 20 Bigrams in Durban Dataset During Civil Unrest\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/9f82e8b395c9c31ad9d783e6.jpg"},{"id":92702063,"identity":"6993c271-d5b4-4065-9bd6-296da5f347c9","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":94039,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 23 Frequency Analysis of Top 20 Bigrams in the #FreeJacobZuma Dataset\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/e8e67b37b91caf04f732f9cb.jpg"},{"id":92702066,"identity":"65744277-c679-40b1-92db-c544c57bcb58","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":93337,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 24 Frequency Analysis of Top 20 Bigrams in the Johannesburg Dataset\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/25db6d1a3a80976e26192354.jpg"},{"id":92702068,"identity":"5b403221-861e-4074-aa39-d8ba84d2b0dc","added_by":"auto","created_at":"2025-10-03 08:39:51","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":103128,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 25 Frequency Analysis of Top 20 Bigrams in the Mahikeng Dataset\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/a05c5f2d1e25d3e4735d8bbc.jpg"},{"id":92702071,"identity":"e86ae045-0582-42e2-a550-e76840737757","added_by":"auto","created_at":"2025-10-03 08:39:51","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":89432,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 26 Frequency Analysis of Top 20 Bigrams in the Mangaung Dataset\u003c/p\u003e","description":"","filename":"8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/d98f9a441bfe2814c68723d5.jpg"},{"id":92702065,"identity":"e5b26ab2-46a1-4ec4-9450-bd3d566e6c4b","added_by":"auto","created_at":"2025-10-03 08:39:50","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":88955,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 27 Frequency Analysis of Top 20 Bigrams in the Phoenix Massacre Dataset\u003c/p\u003e","description":"","filename":"9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/bad7e0fba71347e0edccb77b.jpg"},{"id":92702070,"identity":"ba12551f-5dd5-4a74-a1df-279c9c73ae61","added_by":"auto","created_at":"2025-10-03 08:39:51","extension":"jpg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":86766,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 28 Frequency Analysis of Top 20 Bigrams in the Port Elizabeth Dataset\u003c/p\u003e","description":"","filename":"10.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/eceb50e883c4b3bb3e4526d8.jpg"},{"id":92702223,"identity":"1ea4ff5c-d2f8-44fc-b1d5-f5445bb76146","added_by":"auto","created_at":"2025-10-03 08:47:51","extension":"jpg","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":92894,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 29 Frequency Analysis of Top 20 Bigrams in the Pretoria Dataset\u003c/p\u003e","description":"","filename":"11.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/92fbd4d62af4a9d194c1484d.jpg"},{"id":92702069,"identity":"5720a8fc-c19d-4e6f-b610-5ddb083072b2","added_by":"auto","created_at":"2025-10-03 08:39:51","extension":"jpg","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":88886,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 30 Frequency Analysis of Top 20 Bigrams in the Thohoyandou Dataset\u003c/p\u003e","description":"","filename":"12.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/11e9cc5eb37e46e176274348.jpg"},{"id":92702072,"identity":"4cc5f861-3023-4383-974c-88274db3d0b9","added_by":"auto","created_at":"2025-10-03 08:39:51","extension":"jpg","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":100429,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 31 Frequency Analysis of Top 20 Bigrams in the Umhlathuzi Dataset\u003c/p\u003e","description":"","filename":"13.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/cfb8669999482a1960794a09.jpg"},{"id":94729422,"identity":"75d2aee4-18b6-4ceb-8194-c7c8fc5b7d5c","added_by":"auto","created_at":"2025-10-30 07:04:57","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1888959,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7760005/v1/6054e191-212d-4de1-84aa-db24755c25a8.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Diffusion of protest behaviour: Analysing the July 2021 civil unrest in South Africa through sentiment analysis and topic modelling","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe protest movement that erupted in South Africa in July 2021, following the incarceration of former President Jacob Zuma for contempt of court, represents the latest chapter in the country’s long history of civil unrest (Phungula, 2024; Desai, 2023; Naidoo et al., 2023). South Africa’s protest tradition stretches back to the apartheid era, including the anti-pass law demonstrations that culminated in the Sharpeville-Langa massacre (Dubow, 2015; Coffey, 2022; Hollister, 2023), the 1976 Soweto Uprising against Afrikaans as the language of instruction (Molteno, 1979; Ndlovu, 2006), the Marikana labour strike and massacre in 2012 (Alexander, 2013; Naicker, 2016; Bond \u0026amp; Mottiar, 2013), and the Fees Must Fall student movement that challenged the rising cost of higher education (Langa, 2017; Griffiths, 2019). These protests reflect deeper structural grievances and unresolved conflicts, whether rooted in state repression, racial injustice, labour exploitation, or educational inequality. In this context, protest is not merely a reactionary act but a form of political participation through which citizens articulate resistance and contest state authority within the public sphere.\u003c/p\u003e\n\u003cp\u003eAs earlier sociological theories have suggested, protest is a global phenomenon that is not limited to South Africa or mainly Black communities. Around the world, demonstrations have occurred in various political, economic, and cultural settings. In Spain, for example, activists in Madrid, representing a racially diverse and urban population, occupied Plaza Puerta del Sol on 15 May 2011 to oppose austerity measures imposed by the political elite. Protesters set up temporary camps to demand better welfare systems and less corruption in financial and political institutions. While traditional media provided limited coverage, the movement, mostly driven by youth, used social media to mobilize widely. This made it one of the most significant political mobilizations outside of Spain’s formal labor unions and political groups (della Porta, 2015; Feenstra, 2015; Montero et al., 1997; Kaldor et al., 2015). Similarly, the United States has seen multiple waves of civil unrest throughout its democratic history. Most recently, university students across diverse campuses have protested antisemitism and the Israeli military campaign in Gaza (Tollefson, 2024).\u003c/p\u003e\n\u003cp\u003eStudying protest is crucial since it provides policymakers, researchers, and community members with ways to understand collective action, reveal its motivations, and assess its effects in different socio-political contexts. This inquiry improves our understanding of the complex grievances behind protest movements, such as economic exclusion, political marginalization, and calls for justice and recognition (Tilly, 2004; Gurr, 1968). Instead of seeing protests as spontaneous or isolated occurrences, scholarly research places them in broader historical contexts and structural power relations, which allows for deeper interpretations of their causes and effects (Gamson, 1990). Importantly, empirical research on protest helps develop effective strategies for managing civil unrest. These strategies range from reducing immediate violence to creating long-term plans that tackle systemic inequalities, thus promoting stability in cities and strengthening societal resilience (Davies, 1962; Cloward \u0026amp; Piven, 1977). Scholarly studies of protest movements have used various methods to capture the complex nature of collective action.\u003c/p\u003e\n\u003cp\u003eQuantitative methodologies, particularly surveys and structured questionnaires, enable researchers to measure predictors of protest participation, trace patterns of mobilization, and model the spatial or social diffusion of contentious episodes (Klandermans \u0026amp; Oegema, 1987; Opp \u0026amp; Gern, 1993). These methods support the empirical testing of theories such as relative deprivation, political opportunity structures, and resource mobilisation, facilitating the analysis of protest phenomena at scale (McCarthy \u0026amp; Zald, 1977; McAdam, 1983). Complementing this, qualitative approaches provide the contextual richness necessary to interpret protest through theoretical lenses such as framing, political process, and new social movement theory (Snow \u0026amp; Benford, 2005; Tarrow, 2022; Melucci et al, 1989). These frameworks are applied through case studies, ethnographies, and discourse analyses that investigate activist methods, identity formation, and the interpretive frameworks that support collective mobilization.\u003c/p\u003e\n\u003cp\u003eMachine learning is one of the key areas of artificial intelligence. It offers a strong computational framework for analysing complex datasets and gaining insights into various social issues (Bishop, 2006; Goodfellow et al., 2016). Its main feature is that systems can learn from data by identifying patterns, making predictions, and improving performance on tasks without needing to be explicitly programmed for every situation (Samuel, 1959; Mitchell, 1997). Machine learning methods can be categorised into three primary types: supervised learning, unsupervised learning, and reinforcement learning (Sutton \u0026amp; Barto, 2018; Murphy, 2012). Unsupervised learning finds hidden structures in unlabelled data, usually through clustering. Reinforcement learning teaches agents to make the best sequential decisions based on feedback rewards. In contrast, supervised learning is commonly used in computer science and is increasingly applied in social science research. It allows researchers to learn from labelled input-output pairs to predict or classify future observations, providing a scalable alternative to traditional qualitative and survey methods.\u003c/p\u003e\n\u003cp\u003eIn supervised learning, a machine learning model is trained on a dataset comprising input instances (features) paired with corresponding output labels (target variables) (Hastie et al., 2009). The model, similar to a student, learns the relationship between features and labels by repeatedly processing these input-output pairs. During training, an optimisation algorithm adjusts the model’s internal parameters to reduce a loss function, which measures the difference between the predicted and actual labels (Goodfellow et al., 2016; Shalev-Shwartz \u0026amp; Ben-David, 2014). For example, in sentiment classification, the model connects text features like unigrams, bigrams, or other syntactic patterns with sentiment categories such as ‘positive,’ ‘negative,’ or ‘neutral.’ The model’s effectiveness is evaluated using both the training set and a separate validation set to ensure reliability and prevent overfitting. Once validated, it is tested on new data to check its generalizability (Bishop, 2006). A well-trained model can be utilised for various classification tasks, including spam detection and disease diagnosis. This study categorises social media discussions related to civil unrest.\u003c/p\u003e\n\u003cp\u003eSupervised learning supports numerous important applications in Natural Language Processing (NLP), which aims to enable computers to understand and process human language (Manning \u0026amp; Schütze, 1999; Jurafsky \u0026amp; Martin, 2009). One such task is sentiment analysis, where a model trained on a labelled dataset (e.g., annotated tweets) can determine the emotional tone of new text inputs (Pang \u0026amp; Lee, 2008; Liu, 2012, 2022). Besides sentiment classification, supervised models are commonly used for topic classification (assigning texts to specific thematic areas), named entity recognition (identifying individuals, organizations, or locations), and hate speech detection. These tasks provide detailed insights into social discourse, enabling systematic analysis of public opinion, ideological positions, and behaviour patterns in digital communication (Aggarwal \u0026amp; Zhai, 2012; Fortuna \u0026amp; Nunes, 2018).\u003c/p\u003e\n\u003cp\u003eMooijman et al. (2018) employed machine learning techniques on social media data to investigate the rise of violent protests, specifically the clashes between protesters, police, and counter-protesters during the Baltimore unrest. The research examined how moral beliefs impact individual actions and social media networks in protest situations. It was discovered that how people view the moral legitimacy of a protest, both personally and through social reinforcement, is a crucial factor in driving collective mobilisation. The authors noted that deeply held religious or moral beliefs, when expressed in digital social networks, can increase individuals’ likelihood of engaging in civil disobedience during moments of moral outrage.\u003c/p\u003e\n\u003cp\u003eIn their study of the Baltimore protests, Mooijman et al. (2018) manually labelled 4,800 tweets as either moral or non-moral, creating a training dataset for a neural network that classified an additional 18 million tweets. The aim was to trace the moral arguments present in public discussions around the arrest and subsequent death of Freddie Gray, a 25-year-old Black man who suffered a fatal spinal cord injury while in police custody and was denied medical attention. Although six police officers were charged with serious offences, including second-degree murder and manslaughter, public outrage mounted over the perceived failure of the justice system. The study found that spikes in moral rhetoric were temporally aligned with the escalation of violence and correlated with increased protest-related arrests on an hourly basis. These findings demonstrate the potential of machine learning to operationalise hypotheses in protest research, ranging from measuring public sentiment and participation to uncovering the moral underpinnings of contentious political action. The Baltimore case underscores the complex intersections of racial injustice, police brutality, and institutional distrust, as reflected and amplified through social media discourse.\u003c/p\u003e\n\u003cp\u003eIn the study of protest movements, machine learning enables the processing and analysis of large-scale data, whether in textual, audio, or visual formats. This paper focuses on the protest activity that erupted in July 2021 following the incarceration of former President Jacob Zuma, using Twitter data to examine the evolving discourse and sentiment dynamics. The study contributes to social media analytics and protest research by demonstrating how social media data can be systematically collected and analysed to understand protest behaviour in South Africa. It presents sentiment as a mediating variable in the development and escalation of protest, traces the temporal spread of unrest through sentiment time-series analysis, and reveals provincial differences in participation through the lens of collective behaviour theory. These findings offer practical value to stakeholders, including policymakers and crisis response teams, by providing empirical evidence for anticipating, mitigating, and preventing episodes of civil unrest, thereby supporting safer and more resilient urban governance. The structure of the paper is as follows: Section 2 reviews the literature on theories of collective behaviour; Section 3 outlines the methodology; Section 4 presents the findings; and Section 5 concludes the study with implications and recommendations.\u003c/p\u003e"},{"header":"Literature Review","content":"\u003cp\u003eOne important theoretical framework for understanding protest behaviour is the political opportunity structure (POS) model. This model examines how institutional arrangements and political settings influence collective action. POS theory suggests that the extent to which political systems enable or limit public participation influences the rise and direction of protest movements. According to this perspective, open political systems that allow public demands to be expressed through easy and responsive institutional channels may boost protest activity. They do this by legitimising dissent and making it easier to mobilise. On the other hand, more restrictive or closed systems can also trigger protest. When citizens face institutional barriers, they may turn to non-institutional methods to express their discontent. Some scholars advocate for a hybrid interpretation, suggesting that a mix of openness and constraint can create optimal conditions for protest. Eisinger's (1973) empirical study of American cities supports this perspective, revealing that protest activity was highest in localities characterised by a hybrid political structure, where limited openness coexisted with institutional friction, intensifying citizen mobilisation.\u003c/p\u003e\n\u003cp\u003eEmpirical evidence suggests that aggression frequently arises as a sociological response to deprivation, and that the application of military force is perceived as a form of deprivation, thereby intensifying frustration-aggression dynamics (Gurr, 1968). Within Gurr’s theory, relative deprivation is mediated by various psychological and social variables, including perceived injustice, group identity, political efficacy, and social cohesion, which collectively influence the magnitude of civil strife. While theoretically conceptualised as linear or sublinear influences in Gurr’s model, these mediating variables can be operationalised in machine learning frameworks as independent variables that predict protest behaviour. In such models, the magnitude of strife functions as the dependent variable or output. This computational framing allows researchers to empirically test the explanatory power of these variables using large-scale social data, an approach elaborated upon in the methodological section of this study.\u003c/p\u003e\n\u003cp\u003eBetween 1961 and 1965, Gurr’s seminal study analysed 1,100 strife events across 104 U.S. cities, identified from local newspaper sources. These events were systematically hand-coded, and the resulting variables were subjected to factor analysis. Six indicators of relative deprivation were consolidated into a composite index and statistically correlated with civil strife predictors. Gurr proposed a curvilinear relationship between the independent variable—relative deprivation, operationalised through coercive potential, institutionalisation, facility, and legitimacy—and various forms of civil unrest, including conspiracy, internal war, and turmoil. Notably, coercive force did not follow this curvilinear trend. Conceptually, the schematic model presented by Gurr parallels the architecture of a neural network, where independent variables (relative deprivation indicators) are trained to predict outputs (dimensions of civil strife). The model learns the relationship between these variables, ultimately revealing that more significant relative deprivation corresponds to increased unrest intensity. Importantly, Gurr’s framework excludes revolution from the scope of strife, focusing instead on forms of unrest that emerge within existing political systems.\u003c/p\u003e\n\u003cp\u003eThe basis for protest behaviour is often explained through grievance theory. This framework suggests that collective action typically arises from ongoing, unresolved societal complaints. These complaints often stem from systemic inequalities or perceived unfairness. Issues such as division among elites, problems within institutions, or the exclusionary nature of political systems can exacerbate these grievances.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAccording to the theory, protests occur when these grievances align with openings in the political landscape, such as leadership gaps, conflicts among elites, or a decline in trust in the regime. This situation allows previously marginalised voices to organise. Economic factors, such as high unemployment, poverty, or austerity measures, can exacerbate feelings of dissatisfaction, particularly when government structures fail to address pressing public issues. Therefore, grievance theory highlights how the interaction between structural limits and lack of political openness can trigger mass mobilisation, especially when institutional redress lacks effectiveness.\u003c/p\u003e"},{"header":"Methodology","content":"\u003cp\u003e\u003cem\u003eThis study takes a computational social science approach, using natural language processing (NLP) techniques to analyse digital trace data from social media. The methodology includes several connected parts. It starts with collecting Twitter data on the July 2021 protests in South Africa. After acquiring the data, initial exploratory analyses cleaned and organised it, which helped with sentiment classification and topic extraction. Researchers annotated the dataset to train supervised learning models for sentiment classification. They then applied unsupervised topic modelling algorithms to find the main themes in the discussions. Temporal analysis was followed to track the development and spread of the unrest over time. The insights from these NLP tasks were compared with traditional media coverage and official reports to ensure they were accurate and added depth to the interpretation. This mixed-method strategy offers a solid framework for examining the digital aspects of protest movements and understanding how online discussions relate to offline events.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cstrong\u003eData Collection\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eDuring the early days of the July 2021 unrest, this study collected Twitter data using a specific set of hashtags related to the protests. Hashtags like #JulyUnrest, #Looting, #FreeJacobZuma, #LootingSA, #ShutdownSA, and #PhoenixMassacre were used to gather relevant tweets from Twitter. The data collection method took geography into account, allowing us to categorise tweets by province. We sourced tweets from all ten provinces, with a strong focus on KwaZulu-Natal (Durban), where the protests began, and Gauteng (Johannesburg and Pretoria), which later saw significant unrest. We also collected additional data from other provinces, including North West (Mahikeng), Western Cape (Cape Town), Eastern Cape (Port Elizabeth), Limpopo (Thohoyandou), Free State (Mangaung), and smaller municipalities in KwaZulu-Natal, such as Umhlathuze. A distinct dataset labelled \u0026quot;FreeJacobZuma\u0026quot; was created due to the high frequency of tweets containing that hashtag. The study generated 11 provincial and thematic subsets of data, covering the protest activity from 8 July to 22 July 2021. Table 6 summarises the dataset names, tweet volumes, and relative proportions across subsets.\u003c/p\u003e\n\u003cp id=\"_Toc200940945\"\u003eTable 6 Name of datasets and their proportion\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDatasets\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNumber of samples\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePercentages (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCape town\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1427\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1.97\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eFreeJacobZuma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e24747\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e34.21\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eJohannesburg\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e5448\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e7.53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMahikeng\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e77\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.11\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMangaung\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e159\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.22\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePhoenixMassacre\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e36333\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e50.23\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePort Elizabeth\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e285\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.39\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePretoria\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1952\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e2.70\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eThoyandou\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e41\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eUmhlathuzi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.11\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTotal\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e72332\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e100\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eData Preprocessing\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eExtensive preprocessing was necessary to ensure data quality and analytical relevance before the harvested tweets could be used for sentiment classification and topic modelling. In adherence to Twitter\u0026rsquo;s Developer Policy and to protect user privacy, all tweets were collected without personally identifiable information. The raw tweets often contained unstructured and noisy elements such as numerals, abbreviations, internet slang, typographical errors, special characters, hyperlinks, hashtags, and other non-informative tokens. These components were systematically removed to improve textual coherence.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFollowing this cleaning process, the tweets were tokenised, breaking each tweet into discrete word units. Then, they were lemmatised using the Natural Language Toolkit (NLTK) library to standardise the words to their base forms. This normalisation process ensured consistency in word representation across the corpus. The clean textual data was then vectorised using the Term Frequency-Inverse Document Frequency (TF-IDF) method, which transforms the corpus into a numerical format, emphasising informative terms while down-weighting frequent but uninformative words. The resulting matrix was used as input for subsequent machine learning tasks, such as sentiment classification and topic extraction.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eData Labelling\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSupervised machine learning needs a dataset of input-output pairs, where each input (in this case, a tweet) has a corresponding output label that defines sentiment polarity. In this study, each tweet received a single sentiment label, which can either be positive, negative, or neutral, with no overlap between categories. This means a tweet could only belong to one of the three sentiment classes. This classification setup allowed the training of a machine learning model to learn the relationship between text inputs and their corresponding sentiment outputs. The model updated its internal parameters by minimising prediction errors across the training samples.\u003c/p\u003e\n\u003cp\u003eOnce trained, the model could generalize and classify new tweets into the right sentiment class with fewer errors. To support this supervised training process, about 5,000 tweets (6.91% of the 72,332 tweets collected) were manually labelled by human raters, creating a high-quality dataset. This labelled subset was the basis for training and evaluating the sentiment classification models used in the study.\u003c/p\u003e\n\u003cp\u003eThree independent annotators manually reviewed the tweet dataset to ensure reliable sentiment labelling and assigned each tweet a positive, negative, or neutral sentiment label. Annotators worked independently to minimise potential bias or influence from others, thereby helping to maintain objectivity in the labelling process. A detailed annotation guide was provided to standardise the labelling procedure and help annotators deal with ambiguous or complex tweets. The guide also outlined a method for handling cases where annotators disagreed.\u003c/p\u003e\n\u003cp\u003eWhen all three annotators disagreed on a tweet\u0026apos;s sentiment, a fourth party was involved to resolve the disagreement. This adjudicator reviewed the tweet and the labels given, making a final decision. The final output included a labelled dataset with the original tweets, individual annotations, and a consolidated master sheet that showed the adjudicated sentiment label. This process ensured consistency and validity in the labelling, which is vital for training effective supervised machine learning models.\u003c/p\u003e\n\u003cp\u003eTo evaluate the reliability of the sentiment labels created by the annotators, the study calculated inter-annotator agreement using Fleiss\u0026rsquo; Kappa statistic. Previously, tweets with missing, incomplete, or misspelt labels, as well as those that did not meet the annotation guidelines, were removed from the evaluation. This filtering process dropped the total number of usable labelled tweets from 4,916 to 4,847. The Kappa score from these 4,847 tweets was 0.2737, indicating a fair level of agreement among the three annotators.\u003c/p\u003e\n\u003cp\u003eAlthough the score indicates a reasonable level of consistency in the annotators\u0026apos; judgments, it also highlights significant variation in their interpretations. To address these variations, the final label for each tweet used in training the sentiment analysis model was decided by a majority vote among the three annotators. This method balances personal judgment with systematic aggregation, ensuring the labelled dataset remains solid for supervised learning.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eData Modelling\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFollowing the establishment of fair inter-annotator agreement, the study used a majority voting system to finalise the sentiment labels for each tweet. For each sample, the selected label was the one agreed upon by at least two annotators. This method effectively draws on the judgment of the annotators, reducing individual bias and improving the overall reliability of the labelling process. When annotators disagreed and no clear majority emerged, an adjudicator reviewed the tweet and assigned the final label. This process ensured that all 4,916 samples could be retained for training, including those with initially disputed labels. The finalised dataset, comprised of tweet-input and sentiment-label pairs, was used to train a supervised machine learning model to classify unseen tweets as positive, negative, or neutral. Prior to training, the distribution of sentiment classes was assessed to verify class balance, thereby avoiding model bias toward majority classes. Table 7 presents the frequency distribution of sentiment labels in the dataset.\u003c/p\u003e\n\u003cp id=\"_Toc200940946\"\u003eTable 7 Sentiment label distribution\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eSentiment class\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eSentiment samples\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eSentiment proportion %\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNegative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e3063\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e62.31\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePositive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1099\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22.36\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNeutral\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e752\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15.30\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTotal samples\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e4916\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e99.97\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eSentiment distribution of class labels\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eFrom the table above, the sentiment distribution of the annotated dataset reveals a significant class imbalance. Table 7 shows that of the 4,916 tweets, 3,063 (62.31%) were classified as expressing negative sentiment. In contrast, only 1,099 tweets (22.36%) were labelled positive, and 752 (15.30%) were neutral. This skewed distribution suggests that negative sentiment predominates the dataset, posing a challenge for training a machine learning model. When one class significantly outweighs others, models tend to be biased toward the majority class, compromising the accuracy and fairness of the classification task. To address this issue, the study employed the Synthetic Minority Over-sampling Technique (SMOTE), which generates synthetic samples for the minority classes. By rebalancing the dataset, SMOTE ensures that the model does not overlearn from the overrepresented negative class. This improves the generalisation and predictive accuracy of the sentiment classifier.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo create the sentiment classification model, five popular supervised learning algorithms were chosen. The models chosen are Logistic Regression (LR), Gaussian Na\u0026iuml;ve Bayes (GaussianNB), K-Nearest Neighbours (KNN), Support Vector Machines (SVM), and Extreme Gradient Boosting (XGBoost). These models were selected for their effectiveness in multi-class text classification tasks. The preprocessed and SMOTE-rebalanced dataset was divided into training and testing subsets. The training data allowed the models to learn the connection between tweets (input features) and their corresponding sentiment classes (targets). As a multi-class classification problem, the model predicted one of three mutually exclusive sentiment labels for each tweet: positive, negative, or neutral. This experimental setup enabled an empirical comparison of classifier performance, evaluated using the metrics described in the following section.\u003c/p\u003e\n\u003cp id=\"_Toc200940947\"\u003eTable 8 Classifier performance of Sentiment models\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eModel performance\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(F1-score, macro)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBase models\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eLogistic regression\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGaussainNB\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eKNN\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.46\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.48\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eSVM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.77\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eTable 8 presents the classification performance of the machine learning algorithms used to develop the sentiment model. Logistic regression, Gaussian Naive Bayes, and support vector machines achieved approximately 80% accuracy in classifying tweets, demonstrating strong predictive performance. XGBoost followed closely with a slightly lower accuracy of 76%. By contrast, the K-nearest neighbours (KNN) algorithm exhibited the weakest performance, with a notably high error rate and an accuracy of only 54%. These performance metrics were validated using five-fold cross-validation, a robust technique that partitions the dataset into five equal subsets. In each iteration, four subsets are used for training while the remaining subset is used for testing. This process is repeated five times, ensuring each subset is used once as a test set. The average F1-score across all folds provides a reasonable estimate of the model\u0026apos;s generalizability and helps mitigate overfitting.\u003c/p\u003e\n\u003cp\u003eThis study examined a range of established topic modelling algorithms to identify hidden themes in the tweet collection, where each tweet served as a document. Commonly used models in this area include Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorisation (NMF), and Latent Semantic Analysis (LSA). The study compared these models to see how well they could extract clear and understandable topics related to protest discussions on Twitter. However, it presented detailed results only from the most effective model. The topics revealed by this model highlighted the concerns and key issues of the protesters, shedding light on the reasons behind collective action. These findings were viewed through the lens of collective behaviour, allowing for a better understanding of how social unrest develops in response to shared concerns.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eThis section presents the study\u0026apos;s results. The analysis begins with an exploratory analysis of the dataset used to study the protest movement. It outlines how well the sentiment classification model performed in assessing public sentiment, with a focus on perceptions and the emotional tone of discussions related to protests. Subsequently, the section outlines the application of topic modelling to extract dominant themes, offering insight into the protesters\u0026apos; key grievances and focal points.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAn initial dataset exploration offers preliminary insights into the July 2021 protest movement as observed on Twitter. Descriptive statistics summarise and quantify the characteristics of the data collected, highlighting structural and contextual patterns within the corpus. Common statistical descriptors such as the total number of tweets, average tweet length, standard deviation, range, interquartile range, and sentiment label distribution are typically employed to characterise the dataset. Additional descriptive features such as the frequency of hashtag usage, retweet counts, and the number of likes may enrich the analysis. However, these indicators were sparsely distributed in this dataset, and the number of unique users was omitted in compliance with Twitter\u0026rsquo;s data privacy guidelines.\u003c/p\u003e\n\u003cp id=\"_Toc200940948\"\u003eTable 9 Geographical distribution of dataset\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cbr\u003eDatasets\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNumber of samples\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePercentages (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eProvince\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCape town\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1427\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eWestern cape\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDurban\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1787\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e2.47\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eKwaZulu-Natal\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eFreeJacobZuma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e24747\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e34.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNot applicable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eJohannesburg\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e5448\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e7.53\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGauteng\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMahikeng\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e77\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.11\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNorth West\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMangaung\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e159\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eFree State\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePhoenixMassacre\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e36333\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e50.23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eKwaZulu-Natal\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePort Elizabeth\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e285\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eEastern Cape\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePretoria\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1952\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e2.70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGauteng\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eThohoyandou\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e41\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eLimpopo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eUmhlathuzi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e0.11\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eKwaZulu-Natal\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTotal\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e72332\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e100\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eTable 9 shows the geographical distribution of tweets across different South African provinces and cities. Conversations referencing Phoenix constituted the largest dataset share, accounting for approximately 50.23% of all tweets collected. This significant volume is attributed to the violence and fatalities associated with the Phoenix massacre during the unrest. The FreeJacobZuma dataset was followed by 34.21%, reflecting the widespread public support and online mobilisation advocating for the release of former President Jacob Zuma. Johannesburg contributed 7.53% of the total conversation, while Pretoria and Cape Town accounted for 2.70% and 1.97%, respectively. Smaller percentages of tweets originated from other cities that were less directly involved in the protest action. This geographic variation in tweet volume offers insight into regional engagement and participation in the protest discourse, suggesting a possible correlation between online attention and on-the-ground mobilisation.\u003c/p\u003e\n\u003cp id=\"_Toc200940949\"\u003eTable 10 Twitter dataset descriptive statistics\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd colspan=\"4\" valign=\"top\"\u003e\n \u003cp\u003eMean\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"4\" valign=\"top\"\u003e\n \u003cp\u003eStandard deviation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eWords\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eCharacters\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eWords\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eCharacters\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDatasets\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBefore cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAfter cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBefore cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAfter cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBefore cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAfter cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBefore cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAfter cleaning\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCape town\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e163\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e134\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e89\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e82\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDurban\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e152\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e120\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e83\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e87\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eFreeJacobZuma\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e131\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e83\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e73\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eJohannesburg\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e147\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e117\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e78\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMahikeng\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e159\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e128\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e77\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e69\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMangaung\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e147\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e116\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e88\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e81\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePhoenixMassacre\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e154\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e102\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e87\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e79\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePort Elizabeth\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e140\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e108\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e71\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePretoria\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e143\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e113\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e76\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eThohoyandou\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e161\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e136\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e78\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eUmhlathuzi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e16\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e117\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e74\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eTable 10 presents summary statistics that describe the key characteristics of the Twitter dataset, specifically the average and standard deviation of word and character counts before and after data cleaning. These metrics offer insight into the expressiveness and variability of tweets across geographic locations. For instance, tweets from Thohoyandou and Cape Town exhibit the highest average word counts, such as 27 and 26 words per tweet, respectively, suggesting a relatively more expressive discourse in those regions. Mahikeng follows with an average of 25 words, while tweets from Mangaung and Durban average 24 words. Other provinces generally have an average of 22 to 23 words. This reflects the overall average across all regions.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eInterestingly, the FreeJacobZuma dataset shows a clear difference, with an average of 17 words per tweet. This shorter length suggests that expressions of frustration or calls to action might be shorter, possibly indicating stronger emotions. The average character counts in all subsets stay within Twitter\u0026rsquo;s 280-character limit. This shows that users can usually express their thoughts within the platform\u0026rsquo;s limits. The standard deviation for word counts is between 13 and 15, while character counts range from 76 to 89. This shows a moderate variation in tweet lengths across the dataset. These patterns help us understand the structure and dynamics of protest-related discussions.\u003c/p\u003e\n\u003cp\u003eThe statistics for word and character counts after cleaning show little change from the values before cleaning. This suggests that the text preprocessing, which included removing stop words, punctuation, and other non-essential tokens, did not significantly alter the main content of the tweets. The retained linguistic structure suggests that the preprocessing preserved the semantic integrity of the dataset, ensuring that the machine learning models were trained on meaningful and representative textual information. This outcome affirms the efficacy of the cleaning procedure, demonstrating that it effectively filtered noise without discarding valuable content necessary for sentiment analysis and topic modelling tasks.\u003c/p\u003e\n\u003cp id=\"_Toc200940950\"\u003eTable 11 Sentiment distribution of tweets\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNumber of samples\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eProportion (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNegative labels\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e3063\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e62.33\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePositive labels\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e1099\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22.36\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNeutral labels\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e752\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e15.30\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTotal\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e4914\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e100\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eTable 11 presents the sentiment distribution of the labelled subset of tweets, representing 6.9% (5,000 out of 72,332) of the total dataset. Within this annotated sample, negative sentiment accounted for 62.33% (3,063 tweets), positive sentiment constituted 22.36% (1,099 tweets), and neutral sentiment comprised 15.3% (752 tweets). Figure 15 displays a histogram visualising this distribution, where red bars indicate negative sentiments, green bars represent positive sentiments, and blue bars denote neutral sentiments. The histogram underscores the dataset\u0026apos;s significant class imbalance: for every one positive tweet, there are approximately three negative tweets, and for every neutral tweet, there are roughly eight negative tweets. This imbalance poses potential challenges for sentiment classification tasks and necessitates rebalancing techniques to ensure fair model training.\u003c/p\u003e\n\u003cp\u003eTo train the sentiment analysis model, five widely recognised classifiers were employed: Logistic Regression, Gaussian Naive Bayes (GaussianNB), K-Nearest Neighbours (KNN), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost). These models were selected for their established effectiveness in multi-class classification tasks. Table 12 presents the performance metrics of each classifier after being trained on the labelled dataset. The models were evaluated on their ability to accurately classify tweets into three sentiment categories: positive, negative, and neutral.\u003c/p\u003e\n\u003cp id=\"_Toc200940951\"\u003eTable 12 Performance of the sentiment classification model\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"326\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 172px;\"\u003e\n \u003cp\u003eModel performance\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(F1-score, macro)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eBase models\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eLogistic regression\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003e0.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.79\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eGaussainNB\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.85\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eKNN\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003e0.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.46\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eDecisionTree\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003e0.67\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.71\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eSVM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003e0.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.87\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 154px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 99px;\"\u003e\n \u003cp\u003e0.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 73px;\"\u003e\n \u003cp\u003e0.78\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eTable 12 presents the performance metrics of the sentiment classification models. Logistic Regression, Gaussian Naive Bayes, and Support Vector Machines each achieved approximately 80% classification accuracy, demonstrating their effectiveness in sentiment prediction. XGBoost performed slightly lower, with an accuracy of 76%. In contrast, K-Nearest Neighbours (KNN) exhibited the weakest performance, with an accuracy of 54%, indicating a high error rate and limited generalisability for this task. The results reflect F1-scores obtained through five-fold cross-validation, a technique that mitigates overfitting by dividing the dataset into five equal parts. In each iteration, four folds are used for training while the remaining fold is reserved for validation. This process is repeated across all five folds, ensuring a robust estimate of model performance across different data partitions.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe sentiment analysis model developed using the best-performing classifier was employed to label the remaining 93.9% (67,915 out of 72,332) of previously unlabelled tweets. Following the preprocessing steps, including cleaning and deduplication, the unlabelled dataset was reduced to 67,326 tweets. Table 13 presents the sentiment prediction results generated by the selected model, specifically the Support Vector Machine (SVM) classifier. The SVM model classified a substantial portion of the tweets as expressing negative sentiment, reinforcing the dominant tone observed in the annotated subset. Figure 16 offers a histogram visualisation of the SVM-based sentiment distribution, clearly representing the model\u0026rsquo;s output across the negative, positive, and neutral categories.\u003c/p\u003e\n\u003cp id=\"_Toc200940952\"\u003eTable 13 Distribution of Sentiment Labels in SVM Model Predictions on Unlabelled Tweets\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003eNumber of samples\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003eProportion (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003eNegative labels\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e56070\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e83.28\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003ePositive labels\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e1449\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e2.15\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003eNeutral labels\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e9807\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e14.57\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003eTotal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e67326\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 129px;\"\u003e\n \u003cp\u003e100\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eAs illustrated in Table 13 and Figure 16, the SVM model predicted a dominant proportion of tweets, 83.28% (56,070 out of 67,326), as negative. This aligns with the earlier annotation in Table 11, where 60.33% of the manually labelled tweets were also negative. The consistency across labelled and predicted datasets, reinforced using SMOTE for class rebalancing and five-fold cross-validation to minimise bias, suggests a strong underlying narrative of discontent within the data. It is plausible that the prevalence of negative sentiment stems largely from the Phoenix subset, which constituted 50.23% of the entire dataset (as shown in Table 11). This region witnessed a significant escalation of the July 2021 unrest, marked by racially charged violence between Black South Africans and Indian residents. Thus, the overwhelmingly negative sentiment captured by the SVM model reflects the upheaval, looting, and civil tension that dominated Twitter discourse during the period.\u003c/p\u003e\n\u003cp\u003eWithout prejudice to the predominantly negative sentiment identified by the classifier, the study further explored the dataset\u0026apos;s content by examining the most frequently occurring words. A bigram histogram was generated to visualise the top word pairs, offering insight into the dominant themes and summarising the discourse captured in the dataset. The unlabelled subsets underwent a comprehensive cleaning process to ensure the analysis provided meaningful and interpretable results. This included the removal of URLs, HTML tags, emojis, emoticons, usernames, hashtags, special characters, punctuation marks, and stop words. The resulting cleaned corpus served as the basis for identifying salient terms that reveal the underlying concerns and topics discussed during the protest.\u003c/p\u003e\n\u003cp\u003eThe bigram histogram presented in Figure 17 visualises the top 20 word pairs from the Cape Town dataset. Notably, the name \u0026quot;Jacob Zuma\u0026quot; appeared more than 50 times, reflecting the centrality of his jailing to the protest discourse. These words show that former President Jacob Zuma is at the centre of the discussion surrounding the protest movement. Variations of the name also occurred multiple times with differing frequencies, indicating diverse contexts of mention. Other frequently occurring bigrams include \u0026ldquo;South Africa,\u0026rdquo; \u0026ldquo;taxi violence,\u0026rdquo; and \u0026ldquo;Cape Town,\u0026rdquo; each appearing more than 40 times, suggesting these were significant components of the localised protest narrative. Less frequent but still relevant bigrams, each appearing fewer than 20 times, include \u0026ldquo;inciting violence,\u0026rdquo; \u0026ldquo;state capture,\u0026rdquo; \u0026ldquo;Cyril Ramaphosa,\u0026rdquo; and \u0026ldquo;people loot,\u0026rdquo; pointing to broader socio-political concerns linked to governance, leadership, and public frustration. These words show that former President Jacob Zuma is at the centre of the discussion surrounding the protest movement.\u003c/p\u003e\n\u003cp\u003eThe bigram histogram from Durban, the epicentre of the July 2021 protest, presented in Figure 18, identifies \u0026ldquo;Jacob Zuma\u0026rdquo; as the most frequently occurring term, appearing over 60 times in the dataset. Other notable phrases, such as \u0026ldquo;South Africa,\u0026rdquo; \u0026ldquo;Black people,\u0026rdquo; and \u0026ldquo;people looting,\u0026rdquo; appear with lower frequency counts, each under 50. The name \u0026ldquo;Bheki Cele,\u0026rdquo; South Africa\u0026rsquo;s Minister of Police, surfaces in connection with widespread criticism of the security cluster\u0026rsquo;s inability to control the unrest, as reports indicated the police force was overwhelmed by looters. Specific locations, such as KwaMashu Shopping Centre, were mentioned in relation to the extensive damage, including looting and arson, as confirmed by local media sources. Additionally, \u0026ldquo;Sihle Zikalala,\u0026rdquo; the Premier of KwaZulu-Natal and a senior member of the ruling African National Congress (ANC), was frequently cited in public discourse surrounding the province\u0026rsquo;s response to the unrest.\u003c/p\u003e\n\u003cp\u003eFigure 19 presents a bar chart based on the Free Jacob Zuma subset of the dataset. The terms \u0026ldquo;President Zuma\u0026rdquo; and \u0026ldquo;South Africa\u0026rdquo; appear exceptionally frequently, each occurring over 600 times, highlighting the centrality of Zuma\u0026rsquo;s imprisonment to the protest discourse. The widespread use of the hashtag #FreeJacobZuma underscores the public\u0026rsquo;s demand for his release and reflects his symbolic role in mobilising collective action during the unrest.\u003c/p\u003e\n\u003cp\u003eAs the protest spread from Durban, it began to diffuse into other urban centres within Gauteng Province. Figure 19 displays the most frequently occurring terms in the Johannesburg dataset, where \u0026ldquo;South Africa\u0026rdquo; and \u0026ldquo;Jacob Zuma\u0026rdquo; emerged as the most prominent, each appearing over 100 times. Other notable expressions, occurring at roughly half the frequency, included \u0026ldquo;inciting violence,\u0026rdquo; \u0026ldquo;people looting,\u0026rdquo; \u0026ldquo;black people,\u0026rdquo; \u0026ldquo;South Africans,\u0026rdquo; and \u0026ldquo;free Zuma.\u0026rdquo; Additional terms with moderate frequency\u0026mdash;under 50 mentions\u0026mdash;were \u0026ldquo;President Zuma,\u0026rdquo; \u0026ldquo;security cluster,\u0026rdquo; \u0026ldquo;Cyril Ramaphosa,\u0026rdquo; and \u0026ldquo;stop looting,\u0026rdquo; reflecting both political and public safety concerns circulating during the unrest.\u003c/p\u003e\n\u003cp\u003eFigure 21 presents a histogram of the most frequently occurring bigrams in the Mahikeng dataset. The phrases \u0026ldquo;stolen goods,\u0026rdquo; \u0026ldquo;South Africa,\u0026rdquo; \u0026ldquo;South African,\u0026rdquo; and \u0026ldquo;President Zuma\u0026rdquo; appeared three times, indicating their centrality to the discourse. Other notable bigrams appearing twice include \u0026ldquo;country right,\u0026rdquo; \u0026ldquo;Zuma president,\u0026rdquo; \u0026ldquo;I\u0026rsquo;m happy,\u0026rdquo; and \u0026ldquo;stop looting.\u0026rdquo; The phrase \u0026ldquo;people hungry\u0026rdquo; reveals an underlying concern about the province\u0026apos;s food insecurity and economic deprivation. Similarly, \u0026ldquo;let protect\u0026rdquo; suggests calls for safeguarding public infrastructure, while \u0026ldquo;anger towards\u0026rdquo; signals a directed frustration, though the object of that frustration varies across the dataset.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFigure 22 presents the histogram of the most frequently occurring phrases in the Mangaung dataset. \u0026ldquo;Jacob Zuma\u0026rdquo; emerged as the most frequently used phrase, appearing six times, underscoring the centrality of his imprisonment to the discourse. The phrase \u0026ldquo;looting won\u0026rsquo;t\u0026rdquo; reflects local disapproval of looting activities, suggesting a perceived disconnect between such actions and their influence on judicial outcomes. Other frequently appearing phrases, such as \u0026ldquo;won\u0026rsquo;t impact,\u0026rdquo; \u0026ldquo;impact SA,\u0026rdquo; \u0026ldquo;SA economy,\u0026rdquo; \u0026ldquo;Free State,\u0026rdquo; \u0026ldquo;people looting,\u0026rdquo; \u0026ldquo;South Africas,\u0026rdquo; and \u0026ldquo;black people\u0026rdquo;, each appeared more than three times. These expressions highlight both geographical references and socio-political concerns. Less frequent phrases, occurring approximately twice, add further nuance to the conversation, indicating mixed public sentiment around the protest, its perceived consequences, and broader national identity.\u003c/p\u003e\n\u003cp\u003eFigure 23 presents the frequency distribution of the most used phrases in the Phoenix Massacre dataset. The phrase \u0026ldquo;Black people\u0026rdquo; appeared over 3,000 times, indicating the racial dimension central to the discourse. \u0026ldquo;South Africa\u0026rdquo; was mentioned approximately 1,000 times, reflecting the national significance of the event. Phrases such as \u0026ldquo;people killed,\u0026rdquo; \u0026ldquo;Bheki Cele,\u0026rdquo; and \u0026ldquo;Cyril Ramaphosa\u0026rdquo; appeared around 500 times each, highlighting the public\u0026rsquo;s focus on the loss of life and the role of state leadership in managing the crisis. Additional frequently occurring terms, as visualised in the histogram, further underscore the intensity and specificity of public engagement surrounding the events in Phoenix.\u003c/p\u003e\n\u003cp\u003eFigure 24 displays the histogram of the top 20 most frequently occurring phrases in the Port Elizabeth dataset. The overall low frequency of terms suggests minimal engagement from the Port Elizabeth community in the July 2021 protest. The most prominent phrase, \u0026ldquo;Jacob Zuma,\u0026rdquo; appeared 15 times, reflecting the centrality of the former president in the discourse. \u0026ldquo;South Africans\u0026rdquo; and \u0026ldquo;let burn\u0026rdquo; appeared five times each, while other expressions such as \u0026ldquo;South Africa,\u0026rdquo; \u0026ldquo;Free Zuma,\u0026rdquo; \u0026ldquo;former president,\u0026rdquo; and \u0026ldquo;President Jacob\u0026rdquo; occurred with varying lower frequencies. Notably, the histogram also captures discussions around \u0026ldquo;taxi violence,\u0026rdquo; references to \u0026ldquo;Zuma jail,\u0026rdquo; and the provocative phrase \u0026ldquo;must loot,\u0026rdquo; indicating the presence of isolated incitements despite the region\u0026rsquo;s overall limited participation.\u003c/p\u003e\n\u003cp\u003eFigure 25 presents the frequency distribution of the top 20 bigrams in the Pretoria dataset. Only two phrases, \u0026ldquo;Jacob Zuma\u0026rdquo; and \u0026ldquo;South Africa\u0026rdquo;, appeared more than 40 times, underscoring the central role of Zuma\u0026rsquo;s imprisonment in triggering the unrest and the national context in which it unfolded. Phrases occurring more than 20 times include \u0026ldquo;people looting,\u0026rdquo; \u0026ldquo;President Zuma,\u0026rdquo; \u0026ldquo;former President,\u0026rdquo; \u0026ldquo;South Africans,\u0026rdquo; \u0026ldquo;inciting violence,\u0026rdquo; \u0026ldquo;black people,\u0026rdquo; \u0026ldquo;Zuma must,\u0026rdquo; and \u0026ldquo;Zuma arrested.\u0026rdquo; Additional expressions with lower frequency, such as \u0026ldquo;rule law,\u0026rdquo; \u0026ldquo;looting shops,\u0026rdquo; \u0026ldquo;people hungry,\u0026rdquo; \u0026ldquo;poor people,\u0026rdquo; \u0026ldquo;go loot,\u0026rdquo; and \u0026ldquo;looting burning,\u0026rdquo; further highlight key themes in the discourse. Notably, phrases like \u0026ldquo;people hungry\u0026rdquo; and \u0026ldquo;poor people\u0026rdquo; suggest that economic hardship may have contributed to looting participation, while \u0026ldquo;looting burning\u0026rdquo; reflects the intensity of unrest through the destruction of property.\u003c/p\u003e\n\u003cp\u003eFigure 26 presents the variation in frequently occurring bigrams within the Thohoyandou dataset. The most prominent phrase is \u0026ldquo;mislead JZ,\u0026rdquo; which appears more than six times, indicating a potential narrative that accuses individuals or groups of misleading former President Jacob Zuma. \u0026ldquo;Former President\u0026rdquo; occurred approximately three times, reaffirming the centrality of Zuma in the discussions. Phrases such as \u0026ldquo;hong hong\u0026rdquo; were likely a local language expression, and \u0026ldquo;rights limited\u0026rdquo; reflects localised discourse and concerns about civil liberties. Mentions of \u0026ldquo;Edward Zuma\u0026rdquo; may carry sarcastic undertones aimed at the former president\u0026rsquo;s family. Expressions like \u0026ldquo;dilemma poor,\u0026rdquo; \u0026ldquo;poor people,\u0026rdquo; \u0026ldquo;people know,\u0026rdquo; \u0026ldquo;know feels,\u0026rdquo; \u0026ldquo;feels deprived,\u0026rdquo; and \u0026ldquo;deprived livelihood\u0026rdquo; articulate the socio-economic challenges facing the population, though some terms appeared misspelt.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAdditional phrases such as \u0026ldquo;livelihood jobs,\u0026rdquo; \u0026ldquo;job retails,\u0026rdquo; and \u0026ldquo;retail loss\u0026rdquo; point to economic consequences, including employment disruption and business losses during the protest. Finally, phrases like \u0026ldquo;looting men,\u0026rdquo; \u0026ldquo;men black,\u0026rdquo; and \u0026ldquo;black majority,\u0026rdquo; despite inconsistencies in spelling, suggest that discussions in the dataset emphasised the participation of Black men in the looting incidents, reflecting community-level observations during the July 2021 unrest.\u003c/p\u003e\n\u003cp\u003eFigure 27 illustrates the low-frequency distribution of phrases in the Umhlathuzi dataset, indicating minimal discourse related to the July 2021 unrest. Most phrases appear only once or twice, suggesting limited public engagement on the topic in this region. Among the few recurring phrases are \u0026ldquo;South Africa\u0026rdquo; and \u0026ldquo;President Zuma,\u0026rdquo; each appearing twice, which may indicate a peripheral awareness of the national events. Phrases such as \u0026ldquo;white sovereignty\u0026rdquo; and \u0026ldquo;white acts\u0026rdquo; introduce a racial undertone, although their low frequency renders them statistically insignificant in this dataset. Additionally, the phrase \u0026ldquo;lorch donda\u0026rdquo; serves as a casual expression, whereas several other terms are in local languages or do not align with the broader protest narrative. The sparse and disconnected nature of the discussions suggests that the unrest did not greatly impact Umhlathuzi or that the protest movement did not spread into this province or area.\u003c/p\u003e\n\u003cp\u003eThe limited presence of protest-related conversations in Umhlathuzi shows how topic modelling can serve as a tool for real-time monitoring. This can help identify civil unrest early and prevent it. State actors can identify emerging hotspots of discontent by monitoring the frequency with which specific keywords and phrases co-occur. These keywords might include names of political figures, mentions of violence, or social and economic complaints. This proactive method enables timely actions, such as effective communication, community involvement, or deploying resources to reduce unrest. For instance, the early presence of terms like \u0026ldquo;President Zuma\u0026rdquo; or \u0026ldquo;white sovereignty\u0026rdquo; in otherwise unaffected regions could be precursors to ideological mobilisation. Therefore, integrating topic modelling into digital surveillance frameworks provides a scalable and non-intrusive mechanism for understanding the geographical and thematic evolution of protest movements, thereby enhancing national crisis preparedness and response strategies.\u003c/p\u003e\n\u003cp\u003eThis study adopts a semi-supervised learning approach to extend sentiment annotation beyond the gold-standard human-labelled subset. After training and evaluating the models, five widely used classifiers were selected. These selected classifiers were Logistic Regression, Gaussian Naive Bayes, K-Nearest Neighbours (KNN), Support Vector Machines (SVM), and XGBoost. A comparative performance analysis was conducted using accuracy and F1-scores derived from five-fold cross-validation. The results revealed SVM as the most effective model for multi-class sentiment classification in this context, exhibiting the highest generalisation performance on training and unseen data. Leveraging this outcome, the trained SVM model was used to classify the remaining unlabelled 93.9% of the protest-related tweets. This semi-supervised pipeline, in which a relatively small, annotated dataset guides the labelling of a much larger corpus, exemplifies a scalable strategy for sentiment monitoring when manual annotation is prohibitively time-consuming or expensive. Comparative analysis ensures that only the most reliable classifier is deployed for large-scale sentiment estimation, thereby reducing the risks of misclassification that could compromise subsequent analysis of protest dynamics.\u003c/p\u003e\n\u003cp\u003eA probability threshold was introduced during sentiment prediction to refine the interpretability of sentiment trends further and prevent false positives in identifying potential unrest. Specifically, only sentiment classifications with confidence scores exceeding 0.8 were considered in downstream analyses. This thresholding strategy filters uncertain predictions, ensuring that only high-confidence sentiment outputs influence protest diffusion tracking. This prevents overreaction to sporadic or weakly expressed sentiments that may not reflect the public\u0026apos;s mood. Moreover, this method offers a robust early warning system by correlating temporally aggregated sentiment shifts with topic modelling outputs. Intervention strategies can be triggered when consistently high levels of negative sentiment coincide with discourse around grievance-related themes, such as economic hardship, state repression, or racial tension. Thus, combining semi-supervised classification, comparative model selection, and probabilistic thresholding provides a replicable framework for real-time sentiment monitoring in crisis contexts.\u003c/p\u003e\n\u003cp\u003eNon-negative Matrix Factorisation (NMF) was employed in this study as the topic modelling algorithm to extract latent themes within tweets related to the July 2021 protest movement. The model was applied to tweets originating from KwaZulu-Natal, the epicentre of the unrest, and tweets from provinces where the protest later diffused, specifically Johannesburg and Pretoria (see Tables 9, 10, and 11, respectively). NMF works by factorising the document-term matrix into two non-negative matrices: one representing topics as rows and the other representing documents as columns. In this framework, each tweet is viewed as a composite of multiple topics, where each topic is represented by a distribution of terms. The dimensionality reduction inherent in NMF transforms tweets into sparse, non-negative vectors, making it especially suitable for modelling social media content, which is typically characterised by short text length and thematic fragmentation.\u003c/p\u003e\n\u003cp id=\"_Toc200940953\"\u003eTable 14 Thematic analysis of KwaZulu-Natal tweets\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSN\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eKwaZulu-Natal Topics\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eLooting, stop, looters, stores, shops, busy, burning, continues, said, mall.\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eThis topic centres around the looting and destruction of shops and malls during the unrest. It describes on the ongoing chaos, with mentions of looting, burning, and the government\u0026rsquo;s call to stop the looters.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eZuma, free, released, arrested, release, dudu, support\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eSupporters of Zuma, including his daughter Dudu, continue to call for his release, saying he should no longer remain arrested, while many think the movement to free him will only grow stronger.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eloot, didn\u0026rsquo;t, left, food, love, want, responsibly, looters, excuse, money\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eMany looters didn\u0026rsquo;t leave behind any food or money, using their love for material things as an excuse, while others argue they acted out of desperation and not irresponsibility.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003epeople, black, indians, white, longer, phoenix, killed, saying, stores, thing\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eTensions have escalated in Phoenix, with reports of violence between black and Indian communities, where people are no longer just saying, but acting, leading to stores being destroyed and lives being lost.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esandf, deployed, kzn, need, come, deploying, really, deployment, soldiers, good\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eThe deployment of SANDF soldiers in KZN has been seen as necessary, with many saying that the presence of the military is really needed to restore order, and that the deployment is a good move.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eviolence, chose, taxi, wena, inciting, kzn, acts, abantu, phoenix, cause\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eTaxi drivers in KZN have been accused of inciting violence, particularly in Phoenix, where acts of brutality against abantu (people) have caused widespread fear and chaos\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003epolice, security, looters, minister, army, cluster, private, area, mall, kwamashu\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eThe police, together with private security forces, have been deployed to areas like KwaMashu malls, with the Minister urging the army and security cluster to protect the malls from looters.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003ejacob, president, cyril, zumas, ramaphosa, free, release, arrest, protests, kzn\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eThe arrest of former president Jacob Zuma has led to widespread protests in KZN, with calls for President Cyril Ramaphosa to release him, intensifying the crisis.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003edon\u0026rsquo;t, unrest, saps, know, time, protest, think, im, want, country\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eMany citizens don\u0026rsquo;t know how long the unrest will last, but they think it\u0026apos;s time for the SAPS to take control, as protests continue, and people express their frustration with the state of the country\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esouth, africa, durban, happening, right, protesting, drive, kwazulunatal, lotus, africans\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eIn Durban and across South Africa, protesting is happening right now, with Africans driving the unrest in areas like KwaZulu-Natal and Lotus\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp id=\"_Toc200940954\"\u003eTable 15 Thematic analysis of Johannesburg tweets\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSN\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eJohannesburg Topics\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003elooting, stop, police, looters, shops, say, busy, poor, gauteng, criminality\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eLooting in Gauteng has escalated, with the police struggling to stop looters targeting shops, as criminality thrives amidst the busy and poor conditions.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003ezuma, free, jacob, jail, prison, release, president, ramaphosa, longer, he\u0026rsquo;s\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eDebates continue over Jacob Zuma\u0026apos;s imprisonment, with some calling for his release, while President Ramaphosa is pressured to take a stance\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eloot, come, want, na, responsibly, wan, money, didn\u0026rsquo;t, cup, looters\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eLooters, driven by desperation, want money and resources, though they didn\u0026rsquo;t consider the consequences of their actions, ignoring calls to act responsibly\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eviolence, inciting, taxi, stop, public, arrested, account, incitement, report, peace\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eViolence continues as some incite further unrest, particularly among taxi drivers, leading to arrests and reports calling for peace\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003epeople, black, hungry, steal, white, stealing, unemployed, government, going, jobs\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eEconomic inequality is evident as hungry and unemployed black people resort to stealing, while the government struggles to create jobs and prevent further unrest.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esandf, deployed, deploy, members, saps, need, police, cops, government, situation\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eThe government has deployed SANDF and SAPS members to address the escalating situation, emphasizing the need for a stronger police presence\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esouth, africa, unrest, africans, african, pray, peace, world, jacob, week\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eAs unrest continues in South Africa, Africans across the world pray for peace, reflecting on a turbulent week centred around Jacob Zuma\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esecurity, don\u0026rsquo;t, country, president, amp, like, know, anc, think, time\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eSecurity concerns are mounting in the country, with many questioning the ANC\u0026apos;s leadership and the president\u0026rsquo;s ability to manage the situation in time\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eprotest, peaceful, free, want, criminality, say, level, join, there\u0026rsquo;s, yes\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eCalls for peaceful protest are growing, as people want to express their desire for change without descending into criminality, urging others to join the movement.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 36px;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003emall, glen, jabulani, soweto, maponya, protea, shoprite, ridge, durban, joburg\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 442px;\"\u003e\n \u003cp\u003eMalls like Maponya, Jabulani, and Shoprite in Soweto, Durban, and Joburg have been severely impacted by the unrest, with Glen and Protea Ridge among those affected\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp id=\"_Toc200940955\"\u003eTable 16 Thematic analysis of Pretoria tweets\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"659\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003eSN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003ePretoria Topics\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eInterpretation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003ezuma, jacob, jail, prison, anymore, free, arrested, think, I\u0026rsquo;m, released\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eJacob Zuma is in jail, and people are debating whether he should remain in prison or be released, with many thinking his arrest is unjust.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003elooting, stop, started, shops, police, south, burning, busy, they\u0026rsquo;re, say\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eLooting has erupted across South Africa, with shops burning and the police struggling to stop the unrest that has already started.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eloot, didn\u0026rsquo;t, na, left, need, right, hope, burn, gon, vho\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eLooters didn\u0026rsquo;t leave much behind, driven by a desperate need, hoping they can get what they need before everything burns.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eviolence, inciting, stop, role, taxi, arrested, wena, incite, started, need\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eViolence has escalated, with individuals inciting others, particularly within the taxi industry, leading to arrests as authorities attempt to stop the chaos.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003epeople, hungry, black, poor, going, busy, like, died, arrested, killing\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eThe unrest is deeply rooted in social inequality, with hungry and poor black people feeling like they have no other option, leading to deaths and arrests\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003edon\u0026rsquo;t, steal, know, care, play, things, think, want, understand, let\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eThere\u0026rsquo;s a growing sense of apathy, with many not caring about the consequences of stealing and feeling like people just don\u0026rsquo;t understand or care anymore.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003epresident, law, like, africa, south, jacob, ramaphosa, zumas, release, want\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eThe president is under pressure to uphold the law, with many in South Africa wanting Ramaphosa to release Jacob Zuma\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003eprotest, stealing, going, understand, jail, sa, tht, today, unemployment, funded\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eEconomic desperation has led to protests, with people resorting to stealing and feeling like they\u0026rsquo;re heading for jail if things don\u0026rsquo;t improve in South Africa\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esecurity, minister, state, need, police, tech, smarti, systems, that\u0026rsquo;s, guard\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eThe need for enhanced security is clear, with the Minister emphasizing the importance of police and advanced tech systems to guard against further unrest\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 35px;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 180px;\"\u003e\n \u003cp\u003e\u003cem\u003esandf, saps, deployed, kzn, amp, I\u0026rsquo;m, deployment, police, deploy, ba\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 444px;\"\u003e\n \u003cp\u003eThe SANDF and SAPS have been deployed in KZN, with observers noting how the deployment of these forces will impact the situation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe diffusion of collective violence refers to how civil unrest and protest activity propagate in waves, facilitated by communication networks, protest opportunities, and underlying socio-economic inequalities that mobilise public demand for change. This study\u0026apos;s findings indicate that social media played a central role in the spread of protest during the July 2021 unrest in South Africa. With an estimated internet penetration rate of 72% and widespread smartphone use, platforms like Twitter became crucial for sharing information and gaining support.\u003c/p\u003e\n\u003cp\u003eThe news of the Constitutional Court\u0026rsquo;s decision to imprison former President Jacob Zuma spread quickly online, drawing public sympathy and leading to protests that started in KwaZulu-Natal and spread to almost all provinces. This mobilisation was fuelled by political instability within the ruling party and public frustration over the slow judicial processes related to the State Capture investigation. For many economically disadvantaged citizens, still dealing with inequalities from the apartheid era, Zuma\u0026rsquo;s arrest represented larger issues.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWhat began as a solidarity march at Zuma\u0026rsquo;s Nkandla homestead, where supporters tried to prevent his arrest, quickly turned into widespread looting and property damage in KwaZulu-Natal, before moving to Johannesburg and Pretoria over two weeks. The government\u0026apos;s efforts to restore order led to 354 deaths and more than 8,263 arrests. It is important to mention that attempts to stop the violence were ineffective until Zuma was released, and tensions increased in Phoenix, where clashes between Indian South Africans and Black residents resulted in about 38 deaths.\u003c/p\u003e\n\u003cp\u003eThe South African Human Rights Commission (SAHRC) conducted an inquiry into the July 2021 unrest through a National Hearing Panel, which reviewed 54 oral testimonies and 120 written submissions. The Commission\u0026apos;s findings suggest that the unrest was staged by a group of resourceful primary actors who organised secondary participants, many of whom participated in theft and looting at commercial centres. The report noted that the timing of former President Jacob Zuma\u0026rsquo;s imprisonment and the protests was coincidental, with no direct evidence linking the two events. Although socio-economic hardship was not identified as the primary catalyst for the unrest, the report acknowledged that systemic poverty, inequality, and a widespread lack of trust in the government\u0026apos;s capacity to improve living conditions created fertile ground for secondary actors to vent their frustrations through participation in the unrest. Among its recommendations, the report emphasised the need to develop responsive social media monitoring and intervention mechanisms to facilitate rapid responses to emerging national security threats.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eHow a government responds to public participation in political actions, such as protests, reflects the strength and maturity of its democratic institutions. Responsive governance supports dissent and enables government officials to receive direct feedback on how policies impact society. The diverse interest in protest dynamics, encompassing fields such as sociology, social psychology, political science, economics, and computer science, emphasises the importance of protest studies in understanding, preventing, and managing civil unrest. These insights are crucial for developing effective strategies to restore and maintain urban safety during periods of social disruption.\u003c/p\u003e\n\u003cp\u003eThe analysis of protests related to the imprisonment of former President Jacob Zuma reveals deeper issues of political mismanagement within South Africa’s ruling party. This mismanagement created opportunities for collective mobilisation. By using machine learning techniques, such as sentiment analysis, topic modelling, and data visualisation, the study found that the public generally opposed Zuma’s arrest. The strong presence of negative sentiment in the data shows the lasting legitimacy and support he had among parts of the population, which fuelled protests from KwaZulu-Natal to Pretoria.\u003c/p\u003e\n\u003cp\u003eThe quick spread of information through social media changed a peaceful march into widespread looting, violence, and destruction. Bigrams taken from various provincial datasets reveal deep political tensions within the African National Congress (ANC), echoing historical rivalries like those with the Pan Africanist Congress (PAC) during apartheid. These patterns also uncover underlying economic grievances within marginalised and impoverished communities. The study thus highlights the dual role of social media as a tool for mobilisation and a means for spreading unrest, following the path of the protests from KwaZulu-Natal to the urban centres of Johannesburg and Pretoria.\u003c/p\u003e\n\u003cp\u003eSeveral interconnected factors contributed to the July 2021 unrest in South Africa, stemming from political power struggles, ongoing judicial inquiries, and deep-seated socio-economic issues. The study identifies the structure of the protest's mobilisation and explains its trajectory using collective behaviour theory. The protests initially served as a platform for political expression and civic dissent, but they turned into widespread chaos marked by racial tensions, property destruction, and loss of life. Topic modelling of Twitter discussions revealed economic grievances as key themes, showing that the unrest went beyond loyalty to Jacob Zuma and reflected broader frustrations among marginalised communities.\u003c/p\u003e\n\u003cp\u003eWhat started as a peaceful mobilization in support of Zuma quickly turned into looting and vandalism across several provinces, continuing for two weeks. The clashes in the Phoenix community highlighted long-standing racial divisions, worsening the unrest. These findings line up with the South African Human Rights Commission’s investigation, which found no direct link between Zuma’s imprisonment and the violence. However, socio-economic inequality and a lack of public trust in government drove participation. Based on these observations, the life cycle of the July 2021 protest can be understood in four stages: protest, looting, destruction, and change. “Protest” includes the initial civic mobilization; “looting” refers to the opportunistic exploitation of the unrest; “destruction” describes the violence and fatalities that followed; and “change” indicates the shifts in public awareness or government response triggered by the collective action.\u003c/p\u003e\n\u003cp\u003eThe study shows the benefits of combining machine learning with social science research by using computational methods to analyse large amounts of text and gain insights into complex social issues. Through sentiment analysis and topic modellingg, the research demonstrates how monitoring public perception on social media can provide an essential feedback mechanism for policymakers, businesses, and government officials. These insights can inform policy formation, judicial reviews, and crisis management decisions, ultimately contributing to the creation of safer and more responsive urban environments.\u003c/p\u003e\n\u003cp\u003eThe visualizations produced in the study improve understanding by revealing underlying patterns and relationships within the discussions. However, the study has limitations. Its focus on Twitter data from certain South African provinces during the July 2021 unrest may limit the applicability of the findings and broader causal connections. Future research could benefit from incorporating data from multiple social media platforms and expanding the geographical scope to provide a more comprehensive view of protest dynamics.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch3\u003eConflict of Interest Statement\u003c/h3\u003e\n\u003cp\u003eOn behalf of all authors, the corresponding author states that there is no conflict of interest.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAggarwal, C. C., \u0026amp; Zhai, C. (2012). A survey of text clustering algorithms. Mining text data, 77-128.\u003c/li\u003e\n\u003cli\u003eAlexander, P. (2013). Marikana, turning point in South African history. Review of African Political Economy, 40(138), 605-619.\u003c/li\u003e\n\u003cli\u003eBishop, C. M., \u0026amp; Nasrabadi, N. M. (2006). Pattern recognition and machine learning (Vol. 4, No. 4, p. 738). New York: springer.\u003c/li\u003e\n\u003cli\u003eBond, P., \u0026amp; Mottiar, S. (2013). Movements, protests and a massacre in South Africa. Journal of Contemporary African Studies, 31(2), 283-302.\u003c/li\u003e\n\u003cli\u003eCloward, R. A., \u0026amp; Piven, F. F. (1977). The acquiescence of social work. Society, 14(2), 55-63.\u003c/li\u003e\n\u003cli\u003eCoffey, R. (2022). The Sharpeville Massacre, 1960: African Activism and the Press. In The British Press, Public Opinion and the End of Empire in Africa: The\u0026apos;Wind of Change\u0026apos;, 1957-60 (pp. 169-212). Cham: Springer International Publishing.\u003c/li\u003e\n\u003cli\u003eDaniel Jurafsky and James H. Martin. 2025. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models, 3rd edition. Online manuscript released January 12, 2025. https://web.stanford.edu/~jurafsky/slp3.\u003c/li\u003e\n\u003cli\u003eDavies, J. C. (1962). Toward a theory of revolution. American sociological review, 5-19.\u003c/li\u003e\n\u003cli\u003eDella Porta, D. (2015). Social movements in times of austerity: Bringing capitalism back into protest analysis. John Wiley \u0026amp; Sons.\u003c/li\u003e\n\u003cli\u003eDesai, A. (2023). Geographies of racial capitalism: the 2021 July riots in South Africa. Ethnic and Racial Studies, 46(16), 3542-3561.\u003c/li\u003e\n\u003cli\u003eDubow, S. (2015). Were there political alternatives in the wake of the Sharpeville-Langa violence in South Africa, 1960?. The Journal of African History, 56(1), 119-142.\u003c/li\u003e\n\u003cli\u003eEisinger, P. K. (1973). The conditions of protest behavior in American cities. American political science review, 67(1), 11-28.\u003c/li\u003e\n\u003cli\u003eFeenstra, R. A. (2015). Activist and citizen political repertoire in Spain: A reflection based on civil society theory and different logics of political participation. Journal of civil society, 11(3), 242-258.\u003c/li\u003e\n\u003cli\u003eFortuna, P., \u0026amp; Nunes, S. (2018). A survey on automatic detection of hate speech in text. Acm Computing Surveys (Csur), 51(4), 1-30.\u003c/li\u003e\n\u003cli\u003eGamson, W. A. (1991, March). Commitment and agency in social movements. In Sociological forum (Vol. 6, No. 1, pp. 27-50). New York: Kluwer Academic Publishers-Plenum Publishers.\u003c/li\u003e\n\u003cli\u003eGoodfellow, I., Bengio, Y., \u0026amp; Courville, A. (2016). Deep learning. Book in preparation for MIT Press. URL\u0026iexcl; http://www. deeplearningbook. org, 1.\u003c/li\u003e\n\u003cli\u003eGriffiths, D. (2019). # FeesMustFall and the decolonised university in South Africa: Tensions and opportunities in a globalising world. International Journal of Educational Research, 94, 143-149.\u003c/li\u003e\n\u003cli\u003eGurr, T. (1968). A causal model of civil strife: a comparative analysis using new indices1. American political science review, 62(4), 1104-1124.\u003c/li\u003e\n\u003cli\u003eHastie, T., Tibshirani, R., Friedman, J., Hastie, T., Tibshirani, R., \u0026amp; Friedman, J. (2009). Overview of supervised learning. \u003cem\u003eThe elements of statistical learning: Data mining, inference, and prediction\u003c/em\u003e, 9-41.\u003c/li\u003e\n\u003cli\u003eHollister, R. W. (2023). The Sharpeville Massacre, Violence, and the Struggles of the African National Congress, 1960-1990. Armstrong Undergraduate Journal of History, 13(1), 62-75.\u003c/li\u003e\n\u003cli\u003eJurafsky, D., \u0026amp; Martin, J. H. (2009). Speech and Language Processing.\u003c/li\u003e\n\u003cli\u003eKaldor, M., Selchow, S., \u0026amp; Murray-Leach, T. (Eds.). (2015). Subterranean politics in Europe. Springer.\u003c/li\u003e\n\u003cli\u003eKlandermans, B., \u0026amp; Oegema, D. (1987). Potentials, networks, motivations, and barriers: Steps towards participation in social movements. American sociological review, 519-531.\u003c/li\u003e\n\u003cli\u003eLanga, M. (2017). Researching the# FeesMustFall movement. Hashtag: An analysis of the# FeesMustFall movement at South African universities, 6-12.\u003c/li\u003e\n\u003cli\u003eLiu, B. (2012). Sentiment analysis: A fascinating problem. In \u003cem\u003eSentiment analysis and opinion mining\u003c/em\u003e (pp. 1-8). Cham: Springer International Publishing.\u003c/li\u003e\n\u003cli\u003eLiu, B. (2022). Sentiment analysis and opinion mining. Springer Nature.\u003c/li\u003e\n\u003cli\u003eMcAdam, D. (1983). Tactical innovation and the pace of insurgency. American sociological review, 735-754.\u003c/li\u003e\n\u003cli\u003eMcCarthy, J. D., \u0026amp; Zald, M. N. (1977). Resource mobilization and social movements: A partial theory. American journal of sociology, 82(6), 1212-1241.\u003c/li\u003e\n\u003cli\u003eMelucci, A., Keane, J., \u0026amp; Mier, P. (1989). Nomads of the present: Social movements and individual needs in contemporary society.\u003c/li\u003e\n\u003cli\u003eMitchell, T. M. (1997). Does Machine Learning Really Work?. AI Magazine, 18(3), 11-20.\u003c/li\u003e\n\u003cli\u003eMolteno, F. (1979). The uprising of 16th June: A review of the literature on events in South Africa 1976. Social Dynamics, 5(1), 54-89.\u003c/li\u003e\n\u003cli\u003eMontero, J. R., Gunther, R., \u0026amp; Torcal, M. (1997). Democracy in Spain: Legitimacy, discontent, and disaffection. Studies in comparative international development, 32, 124-160.\u003c/li\u003e\n\u003cli\u003eMooijman, M., Hoover, J., Lin, Y., Ji, H., \u0026amp; Dehghani, M. (2018). Moralization in social networks and the emergence of violence during protests. Nature human behaviour, 2(6), 389-396.\u003c/li\u003e\n\u003cli\u003eMurphy, K. P. (2012). Machine learning: a probabilistic perspective. MIT press.\u003c/li\u003e\n\u003cli\u003eNaidoo, K., Lewis, S., Essop, H., Koch, G. G., Khoza, T. E., Phahlamohlaka, N. M., \u0026amp; Badriparsad, N. R. (2023). July 2021 civil unrest: South African diagnostic radiography students\u0026rsquo; experiences. Health SA Gesondheid, 28(1).\u003c/li\u003e\n\u003cli\u003eNdlovu, S. M. (2006). The soweto uprising. The road to democracy in South Africa, 2, 1970-1980.\u003c/li\u003e\n\u003cli\u003eOpp, K. D., \u0026amp; Gern, C. (1993). Dissident groups, personal networks, and spontaneous cooperation: The East German revolution of 1989. American sociological review, 659-680.\u003c/li\u003e\n\u003cli\u003ePang, B., \u0026amp; Lee, L. (2008). Opinion mining and sentiment analysis. Foundations and Trends\u0026reg; in information retrieval, 2(1\u0026ndash;2), 1-135.\u003c/li\u003e\n\u003cli\u003ePhungula, N. (2024). Understanding the dynamics of South Africa\u0026rsquo;s July 2021 social unrest. \u003cem\u003eJournal of Nation-Building and Policy Studies\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e(1), 71.\u003c/li\u003e\n\u003cli\u003eSamuel, A. L. (1959). Some studies in machine learning using the game of checkers. IBM Journal of research and development, 3(3), 210-229.\u003c/li\u003e\n\u003cli\u003eShalev-Shwartz, S., \u0026amp; Ben-David, S. (2014). Understanding machine learning: From theory to algorithms. Cambridge university press.\u003c/li\u003e\n\u003cli\u003eSnow, D. A., \u0026amp; Benford, R. D. (2005). Clarifying the relationship between framing and ideology. Frames of protest: Social movements and the framing perspective, 205, 209.\u003c/li\u003e\n\u003cli\u003eSutton, R. S., \u0026amp; Barto, A. G. (2018). \u003cem\u003eReinforcement learning: An introduction\u003c/em\u003e. MIT Press.\u003c/li\u003e\n\u003cli\u003eTarrow, S. (2022). \u003cem\u003ePower in movement\u003c/em\u003e. Cambridge university press.\u003c/li\u003e\n\u003cli\u003eTilly, C. (2004). Social boundary mechanisms. Philosophy of the social sciences, 34(2), 211-236.\u003c/li\u003e\n\u003cli\u003eTollefson, J. (2024). Protests over Israel-Hamas war have torn US universities apart: what\u0026apos;s next?. Nature.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"civil unrest, diffusion of protest behaviour, sentiment analysis, topic modelling","lastPublishedDoi":"10.21203/rs.3.rs-7760005/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7760005/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The jailing of former South African President Jacob Zuma in July 2021 ignited protests in KwaZulu-Natal and quickly spread to Johannesburg and Pretoria. Initially peaceful, these demonstrations escalated into widespread looting and violence, culminating in the deaths of 354 individuals. This study employs a mixed-methods approach by integrating machine learning and qualitative analysis to examine the dynamics of the unrest using Twitter data from multiple South African provinces. Tweets were manually annotated for sentiment (positive, negative, neutral), and inter-annotator agreement was measured using Fleiss' Kappa, yielding a score of 0.27, indicative of fair consensus. The most accurate sentiment classification model labelled the remaining dataset, enabling the temporal tracking of sentiment and protest diffusion. Findings underscore the role of economic inequality, political instability, and racial tensions in fuelling the unrest. Also, the lifecycle of the protests progresses from mobilisation to looting, violence, and change, which was confirmed through social media discourse. The study highlights the utility of social media analytics in complementing investigative journalism and informing state responses to urban crises. It contributes to the growing literature on collective violence and the diffusion of protest behaviour in digitally networked societies.","manuscriptTitle":"Diffusion of protest behaviour: Analysing the July 2021 civil unrest in South Africa through sentiment analysis and topic modelling","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-03 08:39:46","doi":"10.21203/rs.3.rs-7760005/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"026203a9-d627-4d2b-9129-2398c5802951","owner":[],"postedDate":"October 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-11-13T11:08:30+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-03 08:39:46","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7760005","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7760005","identity":"rs-7760005","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0