Exploring the Political Biases in Pakistani Tweets Dataset

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Study is conducted on Exploring political biases in Pakistani twitter data set. The objectives of the study is the way Pakistani twitter users express their biases towards different political parties or personalities and the way they advance their political agenda. Mixed methodology approach is used in the study. Biber’s six dimensions, MAT tagger and Douglas Biber’s “variation across speech and writing” is utilized in comparing the analysis 1 and analysis 2. Analysis 2 based on the Corpus received from the department is identical to Analysis 1 from the Douglas Biber’s “ variation across speech and writing” except few genres which indicates little variations from the standard or expected usage of language. Result of the study shows that participants involved in conversation use nouns, pronouns, adjectives, verbs and long words to express their biases towards different political parties or personalities, and to advance their political agenda. They use the argumentative and persuasive techniques to convince the interlocutors on the twitter.
Full text 73,257 characters · extracted from preprint-html · click to expand
Exploring the Political Biases in Pakistani Tweets Dataset | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Exploring the Political Biases in Pakistani Tweets Dataset Noor Zaman Khan Khan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8325825/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Study is conducted on Exploring political biases in Pakistani twitter data set. The objectives of the study is the way Pakistani twitter users express their biases towards different political parties or personalities and the way they advance their political agenda. Mixed methodology approach is used in the study. Biber’s six dimensions, MAT tagger and Douglas Biber’s “variation across speech and writing” is utilized in comparing the analysis 1 and analysis 2. Analysis 2 based on the Corpus received from the department is identical to Analysis 1 from the Douglas Biber’s “ variation across speech and writing” except few genres which indicates little variations from the standard or expected usage of language. Result of the study shows that participants involved in conversation use nouns, pronouns, adjectives, verbs and long words to express their biases towards different political parties or personalities, and to advance their political agenda. They use the argumentative and persuasive techniques to convince the interlocutors on the twitter. Twitter agenda variation participants argumentative persuasive Figures Figure 1 Figure 2 INTRODUCTION Around the world, Twitter is one of the most important data sources for research on political and social communication, and Twitter datasets have been utilized for various projects examining computational politics, polarization, propaganda, and political trolling. Western nations, including the US, are the subject of the majority of publicly available datasets and study. The population of Pakistan is still substantially understudied in social network analysis and social computing research despite the fact that Twitter is one of the key platforms for political conversation there. In the general election of 2013 Twitter was heavily utilized by Pakistani major political parties to advance their agenda through different social media campaigns. In April 2022 Pakistan had to face the biggest crises of politics and law when president Imran khan had to leave the seat because of a no-confidence vote which created a big political controversy on Twitter and other social media platforms. Political biases and favoritism is still on the peak. There is a social media war prevailing in the arena of Twitter. Political discussions on social media have become increasingly prevalent in Pakistani society, with Twitter being a popular platform for individuals to express their opinions and biases towards different political parties and personalities. This research aims to explore how Pakistani Twitter users express their biases and what strategies or tactics they employ to advance their agenda. Understanding these patterns of online behavior can shed light on the current political climate in Pakistan and provide insights into how social media is impacting political discourse in the country. This Corpus based study will help to know the actual use of language in politics in the Pakistani community. STATEMENT OF THE PROBLEM Social media platforms, particularly Twitter, have become important spaces for political engagement and discussion in Pakistan. However, these platforms often reflect and amplify users’ political biases, leading to polarized discourse and the spread of agenda-driven narratives. Despite the growing influence of Twitter in shaping public opinion, there is limited research on the nature, frequency, and impact of biased expressions by Pakistani Twitter users toward political parties and personalities. Furthermore, little is known about the strategies and techniques employed by users to promote their political agendas and how these biases influence the overall discourse. Understanding these patterns is essential to foster balanced and constructive political discussions online and to mitigate the reinforcement of partisan biases in digital spaces. OBJECTIVES This study seeks to: To examine the linguistic strategies and techniques employed by Pakistani Twitter users to advance their political agendas, including the use of persuasive and argumentative language. To analyze the nature of biased expressions in Pakistani Twitter discourse, focusing on the use of nouns, adjectives, verbs, and other linguistic features that convey favoritism toward specific political parties or personalities. RESEARCH QUESTIONS What linguistic strategies and techniques do Pakistani Twitter users employ to advance their political agendas? How do Pakistani Twitter users express their biases toward different political parties or personalities through their language use, including nouns, adjectives, verbs, and other lexical features? LITERATURE REVIEW The use of social media platforms such as Twitter has become increasingly popular in expressing political biases in Pakistan. Many Pakistani Twitter users use the platform to express their political affiliations and opinions on political parties and personalities. Studies have shown that Twitter is a popular platform for political conversations and discourse, particularly during election periods (Khan, Hussain, & Wahab, 2018 ). Mentions are also used to advance political agendas on Twitter in Pakistan. Users can mention other users on Twitter to draw attention to their message or to challenge them. On Twitter, it is common for political opponents to engage in heated debates and arguments using mentions. For example, during the 2018 general elections in Pakistan, the Twitter accounts of different political parties and personalities engaged in debates using mentions and often engaged in heated debates and arguments (Khan, Hussain, & Wahab, 2018 ). Twitter users in Pakistan often use tactics such as hashtags, retweets, and mentions to advance their agenda. Hashtags are a common way to promote a particular agenda or viewpoint on Twitter. For example, during the 2018 general elections in Pakistan, supporters of the Pakistan Tehreek-e-Insaf (PTI) used hashtags such as #TabdeeliAaRahiHai (“Change is coming”) and #NayaPakistan (“New Pakistan”) to promote their party, while supporters of other parties used hashtags such as #VoteKoIzzatDo (“Respect your vote”) to criticize the ruling party (Abbasi & Rehman, 2020 ). Retweeting is another tactic commonly used to advance political agendas on Twitter in Pakistan. By retweeting a tweet, a user can help amplify a message and spread it to a larger audience. Twitter users in Pakistan often use retweeting to support their political parties or personalities and to criticize their opponents. For example, during the 2021 Senate elections in Pakistan, supporters of the Pakistan Peoples Party (PPP) retweeted a video of a member of the ruling party allegedly buying votes, while the ruling party’s supporters retweeted videos alleging that the opposition parties were involved in horse-trading (Butt et al., 2021 ). Shahzad Ashraf et al. ( 2021 ) analyzed the sharing pattern of information among Pakistani politicians on Twitter. The study aimed to identify the sources of information shared by politicians and how these sources impacted political discourse. The study found that politicians engaged in strategic sharing of information, and the type of information shared varied depending on the political motivation of the politician. Finally, a study conducted by Hassan Mahmood et al. ( 2021 ) explored media bias and toxicity in South Asian political discourse. The study analyzed news articles and Twitter conversations related to South Asian politics and identified the presence of media bias and toxicity in the discourse. The study concluded that media outlets and politicians played a significant role in shaping the discourse and that efforts need to be made to promote healthy and balanced political communication. A study conducted by Naeem Ahmed Ibupoto et al. ( 2022 ) aimed to perform sentiment analysis on current political topics in Pakistan’s Twitter user base. The study used machine learning algorithms to analyze tweets related to different political topics and determine the general sentiment expressed by users. The study found that there were significant variations in sentiment on different political topics and the sentiments expressed on Twitter were influenced by the media and the government. Another study by Zeeshan Rasheed et al. ( 2022 ) focused on collecting and analyzing a Twitter dataset of Pakistani political discourse. The study aimed to reveal the types of political topics discussed on Twitter and the main users who engaged in these discussions. The study found that political discussions on Twitter were largely dominated by mainstream politicians and their supporters. The use of Twitter for political discourse in Pakistan is an important area of study for researchers. The studies discussed in this literature review highlight the impact of sentiment analysis, information sharing, media bias, and toxicity on the discourse. Some Researches have also shown that Pakistani Twitter users express their biases towards different political parties and personalities using a range of tactics such as hashtags, retweets, and mentions. These tactics are commonly used to advance political agendas and promote political views. This research can be useful for policymakers and stakeholders to promote healthy political communication in the country and to help to know the actual use of language in Pakistani political discussions and debates on twitter. Research Gap While previous studies have examined political discourse on Twitter in Pakistan, much of the existing research has primarily focused on sentiment analysis, information sharing, and the use of specific tactics such as hashtags, retweets, and mentions. Although these studies provide insights into general patterns of political engagement, there is limited research specifically investigating the frequency, nature, and strategies of biased expressions by Pakistani Twitter users toward different political parties and personalities. Additionally, few studies have explored the underlying factors that contribute to the formation and reinforcement of these biases or the broader impact of agenda-driven strategies on overall political discourse. Most of the current research emphasizes either individual user behavior or political actors, without offering a comprehensive analysis of how biases influence the structure and tone of online political discussions in Pakistan. Therefore, there is a need for a focused, corpus-based study that examines the linguistic and strategic dimensions of political bias on Pakistani Twitter, providing a deeper understanding of how political agendas are advanced and how discourse is shaped on this platform. METHODOLOGY Research Design and Data Sampling Mixed methodology approach is being used here. We have received a Corpus of tweets from the department based on hashtags.We have also selected some hashtag words upon which the corpus This corpus is to be run on the Mat tagger software. Another standard English is also to be run on the Mat Tagger. Values of corpora are to be compared and frequent linguistic features are to be analyzed using some dimensions or theories. Dimension 1 of Biber’s six dimensions is used here to analyze the text. Mat tagger Multidimensional Analysis Tagger (MAT) is a software tool designed for text analysis. It is used to tag words and phrases in a text document with various categories such as topic, sentiment, emotion, status, and more. MAT uses advanced natural language processing algorithms and machine learning techniques to extract meaning and context from textual data. It is commonly used in social media monitoring, market research, and customer feedback analysis to identify patterns and trends in user-generated content. Biber’s (1988) Dimensions The conflict between involved and informational discourse is dimension 1. High scores on this variable suggest that the text is emotive and interactive, like a casual chat, while low levels suggest that the language is informationally rich, like academic prose. In contrast, a low score on this Dimension indicates that the text uses a lot of nouns, big words, and adjectives. A high score on this Dimension indicates that the text uses a lot of verbs and pronouns. The conflict between narrative and non-narrative concerns is Dimension 2. High scores on this variable suggest that the text is narrative, such as a novel, while low levels indicate that the material is non-narrative. A high score for this Dimension indicates that, among other things, the text uses third person pronouns and several past tenses. Context-Independent Discourse and Context-Dependent Discourse are in opposition to one another in Dimension 3. A high score on this variable indicates that the text is not dependent on the context, such as scholarly prose, while a low score indicates that the text is dependent on the context, such as in the case of a sports broadcast. A high score on this Dimension denotes the text’s frequent use of nominalizations, whereas a low score denotes the text’s frequent use of adverbs, among other qualities. Overt Expression of Persuasion is measured by Dimension 4. High scores for this variable show that, for example in business letters, the author’s point of view and evaluation of likelihood and/or certainty are clearly marked in the text. A high score on this dimension indicates that, among other things, the text uses a lot of modal verbs. The conflict between abstract and non-abstract information is dimension five. High scores for this variable suggest that the text presents information in a formal, technical, and abstract manner, like that found, for instance, in scientific speech. A text with a high score on this Dimension contains numerous passive sentences and conjuncts, among other qualities. Measurement 6 refers to Online Explanation of Information. High ratings for this characteristic suggest that the text is informative in nature but was created under time constraints, as those found in speeches, for example. A text that performs well on this Dimension exhibits numerous post modifications of noun phrases, among other traits. DATA ANALYSIS AND INTERPRETATION Analysis 1 Corpus Analysis from the Book “Variation across speech and writing by Douglas Biber Dimension 1 Figure 1. Dimension 1 from Douglas Biber’s “Variation across speech and writing” Analysis 2 MAT Analysis of corpus based on tweets Figure 2. Dimension 1 from Received corpus from the department Value score of analysis 1 and analysis 2 is compared to know the diversity of language usage from the standard or the expected usage The corpus received from the department that was based on tweets was analyzed using MAT Tagger software. This Analysis is referred to as “Analysis 2”. The values of this corpus we are going to compare with the values of Corpus Analysis from Douglas Biber’s Book “Variation across speech and writing”. This corpus Analysis is referred to as “analysis 1”. Only the values of dimensions 1 are being compared. Here values of the properties of analysis 1 and analysis 2 are being compared here. Both of Analysis are displaying their properties on the X axis and their score on the Y axis. The score of the genre of conversation in analysis 1 and analysis 2 is 36. Value of Broadcast is 3 in analysis 1 and analysis 2. Values of prepared speeches in analysis 1 and analysis 2 are 2 and 1.5 respectively. Value of personal letters in analysis 1 and analysis 2 is 20. Value of general fiction in analysis 1 and analysis 2 is -1. Values of press reportage in analysis 1 and analysis 2 are − 8 and − 7 respectively. Value of academic prose in analysis 1 and analysis 2 is -8. Value of the official document in analysis 1 and analysis 2 is -18. These are differences between analysis 1 and analysis 2 in terms of scores. The closest genre in analysis 1 and analysis 2 is official documents. Table 1 Comparison of Fig. 1 and Fig. 2 in terms of values. Sr no Genres Analysis 1 Analysis 2 Variation score 1 Conversation 36 36 0 2 Broadcasts 3 3 0 3 Prepared speeches 2 1.5 0.5 4 Personal letters 20 20 0 5 General fiction -1 -1 0 6 Press reportage -8 -7 -1 7 Academic prose -8 -8 0 8 Official documents -18 -18 0 Text types or genres found in analysis 1 and analysis 2 show little variations in terms of score of the genres. The main difference is found in only two genres: prepared speeches and press reportage which is 0.5 and − 1 respectively. It seems that the language used in the both corpora is identical except prepared speeches and press reportage language. The language of these two genres show a little deviation. The values of dimension 1 in analysis 1 are greater than values of dimension 1 in analysis 2. It shows that analysis 2 shows deviation of language used in Received corpus from the Analysis 1 that is found in Douglas Biber’s “variation across speech and writing. Table 2 Mean values of Biber’s dimensions found in the Corpus based on the tweets. Dimension 1 Dimension 2 Dimension 3 Dimension 4 Dimension 5 Dimension 6 Closest text type -22.19 -2.39 2.46 -3.46 -3.18 -2.57 Learned exposition Values of Biber’s six dimensions found in the Received corpus based on the tweets data set are displayed in Table 2 . All six dimensions contain different values which suggest the different nature of language used in tweets as political discourse. The dimension we are concerned about is dimension 1 which is The conflict between involved and informational discourse. High scores on this variable suggest that the text is emotive and interactive, like a casual chat, while low levels suggest that the language is informationally rich, like academic prose. In contrast, a low score on this Dimension indicates that the text uses a lot of nouns, big words, and adjectives. A high score on this Dimension indicates that the text uses a lot of verbs and pronouns. So here is the low value of dimension 1 which indicates that the text is informational such as academic prose. Low scores also suggest that text uses lots of nouns, big words and adjectives. RESULT AND DISCUSSION The question on which the study is conducted is “ How do Pakistani twitter users express their biases towards different political parties or personalities?What kind of strategies and techniques do Pakistani Twitter users use to advance their political agenda?” MAT Analysis of Corpus based on tweets shows that most of the language is used in conversation. The Value of conversation is 36 in Dimension 1 which is conflict between involved and informational discourse. If the values of genres are low then the text is informational and if the values of genres are high then the text is interactional and emotive. To make the conversation interactive and emotive nouns and verbs are used. Closest genres in analysis 1 and analysis 2 of dimension 1 is the official document and learned exposition. Learned exposition, academic prose and official documents serve as informational. Nouns and verbs are used in informational text. Value of conversation in Analysis of dimension 1 in both corpora is 36 which is highest which indicates that text is also emotive and interactive. If we compare the mean value of Biber’s six dimesions the mean value of dimensions 1 is the lowest value which is -22.19 that is shown in Table 2 which indicates that the overall nature of text is informational as compared to the nature of other dimensions of Biber. Academic prose and official documents are considered as informational. Participants involved in conversation make the conversation interactive and emotive. Interlocutors in conversation on any platform use persuasive and argumentative strategies or techniques to convince the participants or interlocutors involved in the face to face or online textual conversation. Along with the persuasive and argumentative strategies and techniques Interlocutors use the excess of nouns, pronouns, adjectives, big words and verbs to make the conversation informational, interactive and emotive. So we can say that Pakistani twitter users express their favoritism and biases towards different politicians or parties Using nouns, adjectives ,verbs ,big words such as Imran khan, Azeem leader, change, corruption etc. In tweets they indicate the flaws of opposite parties and qualities of their own party. In this regard they use argumentative and persuasive techniques as used by the interlocutors in the talk shows. They make the online conversation on twitter interactive,emotive and informational using nouns, adjectives, big words, and verbs and videos of the different parties as evidence to support or oppose their conversation as well as parties. MAT analysis shows that the actual usage of language used on twitter platform shows little deviation from the expected usage of language as used according to standard norms of language but their language is organized according to their usage in politics to support or defend the parties or the agendas. CONCLUSION Exploring political biases in Pakistani twitter data set is the topic of the study. The questions on which we have conducted the study are “ How do Pakistani twitter users express their biases towards different political parties or personalities?What kind of strategies and techniques do Pakistani Twitter users use to advance their political agenda?”. Mixed methodology is used in the study. Analysis 1 from the Douglas Biber’s “variation across speech and writing” and analysis 2 of the Corpus received from the department are compared in terms of their scores. Data is analyzed and interpreted using dimensions 1 of Biber’s six dimensions. Scores of genres in dimensions of analysis 2 shows little bit variation from the analysis 1. Mean value of dimension 1 of the analysis 2 is -22.19. Scores of genres in dimension 1 of analysis 2 shows that text is interactional, emotive and informational. Value of dimension 1 is the lowest value which indicates that the overall nature of the text is informational. Participants of the conversation on twitter use nouns, adjectives, verbs and big words to advance political agenda and express their biases towards political parties .they also use persuasive and argumentative strategies to do so. Declarations Author Contribution Noor Zaman Khan conceptualized the study, designed the research framework, and collected the data. The author conducted the corpus-based analysis using MAT Tagger software and applied Biber’s multidimensional analysis. Noor Zaman Khan performed the interpretation of results, drafted the manuscript, and revised it critically for important intellectual content. The author approves the final version of the manuscript and agrees to be accountable for all aspects of the work, ensuring accuracy and integrity. References Abbasi, J. A., & Rehman, N. U. (2020). Tweeting for political change: The impact of hashtags during the 2018 Pakistan general elections. Telematics and Informatics , 49 , 101382. https://doi.org/10.1016/j.tele.2020.101382 Ashraf, S., Khamis, S., & Weldon, L. (2021). Twitter dialogue: An analysis of Pakistani politicians’ information sharing. Social Media + Society , 7 (1), 20563051211003620. https://doi.org/10.1177/20563051211003620 Butt, S. A., Aslam, S., & Asghar, T. (2021). Use of social media for political discourse in Pakistan: Evidence from 2021 Senate elections. TechTrends , 65 , 1–10. https://doi.org/10.1007/s11528-021-00619-7 Ibupoto, N. A., Rasheed, Z., Ashraf, S., Ali, M., Tanweer, R., & Sen, S. (2022). Sentiment analysis on current political topics in Pakistan’s Twitter user bases (pp. 1–21). Preprint. https://doi.org/10.21203/rs.3.rs-2095172/v1 Khan, M. S., Hussain, A., & Wahab, F. (2018). Political discourse on social media: A case study of Twitter use in Pakistan’s 2018 general elections. Journal of Asian Pacific Communication , 28 (2), 222–241. https://doi.org/10.1075/japc.28.2.05kha Mahmood, H., Ayub, A., & Jabeen, U. (2021). Exploring media bias and toxicity in South Asian political discourse. New Media & Society , 23 (8), 2317–2336. https://doi.org/10.1177/1461444820915995 Rasheed, Z., Ibupoto, N. A., Ashraf, S., Ali, M., & Sen, S. (2022). A Twitter dataset for Pakistani political discourse. Data in Brief , 38, 107277. https://doi.org/10.1016/j.dib.2022.107277 . Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8325825","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":560240507,"identity":"eaeba5b4-098d-428b-80dc-d24cfe570b42","order_by":0,"name":"Noor Zaman Khan Khan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA80lEQVRIiWNgGAWjYHCDw4eZGSqANDNzA5E6GI8lMzOcAWlhJFYL8xljZsY2sGb8WszZm499/FGzTd6c7YyxceG82mj+dqCWHxXbcGqx7DmWPEPi2G3DnT3HipNnbjueO+MwYwNjz5nbOLUY3MgxZjBgu8244cbhzYd5tx3LbQBqAbqQgJaEf7ftN9x/YHyYd86x3PlEaTnYdjtxw4Ejxsm8DTW5GwhqOXMsmbGx73byhgPHko1nHDuQuxGo5SBevxxvPsz449tt2w0HDh+WLqipy513/vDBBz8qcGtBB4fB5AGi1QNBHSmKR8EoGAWjYIQAAOggZUo0YbkuAAAAAElFTkSuQmCC","orcid":"","institution":"University of Education","correspondingAuthor":true,"prefix":"","firstName":"Noor","middleName":"Zaman Khan","lastName":"Khan","suffix":""}],"badges":[],"createdAt":"2025-12-10 09:53:20","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8325825/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8325825/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":98752091,"identity":"bf08962c-345f-49ae-abfa-d2157e268c4c","added_by":"auto","created_at":"2025-12-22 09:14:09","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":161489,"visible":true,"origin":"","legend":"","description":"","filename":"Article.docx","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/622fb3e67caf59fd990907bf.docx"},{"id":98779735,"identity":"67283163-985e-4272-960f-9c42b92d72dc","added_by":"auto","created_at":"2025-12-22 12:30:40","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3362,"visible":true,"origin":"","legend":"","description":"","filename":"926eda01a45d4bc688aa3dab5d6c392b.json","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/5b356aed7e4d315fb2e10b50.json"},{"id":98752088,"identity":"57175d49-6045-45bb-97bb-fb7cb79fe820","added_by":"auto","created_at":"2025-12-22 09:14:09","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":46091,"visible":true,"origin":"","legend":"","description":"","filename":"926eda01a45d4bc688aa3dab5d6c392b1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/730db425aa6d427019602626.xml"},{"id":98778233,"identity":"67ddd65b-7649-4218-8d26-e697da4fd5de","added_by":"auto","created_at":"2025-12-22 12:29:02","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":37068,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/dc1e96372632c81652a3be51.png"},{"id":98778159,"identity":"0b4d6c04-5b50-4847-90cb-46f2478e788f","added_by":"auto","created_at":"2025-12-22 12:28:56","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":17239,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/5c222efe81b17e4728178f89.png"},{"id":98752096,"identity":"c3f69e54-55f0-496c-b6b2-738c2193dd99","added_by":"auto","created_at":"2025-12-22 09:14:09","extension":"xml","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":43010,"visible":true,"origin":"","legend":"","description":"","filename":"926eda01a45d4bc688aa3dab5d6c392b1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/28d7fc501cadbd3f33d8e15b.xml"},{"id":98752094,"identity":"53bc445b-9cbb-41fa-9bcf-e5d8f1c5d7f7","added_by":"auto","created_at":"2025-12-22 09:14:09","extension":"html","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":49653,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/67d2c706ae161c23b93f1491.html"},{"id":98752090,"identity":"8b087288-7dad-403c-b4fa-408d222cb2b2","added_by":"auto","created_at":"2025-12-22 09:14:09","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":71351,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDimension 1 from Douglas Biber’s “Variation across speech and writing”\u003c/em\u003e\u003cstrong\u003eAnalysis 2\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/a08a6a497795a427082b0e9f.png"},{"id":98779933,"identity":"8ad6446b-686f-45ec-b036-2a0c406e2f37","added_by":"auto","created_at":"2025-12-22 12:30:56","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":71510,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDimension 1 from Received corpus from the department\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/69ff9acef930c397e74899fb.png"},{"id":99867611,"identity":"a2150bb0-e74b-48d1-addf-1d7cb655e178","added_by":"auto","created_at":"2026-01-09 08:25:17","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":638462,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8325825/v1/e6e123c4-1939-409d-8648-84907980c859.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eExploring the Political Biases in Pakistani Tweets Dataset\u003c/p\u003e","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eAround the world, Twitter is one of the most important data sources for research on political and social communication, and Twitter datasets have been utilized for various projects examining computational politics, polarization, propaganda, and political trolling. Western nations, including the US, are the subject of the majority of publicly available datasets and study. The population of Pakistan is still substantially understudied in social network analysis and social computing research despite the fact that Twitter is one of the key platforms for political conversation there.\u003c/p\u003e \u003cp\u003eIn the general election of 2013 Twitter was heavily utilized by Pakistani major political parties to advance their agenda through different social media campaigns. In April 2022 Pakistan had to face the biggest crises of politics and law when president Imran khan had to leave the seat because of a no-confidence vote which created a big political controversy on Twitter and other social media platforms. Political biases and favoritism is still on the peak. There is a social media war prevailing in the arena of Twitter.\u003c/p\u003e \u003cp\u003ePolitical discussions on social media have become increasingly prevalent in Pakistani society, with Twitter being a popular platform for individuals to express their opinions and biases towards different political parties and personalities. This research aims to explore how Pakistani Twitter users express their biases and what strategies or tactics they employ to advance their agenda. Understanding these patterns of online behavior can shed light on the current political climate in Pakistan and provide insights into how social media is impacting political discourse in the country. This Corpus based study will help to know the actual use of language in politics in the Pakistani community.\u003c/p\u003e\n\u003ch3\u003eSTATEMENT OF THE PROBLEM\u003c/h3\u003e\n\u003cp\u003eSocial media platforms, particularly Twitter, have become important spaces for political engagement and discussion in Pakistan. However, these platforms often reflect and amplify users\u0026rsquo; political biases, leading to polarized discourse and the spread of agenda-driven narratives. Despite the growing influence of Twitter in shaping public opinion, there is limited research on the nature, frequency, and impact of biased expressions by Pakistani Twitter users toward political parties and personalities. Furthermore, little is known about the strategies and techniques employed by users to promote their political agendas and how these biases influence the overall discourse. Understanding these patterns is essential to foster balanced and constructive political discussions online and to mitigate the reinforcement of partisan biases in digital spaces.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eOBJECTIVES\u003c/h2\u003e \u003cp\u003eThis study seeks to:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eTo examine the linguistic strategies and techniques employed by Pakistani Twitter users to advance their political agendas, including the use of persuasive and argumentative language.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eTo analyze the nature of biased expressions in Pakistani Twitter discourse, focusing on the use of nouns, adjectives, verbs, and other linguistic features that convey favoritism toward specific political parties or personalities.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eRESEARCH QUESTIONS\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eWhat linguistic strategies and techniques do Pakistani Twitter users employ to advance their political agendas?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eHow do Pakistani Twitter users express their biases toward different political parties or personalities through their language use, including nouns, adjectives, verbs, and other lexical features?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"LITERATURE REVIEW","content":"\u003cp\u003eThe use of social media platforms such as Twitter has become increasingly popular in expressing political biases in Pakistan. Many Pakistani Twitter users use the platform to express their political affiliations and opinions on political parties and personalities. Studies have shown that Twitter is a popular platform for political conversations and discourse, particularly during election periods (Khan, Hussain, \u0026amp; Wahab, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eMentions are also used to advance political agendas on Twitter in Pakistan. Users can mention other users on Twitter to draw attention to their message or to challenge them. On Twitter, it is common for political opponents to engage in heated debates and arguments using mentions. For example, during the 2018 general elections in Pakistan, the Twitter accounts of different political parties and personalities engaged in debates using mentions and often engaged in heated debates and arguments (Khan, Hussain, \u0026amp; Wahab, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTwitter users in Pakistan often use tactics such as hashtags, retweets, and mentions to advance their agenda. Hashtags are a common way to promote a particular agenda or viewpoint on Twitter. For example, during the 2018 general elections in Pakistan, supporters of the Pakistan Tehreek-e-Insaf (PTI) used hashtags such as #TabdeeliAaRahiHai (\u0026ldquo;Change is coming\u0026rdquo;) and #NayaPakistan (\u0026ldquo;New Pakistan\u0026rdquo;) to promote their party, while supporters of other parties used hashtags such as #VoteKoIzzatDo (\u0026ldquo;Respect your vote\u0026rdquo;) to criticize the ruling party (Abbasi \u0026amp; Rehman, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eRetweeting is another tactic commonly used to advance political agendas on Twitter in Pakistan. By retweeting a tweet, a user can help amplify a message and spread it to a larger audience. Twitter users in Pakistan often use retweeting to support their political parties or personalities and to criticize their opponents. For example, during the 2021 Senate elections in Pakistan, supporters of the Pakistan Peoples Party (PPP) retweeted a video of a member of the ruling party allegedly buying votes, while the ruling party\u0026rsquo;s supporters retweeted videos alleging that the opposition parties were involved in horse-trading (Butt et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eShahzad Ashraf et al. (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) analyzed the sharing pattern of information among Pakistani politicians on Twitter. The study aimed to identify the sources of information shared by politicians and how these sources impacted political discourse. The study found that politicians engaged in strategic sharing of information, and the type of information shared varied depending on the political motivation of the politician.\u003c/p\u003e \u003cp\u003eFinally, a study conducted by Hassan Mahmood et al. (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) explored media bias and toxicity in South Asian political discourse. The study analyzed news articles and Twitter conversations related to South Asian politics and identified the presence of media bias and toxicity in the discourse. The study concluded that media outlets and politicians played a significant role in shaping the discourse and that efforts need to be made to promote healthy and balanced political communication.\u003c/p\u003e \u003cp\u003eA study conducted by Naeem Ahmed Ibupoto et al. (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) aimed to perform sentiment analysis on current political topics in Pakistan\u0026rsquo;s Twitter user base. The study used machine learning algorithms to analyze tweets related to different political topics and determine the general sentiment expressed by users. The study found that there were significant variations in sentiment on different political topics and the sentiments expressed on Twitter were influenced by the media and the government.\u003c/p\u003e \u003cp\u003eAnother study by Zeeshan Rasheed et al. (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) focused on collecting and analyzing a Twitter dataset of Pakistani political discourse. The study aimed to reveal the types of political topics discussed on Twitter and the main users who engaged in these discussions. The study found that political discussions on Twitter were largely dominated by mainstream politicians and their supporters.\u003c/p\u003e \u003cp\u003eThe use of Twitter for political discourse in Pakistan is an important area of study for researchers. The studies discussed in this literature review highlight the impact of sentiment analysis, information sharing, media bias, and toxicity on the discourse. Some Researches have also shown that Pakistani Twitter users express their biases towards different political parties and personalities using a range of tactics such as hashtags, retweets, and mentions. These tactics are commonly used to advance political agendas and promote political views. This research can be useful for policymakers and stakeholders to promote healthy political communication in the country and to help to know the actual use of language in Pakistani political discussions and debates on twitter.\u003c/p\u003e\n\u003ch3\u003eResearch Gap\u003c/h3\u003e\n\u003cp\u003eWhile previous studies have examined political discourse on Twitter in Pakistan, much of the existing research has primarily focused on sentiment analysis, information sharing, and the use of specific tactics such as hashtags, retweets, and mentions. Although these studies provide insights into general patterns of political engagement, there is limited research specifically investigating the frequency, nature, and strategies of biased expressions by Pakistani Twitter users toward different political parties and personalities. Additionally, few studies have explored the underlying factors that contribute to the formation and reinforcement of these biases or the broader impact of agenda-driven strategies on overall political discourse. Most of the current research emphasizes either individual user behavior or political actors, without offering a comprehensive analysis of how biases influence the structure and tone of online political discussions in Pakistan. Therefore, there is a need for a focused, corpus-based study that examines the linguistic and strategic dimensions of political bias on Pakistani Twitter, providing a deeper understanding of how political agendas are advanced and how discourse is shaped on this platform.\u003c/p\u003e"},{"header":"METHODOLOGY","content":"\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eResearch Design and Data Sampling\u003c/h2\u003e \u003cp\u003eMixed methodology approach is being used here. We have received a Corpus of tweets from the department based on hashtags.We have also selected some hashtag words upon which the corpus This corpus is to be run on the Mat tagger software. Another standard English is also to be run on the Mat Tagger. Values of corpora are to be compared and frequent linguistic features are to be analyzed using some dimensions or theories. Dimension 1 of Biber\u0026rsquo;s six dimensions is used here to analyze the text.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eMat tagger\u003c/h2\u003e \u003cp\u003eMultidimensional Analysis Tagger (MAT) is a software tool designed for text analysis. It is used to tag words and phrases in a text document with various categories such as topic, sentiment, emotion, status, and more. MAT uses advanced natural language processing algorithms and machine learning techniques to extract meaning and context from textual data. It is commonly used in social media monitoring, market research, and customer feedback analysis to identify patterns and trends in user-generated content.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eBiber’s (1988) Dimensions\u003c/h3\u003e\n\u003cp\u003eThe conflict between involved and informational discourse is dimension 1. High scores on this variable suggest that the text is emotive and interactive, like a casual chat, while low levels suggest that the language is informationally rich, like academic prose. In contrast, a low score on this Dimension indicates that the text uses a lot of nouns, big words, and adjectives. A high score on this Dimension indicates that the text uses a lot of verbs and pronouns.\u003c/p\u003e \u003cp\u003eThe conflict between narrative and non-narrative concerns is Dimension 2. High scores on this variable suggest that the text is narrative, such as a novel, while low levels indicate that the material is non-narrative. A high score for this Dimension indicates that, among other things, the text uses third person pronouns and several past tenses.\u003c/p\u003e \u003cp\u003eContext-Independent Discourse and Context-Dependent Discourse are in opposition to one another in Dimension 3. A high score on this variable indicates that the text is not dependent on the context, such as scholarly prose, while a low score indicates that the text is dependent on the context, such as in the case of a sports broadcast. A high score on this Dimension denotes the text\u0026rsquo;s frequent use of nominalizations, whereas a low score denotes the text\u0026rsquo;s frequent use of adverbs, among other qualities.\u003c/p\u003e \u003cp\u003eOvert Expression of Persuasion is measured by Dimension 4. High scores for this variable show that, for example in business letters, the author\u0026rsquo;s point of view and evaluation of likelihood and/or certainty are clearly marked in the text. A high score on this dimension indicates that, among other things, the text uses a lot of modal verbs.\u003c/p\u003e \u003cp\u003eThe conflict between abstract and non-abstract information is dimension five. High scores for this variable suggest that the text presents information in a formal, technical, and abstract manner, like that found, for instance, in scientific speech. A text with a high score on this Dimension contains numerous passive sentences and conjuncts, among other qualities.\u003c/p\u003e \u003cp\u003eMeasurement 6 refers to Online Explanation of Information. High ratings for this characteristic suggest that the text is informative in nature but was created under time constraints, as those found in speeches, for example. A text that performs well on this Dimension exhibits numerous post modifications of noun phrases, among other traits.\u003c/p\u003e\n\u003ch3\u003eDATA ANALYSIS AND INTERPRETATION\u003c/h3\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis 1\u003c/h2\u003e \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e \u003ch2\u003eCorpus Analysis from the Book \u0026ldquo;Variation across speech and writing by Douglas Biber\u003c/h2\u003e \u003cp\u003e \u003cb\u003eDimension 1\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 1. Dimension 1 from Douglas Biber\u0026rsquo;s \u0026ldquo;Variation across speech and writing\u0026rdquo;\u003c/em\u003e \u003cb\u003eAnalysis 2\u003c/b\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eMAT Analysis of corpus based on tweets\u003c/h2\u003e \u003cp\u003e \u003cem\u003eFigure 2. Dimension 1 from Received corpus from the department\u003c/em\u003e \u003c/p\u003e \u003cp\u003eValue score of analysis 1 and analysis 2 is compared to know the diversity of language usage from the standard or the expected usage The corpus received from the department that was based on tweets was analyzed using MAT Tagger software. This Analysis is referred to as \u0026ldquo;Analysis 2\u0026rdquo;. The values of this corpus we are going to compare with the values of Corpus Analysis from Douglas Biber\u0026rsquo;s Book \u0026ldquo;Variation across speech and writing\u0026rdquo;. This corpus Analysis is referred to as \u0026ldquo;analysis 1\u0026rdquo;. Only the values of dimensions 1 are being compared. Here values of the properties of analysis 1 and analysis 2 are being compared here. Both of Analysis are displaying their properties on the X axis and their score on the Y axis. The score of the genre of conversation in analysis 1 and analysis 2 is 36. Value of Broadcast is 3 in analysis 1 and analysis 2. Values of prepared speeches in analysis 1 and analysis 2 are 2 and 1.5 respectively. Value of personal letters in analysis 1 and analysis 2 is 20. Value of general fiction in analysis 1 and analysis 2 is -1. Values of press reportage in analysis 1 and analysis 2 are \u0026minus;\u0026thinsp;8 and \u0026minus;\u0026thinsp;7 respectively. Value of academic prose in analysis 1 and analysis 2 is -8. Value of the official document in analysis 1 and analysis 2 is -18. These are differences between analysis 1 and analysis 2 in terms of scores. The closest genre in analysis 1 and analysis 2 is official documents.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of Fig.\u0026nbsp;1 and Fig.\u0026nbsp;2 in terms of values.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSr no\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGenres\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAnalysis 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAnalysis 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eVariation score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eConversation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBroadcasts\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrepared speeches\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePersonal letters\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGeneral fiction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePress reportage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAcademic prose\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOfficial documents\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eText types or genres found in analysis 1 and analysis 2 show little variations in terms of score of the genres. The main difference is found in only two genres: prepared speeches and press reportage which is 0.5 and \u0026minus;\u0026thinsp;1 respectively. It seems that the language used in the both corpora is identical except prepared speeches and press reportage language. The language of these two genres show a little deviation. The values of dimension 1 in analysis 1 are greater than values of dimension 1 in analysis 2. It shows that analysis 2 shows deviation of language used in Received corpus from the Analysis 1 that is found in Douglas Biber\u0026rsquo;s \u0026ldquo;variation across speech and writing.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMean values of Biber\u0026rsquo;s dimensions found in the Corpus based on the tweets.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDimension 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDimension 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDimension 3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDimension 4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDimension 5\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDimension 6\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eClosest text type\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e-22.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-2.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-3.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-3.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-2.57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eLearned exposition\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eValues of Biber\u0026rsquo;s six dimensions found in the Received corpus based on the tweets data set are displayed in Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. All six dimensions contain different values which suggest the different nature of language used in tweets as political discourse. The dimension we are concerned about is dimension 1 which is The conflict between involved and informational discourse. High scores on this variable suggest that the text is emotive and interactive, like a casual chat, while low levels suggest that the language is informationally rich, like academic prose. In contrast, a low score on this Dimension indicates that the text uses a lot of nouns, big words, and adjectives. A high score on this Dimension indicates that the text uses a lot of verbs and pronouns. So here is the low value of dimension 1 which indicates that the text is informational such as academic prose. Low scores also suggest that text uses lots of nouns, big words and adjectives.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULT AND DISCUSSION","content":" \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003cp\u003eThe question on which the study is conducted is \u0026ldquo; How do Pakistani twitter users express their biases towards different political parties or personalities?What kind of strategies and techniques do Pakistani Twitter users use to advance their political agenda?\u0026rdquo; MAT Analysis of Corpus based on tweets shows that most of the language is used in conversation. The Value of conversation is 36 in Dimension 1 which is conflict between involved and informational discourse. If the values of genres are low then the text is informational and if the values of genres are high then the text is interactional and emotive. To make the conversation interactive and emotive nouns and verbs are used. Closest genres in analysis 1 and analysis 2 of dimension 1 is the official document and learned exposition. Learned exposition, academic prose and official documents serve as informational. Nouns and verbs are used in informational text. Value of conversation in Analysis of dimension 1 in both corpora is 36 which is highest which indicates that text is also emotive and interactive. If we compare the mean value of Biber\u0026rsquo;s six dimesions the mean value of dimensions 1 is the lowest value which is -22.19 that is shown in Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e which indicates that the overall nature of text is informational as compared to the nature of other dimensions of Biber. Academic prose and official documents are considered as informational.\u003c/p\u003e \u003cp\u003eParticipants involved in conversation make the conversation interactive and emotive. Interlocutors in conversation on any platform use persuasive and argumentative strategies or techniques to convince the participants or interlocutors involved in the face to face or online textual conversation. Along with the persuasive and argumentative strategies and techniques Interlocutors use the excess of nouns, pronouns, adjectives, big words and verbs to make the conversation informational, interactive and emotive. So we can say that Pakistani twitter users express their favoritism and biases towards different politicians or parties Using nouns, adjectives ,verbs ,big words such as Imran khan, Azeem leader, change, corruption etc.\u003c/p\u003e \u003cp\u003eIn tweets they indicate the flaws of opposite parties and qualities of their own party. In this regard they use argumentative and persuasive techniques as used by the interlocutors in the talk shows. They make the online conversation on twitter interactive,emotive and informational using nouns, adjectives, big words, and verbs and videos of the different parties as evidence to support or oppose their conversation as well as parties. MAT analysis shows that the actual usage of language used on twitter platform shows little deviation from the expected usage of language as used according to standard norms of language but their language is organized according to their usage in politics to support or defend the parties or the agendas.\u003c/p\u003e \u003c/div\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eExploring political biases in Pakistani twitter data set is the topic of the study. The questions on which we have conducted the study are \u0026ldquo; How do Pakistani twitter users express their biases towards different political parties or personalities?What kind of strategies and techniques do Pakistani Twitter users use to advance their political agenda?\u0026rdquo;. Mixed methodology is used in the study. Analysis 1 from the Douglas Biber\u0026rsquo;s \u0026ldquo;variation across speech and writing\u0026rdquo; and analysis 2 of the Corpus received from the department are compared in terms of their scores. Data is analyzed and interpreted using dimensions 1 of Biber\u0026rsquo;s six dimensions. Scores of genres in dimensions of analysis 2 shows little bit variation from the analysis 1. Mean value of dimension 1 of the analysis 2 is -22.19. Scores of genres in dimension 1 of analysis 2 shows that text is interactional, emotive and informational. Value of dimension 1 is the lowest value which indicates that the overall nature of the text is informational. Participants of the conversation on twitter use nouns, adjectives, verbs and big words to advance political agenda and express their biases towards political parties .they also use persuasive and argumentative strategies to do so.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eNoor Zaman Khan conceptualized the study, designed the research framework, and collected the data. The author conducted the corpus-based analysis using MAT Tagger software and applied Biber\u0026rsquo;s multidimensional analysis. Noor Zaman Khan performed the interpretation of results, drafted the manuscript, and revised it critically for important intellectual content. The author approves the final version of the manuscript and agrees to be accountable for all aspects of the work, ensuring accuracy and integrity.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAbbasi, J. A., \u0026amp; Rehman, N. U. (2020). Tweeting for political change: The impact of hashtags during the 2018 Pakistan general elections. \u003cem\u003eTelematics and Informatics\u003c/em\u003e, \u003cem\u003e49\u003c/em\u003e, 101382. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.tele.2020.101382\u003c/span\u003e\u003cspan address=\"10.1016/j.tele.2020.101382\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAshraf, S., Khamis, S., \u0026amp; Weldon, L. (2021). Twitter dialogue: An analysis of Pakistani politicians\u0026rsquo; information sharing. \u003cem\u003eSocial Media\u0026thinsp;+\u0026thinsp;Society\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(1), 20563051211003620. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/20563051211003620\u003c/span\u003e\u003cspan address=\"10.1177/20563051211003620\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eButt, S. A., Aslam, S., \u0026amp; Asghar, T. (2021). Use of social media for political discourse in Pakistan: Evidence from 2021 Senate elections. \u003cem\u003eTechTrends\u003c/em\u003e, \u003cem\u003e65\u003c/em\u003e, 1\u0026ndash;10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11528-021-00619-7\u003c/span\u003e\u003cspan address=\"10.1007/s11528-021-00619-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIbupoto, N. A., Rasheed, Z., Ashraf, S., Ali, M., Tanweer, R., \u0026amp; Sen, S. (2022). \u003cem\u003eSentiment analysis on current political topics in Pakistan\u0026rsquo;s Twitter user bases\u003c/em\u003e (pp. 1\u0026ndash;21). Preprint. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.21203/rs.3.rs-2095172/v1\u003c/span\u003e\u003cspan address=\"10.21203/rs.3.rs-2095172/v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan, M. S., Hussain, A., \u0026amp; Wahab, F. (2018). Political discourse on social media: A case study of Twitter use in Pakistan\u0026rsquo;s 2018 general elections. \u003cem\u003eJournal of Asian Pacific Communication\u003c/em\u003e, \u003cem\u003e28\u003c/em\u003e(2), 222\u0026ndash;241. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1075/japc.28.2.05kha\u003c/span\u003e\u003cspan address=\"10.1075/japc.28.2.05kha\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMahmood, H., Ayub, A., \u0026amp; Jabeen, U. (2021). Exploring media bias and toxicity in South Asian political discourse. \u003cem\u003eNew Media \u0026amp; Society\u003c/em\u003e, \u003cem\u003e23\u003c/em\u003e(8), 2317\u0026ndash;2336. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/1461444820915995\u003c/span\u003e\u003cspan address=\"10.1177/1461444820915995\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRasheed, Z., Ibupoto, N. A., Ashraf, S., Ali, M., \u0026amp; Sen, S. (2022). A Twitter dataset for Pakistani political discourse. \u003cem\u003eData in Brief\u003c/em\u003e, 38, 107277. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.dib.2022.107277\u003c/span\u003e\u003cspan address=\"10.1016/j.dib.2022.107277\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Twitter, agenda, variation, participants, argumentative, persuasive","lastPublishedDoi":"10.21203/rs.3.rs-8325825/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8325825/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eStudy is conducted on Exploring political biases in Pakistani twitter data set. The objectives of the study is the way Pakistani twitter users express their biases towards different political parties or personalities and the way they advance their political agenda. Mixed methodology approach is used in the study. Biber\u0026rsquo;s six dimensions, MAT tagger and Douglas Biber\u0026rsquo;s \u0026ldquo;variation across speech and writing\u0026rdquo; is utilized in comparing the analysis 1 and analysis 2. Analysis 2 based on the Corpus received from the department is identical to Analysis 1 from the Douglas Biber\u0026rsquo;s \u0026ldquo; variation across speech and writing\u0026rdquo; except few genres which indicates little variations from the standard or expected usage of language. Result of the study shows that participants involved in conversation use nouns, pronouns, adjectives, verbs and long words to express their biases towards different political parties or personalities, and to advance their political agenda. They use the argumentative and persuasive techniques to convince the interlocutors on the twitter.\u003c/p\u003e","manuscriptTitle":"Exploring the Political Biases in Pakistani Tweets Dataset","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-22 09:14:04","doi":"10.21203/rs.3.rs-8325825/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2006e024-a0dd-4e71-89f8-49d41072d038","owner":[],"postedDate":"December 22nd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-01-09T08:25:05+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-22 09:14:04","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8325825","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8325825","identity":"rs-8325825","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00