Modeling human cognition and polarization using reinforcement Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Modeling human cognition and polarization using reinforcement Learning P. Pavankumar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3942048/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Some people polarize due to heated arguments but in reality and research studies says that they even polarize with similar information also. People think irrationally and confirms their pre-existing beliefs and strive to look into that information only. With these they avoid such data that contradicts them. Polarization occurs through the bias contribute to the reinforcement of existing beliefs. The people manipulate the new evidence based on the previous beliefs and bias confirmation suggestions. In order to form beliefs, people depend on least samples rather than gathering huge data and processing is also costly. A hypothesis is generated based on the least samples collected and then belief polarization model is used to test our hypothesis. A new belief formation model is developed based on reinforcement learning. In accordance to this a new evaluation metrics is used for polarization and simulations are learned how least size affects polarization. In this work basic assumptions are done with the generated hypothesis results, such that with these least samples polarization may increase. Practical suggestions are done for obtaining polarization based on our findings. Belief Polarization Machine Learning Reinforcement Learning etc Figures Figure 1 Figure 2 Figure 3 1. Introduction Indian people are scared that whether they would get the coronavirus vaccine if it is available since 2021. Such controversy questions and beliefs arises regarding vaccines brings unfortunate results and leads to false outcomes, which results them to miss their life-saving treatments. This thought of belief from people having violent behavior and conflicts about vaccination is common in the society. Are humans responsible for natural calamities. Will caste reservation alone solve the problems of right to equality and reduces the crime. Will the new tax regime and tax slabs improve the Indian economy? People often disagree sharply on consequential issues that guide our lives. Polarization is considered as very dangerous when one category of people totally decline with ideology of other set of peoples. Researchers has come to conclusion that political polarization has more probability in increasing the risk of violent conflicts within and between neighboring countries. However, beliefs are framed based on some small and least samples as gathering the huge data or information is costly. People may be confused with the results from small or least samples as these results itself cannot make a conclusion. For instance, consider people who browse news articles online. Within a short period of time, they can be exposed to hundreds of issues from many topics. However, there is a limit to how much information people can process as reading and understanding articles take time and effort. Therefore, people may only read the headlines or just skim through the articles. Even when people read the articles, they are unlikely to research further and look for related articles. People come to conclusion based on only few samples, thinking that gathering huge data or information is costly. We hypothesize that reliance on smaller samples leads to greater polarization. We discuss people’s tendency to rely on small samples, the theories on why people rely on small samples, and how it translates to belief polarization. We discuss prior models, formalize the definition of polarization, and design our model. We explore a novel approach to modeling belief polarization using machine learning. In this work is the first to provide a machine-learning framework for empirical analysis of belief polarization. 2. Literature Review Increasing polarization increases the risk of conflict, this is the major accepted guessing hypothesis in the society which may leads to violence. Although balance-of-power theorists are still influential, the growing available evidence shows that power parity is rather associated with an increased risk of conflict. There are strong reasons to believe that polarization also reduces the willingness of humans to contribute to common causes, reducing the provision of public goods and hampering the economic prospects of developing countries. Increased divisions within a society may create damage to democracy and may weaken the social progress, we have only just started to understand these processes and their impact upon society [ 1 ]. However, divergent information does not fully explain why people polarize. There is evidence that people are reasonably good at detecting misinformation [ 2 ] and even echo chambers are not sealed off from unbiased information. More importantly, empirical studies indicate that people can polarize even when they consider similar information. clearly show that human learning can be affected by the addition of constants to the payoffs. It seems that the distinction between gains and losses affects learning. Whereas the quantification proposed by Erev and Roth provides a good fit for the current data and to previous probability learning and signal detection experiments, it is not likely to be generally accurate. It seems that in some settings behavior is better approximated under the assumption of a reference point that is initially (or quickly adjusted to be) larger than the worst possible outcome. [ 3 ]. Researchers often attribute this to cognitive biases that influence how people process new evidence. For example, confirmation bias suggests people interpret information in a manner that validates their prior beliefs. However, this explanation is also incomplete as it does not address why people have contrasting priors in the first place. Probability matching describes a decision strategy in which predictions of class membership are proportional to the class probability. For instance, one study [ 4 ] asked the participants to guess which color would appear on the screen next. If they made a correct guess, they would gain some amount of reward, and if they made an incorrect guess, they would lose the same amount. Suppose the color red appears with a probability of 0.7 and the color blue appears with a probability of 0.3. The optimal strategy would be to predict red in every single trial. However, participants generally exhibited probability matching. Researchers suggest this can be explained by people’s reliance on small samples. For instance, if people only considered the past four examples and predicted the common color as the next color, they would predict red and blue with probabilities close to 0.7 and 0.3, respectively. In many models of cognition, researchers assume acquiring and processing samples are costly. Here, the cost may represent time, metabolic energy, memory, or even effort. Modeling human cognition is challenging because there are infinitely many mechanisms that can generate any given observation. Some researchers address this by constraining the hypothesis space through assumptions about what the human mind can and cannot do, while others constrain it through principles of rationality and adaptation [ 5 ]. Large samples are necessary to accurately approximate the probability distribution. However, samples are costly. Therefore, people face a trade-off between accuracy and sample size. Let us consider choosing a route to avoid traffic. How many possible arrangements of traffic across the city should one consider before deciding whether to turn left or right at the next intersection? Clearly, one should not pause at the intersection for hours to consider all the possibilities. Yet, one should also not make a random decision without any consideration. The researchers suggest that in such setting where people face a trade-off, making many quick but locally sub-optimal decisions based on a few samples may be a globally optimal strategy over long periods. Similarly, we hypothesize that people form many beliefs based on a few samples, which may be locally sub-optimal. In many learning or inference tasks human behavior approximates that of Bayesian ideal observer, suggesting that, at some level, cognition can be described as Bayesian inference. However, a number of findings have highlighted an intriguing mismatch between human behavior and standard assumptions about optimality: People often appear to make decisions based on just one or a few samples from the appropriate posterior probability distribution, rather than using the full distribution. Although sampling-based approximations are a common way to implement Bayesian inference, the very limited numbers of samples often used by humans seem insufficient to approximate the required probability distributions very accurately. For one particular belief, however, people’s beliefs may be inaccurate and polarize with others beliefs [ 6 ]. Many studies have explored content-selection algorithms to recommend samples that accurately represent the population [ 7 ]. The second challenge is exposing people to more samples. Some people depend on least samples due to their cost, so it suggest methods of delivering information in a cost-efficient manner. For instance, consider people browsing news articles online. Typically, people only read one or two articles regarding some topic before moving on to the next topic. That is, people are unlikely to research further and look for various perspectives. It is represented that if people are presented with multiple headlines on the same topic from diverse sources, people may be more inclined to consider various perspectives. In addition, need to provide readers with well-designed contents or information of the topic that make the key points easier to understand [ 8 ]. Assuming people have a fixed amount of time and energy to consider information, making the information cost-efficient to understand may be a promising strategy. 3. Methodology Let us discuss how polarization can be measured and start initiating the problem by defining the task and learning model. For example, regarding the safety of newly developed vaccine there are some beliefs formed by medical practitioners which are actually learning agents. Beliefs are formed by the two medical practitioners based on the treatment outcomes that are observed from the set of patients about the vaccine. So, the medical practitioners who are observing the set of patients are known as the training set. For a particular training sample (x,y), where x is a vector of values that describe the treatment such as the patient’s age, the patient’s blood pressure, the dosage of the vaccine, and so on. Let patient’s response to the vaccine is represented as label y. similarly a large y indicates that the treatment was effective and safe. After believing themselves the medical practitioners are given with some test set of new patients which are identical and are asked them to find prediction about vaccine and how will they are responded. We measure polarization based on their predictions on the test set. We propose three different ways of measuring polarization: variance, continuous disagreement, and discrete disagreement. In machine learning, bias-variance decomposition is a way of analyzing a learning agent’s expected error with respect to a particular problem as a sum of three terms: bias, variance, and irreducible noise. Expected Error = bias 2 + variance + noise. The squared error, err(h) = (h(x) − f(x))2, can be decomposed into bias and variance as follows: Bias 2 = [E[(h(x)] − f(x)] 2 variance = E[[h(x) − E[h(x)]] 2 ] The bias represents the extent to which the average prediction over all data sets differs from the target function, and the variance measures the extent to which the predictions for individual data sets vary around their average. With respect to the formation of beliefs, the variance measures how far the agent’s beliefs are from each other and the bias measures how far the agent’s beliefs are from the objective truth. Hence the spread of the beliefs are known as the measure of variance and finally conclude that the larger variance results greater polarization. 4. Experimental Results Multiple groups of learning agents are created based on the no. of training samples that are observed. a few least sizes are considered based on the multiple groups only. Then one group observes 3 samples from 80 agents whereas other group observe 4 samples from 80 agents and this process repeats simultaneously. Now prediction is done from 80 agents which consists of random test samples from the target distribution. This is done after agents learn all the belief functions. Based on the predictions on the test set a new polarization is measured. Let us visualize what a high polarization looks like compared to a low polarization before calculating the different measures of polarization,. From Fig. 1 let us consider a large polarization between two agents on a test set for visualization purpose. Based on target values for visualization consider the first column of the graph which shows 80 test samples. The predictions are done through 2 training samples are looked by two agents are shown in two columns to the right. On the test samples there is a clear difference between the two agents’ predictions. For example, with respect to large samples agent A predicts positive values whereas agent B predicts negative values. Hence, one medical practitioner believes vaccination will be risky for same set of patients. while the other medical practitioner believes vaccination will be safe. From Fig. 2 the predictions are made from more training samples by two agents. The two agent’s predictions are more identical to each other and to the target when the agents observe a huge number of samples,. From Fig. 3 polarization is measured using the variance that shows the expected bias and the expected variance for agents who observe a different number of training samples. The final expected variance decreases simultaneously with respect to the result as agents observe more samples. This suggests that when two people observe more information, their beliefs tend to be more aligned with each other. For example, two medical practitioners who have observed more previous patients are expected to make more similar predictions about new patients. The figure also shows that when agents observe a sufficient number of samples, the expected variance is minimized and additional samples do not reduce the variance. This suggests that once people observe enough information, their beliefs converge. Similarly, the expected bias initially decreases as agents observe more samples. However, the expected bias is notably smaller and plateaus faster than the expected variance. In our model, the variance contributes much more to the total error than the bias when agents rely on small samples. Note that the reason the agents have low bias is that linear agents are suited for solving linear regression tasks. Meaning, our model describes a setting in which people are learning about a topic that they can fully comprehend once they observe enough information. Suppose our agents are given many tasks and have to learn many target distributions, but samples are costly. If the agent’s goal is to be as accurate as possible across many tasks, but samples are costly, they may choose to only observe a few samples. This is because few samples are sufficient to minimize the expected bias. However, small samples lead to high variance, which results in greater polarization. 5. Conclusion To test this hypothesis, we explored a new method of modeling polarization based on an empirical analysis of machine learning. After designing a machine learning framework that emulates human learning, we used simulations to examine the impact of small sample reliance on belief polarization. Our simulation results confirmed our hypothesis that small sample reliance leads to greater polarization. The results also indicated that complex problems, high levels of noise, and low cognitive efforts also lead to greater polarization. Based on our findings, we offer possible suggestions for mitigating polarization in the real world. Our simulations indicate that polarization decreases when agents observe more samples, assuming that people observe information from similar sources. However, we also found that when the problem is too complex, has high noise, or is difficult to process, increasing the sample size only reduces polarization marginally. Therefore, we suggest exposing people to more high-quality samples that can be easily digested at little cost. The first challenge is making sure that the information people observe is of high quality. Given the countless configurations involved in the study of belief formation and the social phenomenon of polarization, we believe a flexible framework like machine learning is appropriate and can provide valuable insights. For instance, our work only explores one of many possible ways in which we can instantiate the learning agent and the learning task. Based on the knowledge that human decisions tend to exhibit linear properties, we expressed the learning agent as a linear model. However, one may explore a different, perhaps non-linear, learning algorithm to depict how humans learn. Similarly, based on the knowledge that many relationships in the world are linear, we expressed the task as a linear problem. However, one may design a unique machine learning task depending on the focus of their study. In addition, our model assumed a supervised learning framework. However, different frameworks such as unsupervised learning, reinforcement learning, or even deep learning may serve as an appropriate expression of how humans learn. As the field of machine learning grows, how we can utilize machine learning to better understand cognition will be a topic of great interest. Declarations Author Contribution study conception and design: Dr.P.Pavankumardata collection: Dr.P.Pavankumaranalysis and interpretation of results: Dr.P.Pavankumardraft manuscript preparation: Dr.P.Pavankumar References Joan Esteban and Gerald Schneider: Polarization and conflict: Theoretical and empirical issues: Introduction. J. Peace Res. 45 (2), 131–141 (2008) David Lazer: The rise of the social algorithm. Science. 348 (6239), 1090–1091 (2015) Mathew, D., Hardy, Bill, D., Thompson, P.M., Krafft, Thomas, L., Griffiths: Bias amplification in experimental social networks is reduced by resampling. arXiv preprint arXiv:2208.07261 Yoella Bereby-Meyer and Ido Erev. On learning to become a successful loser: A comparison of alternative abstractions of learning processes in the loss domain. Journal of mathematical psychology , 42(2–3):266–286, 1998. Falk Lieder and Thomas L Griffiths, “Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources”, Behavioral and brain sciences , 43, 2020. Edward Vul, Noah Goodman, Thomas L Griffiths, and Joshua B Tenenbaum, “One and done? optimal decisions from very few samples”, Cognitive science , 38(4):599–637, 2014. Bhattacharya, Soham, Monoranjan Guchait, and Aravind H. Vijay. "Boosted top quark tagging and polarization measurement using machine learning." Physical Review D 105.4 (2022): 042005. Gentzkow, Matthew, Jesse Shapiro, and Matt Taddy. Measuring polarization in high- dimensional data: Method and application to congressional speech . No. id: 11114. 2016. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3942048","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":272292664,"identity":"d167c76c-fc9b-46ed-a576-261554bd7290","order_by":0,"name":"P. Pavankumar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABFklEQVRIiWNgGAWjYNCCAzBGgQUDP4hOKCBai4EEg2QDSIsBKVoMwBw8WszZzxh+LjhzOJp/9gG2Bz8MJOSNz69O/PDAgEGeX+wAVi2WPTnG0jNuHM6dcS6B3bDHQMJw2423myWADjOcOTsBqxaDA2kJ0jwfDuc2nGFgk+AxkGDcduPsBpCWBIPbOLScf5b8G6RlPlCL5B8DCfvNM85u/oFXy43kY9I8QIdtAGqRBtqSuIG/dxteWyxnPD5mzXMmPXfjGcZ2YxkDieQZN3i3WSQAPYXLL+b8ic23eY5Z5847w3zs4ZsKG9v+/rObb/6osJHnl8bhMASTsQ1CS4BVSmBVjqaFgQ1C8R/AqXoUjIJRMApGJgAA2SRiYoDdDAMAAAAASUVORK5CYII=","orcid":"","institution":"vardhaman college of engineering","correspondingAuthor":true,"prefix":"","firstName":"P.","middleName":"","lastName":"Pavankumar","suffix":""}],"badges":[],"createdAt":"2024-02-09 05:29:25","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-3942048/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3942048/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":51015959,"identity":"e20c8ef9-b651-4929-8485-85a551522eb1","added_by":"auto","created_at":"2024-02-12 18:52:12","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":18771,"visible":true,"origin":"","legend":"\u003cp\u003ePredictions through test set by Two Agents are restricted to 2 training samples\u003c/p\u003e","description":"","filename":"figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3942048/v1/e4dbe71f5be28c0e3943268d.jpg"},{"id":51015961,"identity":"9a8bed70-07c9-410f-a533-067cf53aeeef","added_by":"auto","created_at":"2024-02-12 18:52:12","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":18803,"visible":true,"origin":"","legend":"\u003cp\u003ePredictions through test set by Two Agents are restricted to 8 training samples\u003c/p\u003e","description":"","filename":"figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3942048/v1/73545ceb107a3d9a2ad55ffb.jpg"},{"id":51015960,"identity":"e16d99ac-8ba5-42bc-bd68-4a6b06fff5d2","added_by":"auto","created_at":"2024-02-12 18:52:12","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":24678,"visible":true,"origin":"","legend":"\u003cp\u003eExpected Bias-Variance vs. Sample Size\u003c/p\u003e","description":"","filename":"figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3942048/v1/eb4bacc7cde0ea445c45e8fc.jpg"},{"id":58284354,"identity":"bc01f972-d0cd-4ebf-a5dd-3c4ce2697930","added_by":"auto","created_at":"2024-06-13 11:47:25","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":255756,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3942048/v1/f753bb2c-5b3b-40b2-bf2a-4c246d032c6a.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Modeling human cognition and polarization using reinforcement Learning","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eIndian people are scared that whether they would get the coronavirus vaccine if it is available since 2021. Such controversy questions and beliefs arises regarding vaccines brings unfortunate results and leads to false outcomes, which results them to miss their life-saving treatments. This thought of belief from people having violent behavior and conflicts about vaccination is common in the society. Are humans responsible for natural calamities. Will caste reservation alone solve the problems of right to equality and reduces the crime. Will the new tax regime and tax slabs improve the Indian economy? People often disagree sharply on consequential issues that guide our lives. Polarization is considered as very dangerous when one category of people totally decline with ideology of other set of peoples. Researchers has come to conclusion that political polarization has more probability in increasing the risk of violent conflicts within and between neighboring countries.\u003c/p\u003e \u003cp\u003eHowever, beliefs are framed based on some small and least samples as gathering the huge data or information is costly. People may be confused with the results from small or least samples as these results itself cannot make a conclusion. For instance, consider people who browse news articles online. Within a short period of time, they can be exposed to hundreds of issues from many topics. However, there is a limit to how much information people can process as reading and understanding articles take time and effort. Therefore, people may only read the headlines or just skim through the articles. Even when people read the articles, they are unlikely to research further and look for related articles. People come to conclusion based on only few samples, thinking that gathering huge data or information is costly. We hypothesize that reliance on smaller samples leads to greater polarization.\u003c/p\u003e \u003cp\u003eWe discuss people\u0026rsquo;s tendency to rely on small samples, the theories on why people rely on small samples, and how it translates to belief polarization. We discuss prior models, formalize the definition of polarization, and design our model. We explore a novel approach to modeling belief polarization using machine learning. In this work is the first to provide a machine-learning framework for empirical analysis of belief polarization.\u003c/p\u003e"},{"header":"2. Literature Review","content":"\u003cp\u003eIncreasing polarization increases the risk of conflict, this is the major accepted guessing hypothesis in the society which may leads to violence. Although balance-of-power theorists are still influential, the growing available evidence shows that power parity is rather associated with an increased risk of conflict. There are strong reasons to believe that polarization also reduces the willingness of humans to contribute to common causes, reducing the provision of public goods and hampering the economic prospects of developing countries. Increased divisions within a society may create damage to democracy and may weaken the social progress, we have only just started to understand these processes and their impact upon society [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eHowever, divergent information does not fully explain why people polarize. There is evidence that people are reasonably good at detecting misinformation [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] and even echo chambers are not sealed off from unbiased information. More importantly, empirical studies indicate that people can polarize even when they consider similar information. clearly show that human learning can be affected by the addition of constants to the payoffs. It seems that the distinction between gains and losses affects learning. Whereas the quantification proposed by Erev and Roth provides a good fit for the current data and to previous probability learning and signal detection experiments, it is not likely to be generally accurate. It seems that in some settings behavior is better approximated under the assumption of a reference point that is initially (or quickly adjusted to be) larger than the worst possible outcome. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Researchers often attribute this to cognitive biases that influence how people process new evidence. For example, confirmation bias suggests people interpret information in a manner that validates their prior beliefs. However, this explanation is also incomplete as it does not address why people have contrasting priors in the first place.\u003c/p\u003e \u003cp\u003eProbability matching describes a decision strategy in which predictions of class membership are proportional to the class probability. For instance, one study [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e4\u003c/span\u003e] asked the participants to guess which color would appear on the screen next. If they made a correct guess, they would gain some amount of reward, and if they made an incorrect guess, they would lose the same amount. Suppose the color red appears with a probability of 0.7 and the color blue appears with a probability of 0.3. The optimal strategy would be to predict red in every single trial. However, participants generally exhibited probability matching. Researchers suggest this can be explained by people\u0026rsquo;s reliance on small samples. For instance, if people only considered the past four examples and predicted the common color as the next color, they would predict red and blue with probabilities close to 0.7 and 0.3, respectively.\u003c/p\u003e \u003cp\u003eIn many models of cognition, researchers assume acquiring and processing samples are costly. Here, the cost may represent time, metabolic energy, memory, or even effort. Modeling human cognition is challenging because there are infinitely many mechanisms that can generate any given observation. Some researchers address this by constraining the hypothesis space through assumptions about what the human mind can and cannot do, while others constrain it through principles of rationality and adaptation [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Large samples are necessary to accurately approximate the probability distribution. However, samples are costly. Therefore, people face a trade-off between accuracy and sample size.\u003c/p\u003e \u003cp\u003eLet us consider choosing a route to avoid traffic. How many possible arrangements of traffic across the city should one consider before deciding whether to turn left or right at the next intersection? Clearly, one should not pause at the intersection for hours to consider all the possibilities. Yet, one should also not make a random decision without any consideration. The researchers suggest that in such setting where people face a trade-off, making many quick but locally sub-optimal decisions based on a few samples may be a globally optimal strategy over long periods. Similarly, we hypothesize that people form many beliefs based on a few samples, which may be locally sub-optimal. In many learning or inference tasks human behavior approximates that of Bayesian ideal observer, suggesting that, at some level, cognition can be described as Bayesian inference. However, a number of findings have highlighted an intriguing mismatch between human behavior and standard assumptions about optimality: People often appear to make decisions based on just one or a few samples from the appropriate posterior probability distribution, rather than using the full distribution. Although sampling-based approximations are a common way to implement Bayesian inference, the very limited numbers of samples often used by humans seem insufficient to approximate the required probability distributions very accurately. For one particular belief, however, people\u0026rsquo;s beliefs may be inaccurate and polarize with others beliefs [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMany studies have explored content-selection algorithms to recommend samples that accurately represent the population [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. The second challenge is exposing people to more samples. Some people depend on least samples due to their cost, so it suggest methods of delivering information in a cost-efficient manner. For instance, consider people browsing news articles online. Typically, people only read one or two articles regarding some topic before moving on to the next topic. That is, people are unlikely to research further and look for various perspectives. It is represented that if people are presented with multiple headlines on the same topic from diverse sources, people may be more inclined to consider various perspectives. In addition, need to provide readers with well-designed contents or information of the topic that make the key points easier to understand [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Assuming people have a fixed amount of time and energy to consider information, making the information cost-efficient to understand may be a promising strategy.\u003c/p\u003e"},{"header":"3. Methodology","content":"\u003cp\u003eLet us discuss how polarization can be measured and start initiating the problem by defining the task and learning model. For example, regarding the safety of newly developed vaccine there are some beliefs formed by medical practitioners which are actually learning agents. Beliefs are formed by the two medical practitioners based on the treatment outcomes that are observed from the set of patients about the vaccine. So, the medical practitioners who are observing the set of patients are known as the training set. For a particular training sample (x,y), where x is a vector of values that describe the treatment such as the patient\u0026rsquo;s age, the patient\u0026rsquo;s blood pressure, the dosage of the vaccine, and so on. Let patient\u0026rsquo;s response to the vaccine is represented as label y. similarly a large y indicates that the treatment was effective and safe.\u003c/p\u003e \u003cp\u003eAfter believing themselves the medical practitioners are given with some test set of new patients which are identical and are asked them to find prediction about vaccine and how will they are responded. We measure polarization based on their predictions on the test set. We propose three different ways of measuring polarization: variance, continuous disagreement, and discrete disagreement.\u003c/p\u003e \u003cp\u003eIn machine learning, bias-variance decomposition is a way of analyzing a learning agent\u0026rsquo;s expected error with respect to a particular problem as a sum of three terms: bias, variance, and irreducible noise.\u003c/p\u003e \u003cp\u003eExpected Error\u0026thinsp;=\u0026thinsp;bias\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;+\u0026thinsp;variance\u0026thinsp;+\u0026thinsp;noise.\u003c/p\u003e \u003cp\u003eThe squared error, err(h) = (h(x)\u0026thinsp;\u0026minus;\u0026thinsp;f(x))2, can be decomposed into bias and variance as follows:\u003c/p\u003e \u003cp\u003eBias\u003csup\u003e2\u003c/sup\u003e = [E[(h(x)]\u0026thinsp;\u0026minus;\u0026thinsp;f(x)]\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003cp\u003evariance\u0026thinsp;=\u0026thinsp;E[[h(x)\u0026thinsp;\u0026minus;\u0026thinsp;E[h(x)]]\u003csup\u003e2\u003c/sup\u003e]\u003c/p\u003e \u003cp\u003eThe bias represents the extent to which the average prediction over all data sets differs from the target function, and the variance measures the extent to which the predictions for individual data sets vary around their average. With respect to the formation of beliefs, the variance measures how far the agent\u0026rsquo;s beliefs are from each other and the bias measures how far the agent\u0026rsquo;s beliefs are from the objective truth. Hence the spread of the beliefs are known as the measure of variance and finally conclude that the larger variance results greater polarization.\u003c/p\u003e"},{"header":"4. Experimental Results","content":"\u003cp\u003eMultiple groups of learning agents are created based on the no. of training samples that are observed. a few least sizes are considered based on the multiple groups only. Then one group observes 3 samples from 80 agents whereas other group observe 4 samples from 80 agents and this process repeats simultaneously. Now prediction is done from 80 agents which consists of random test samples from the target distribution. This is done after agents learn all the belief functions. Based on the predictions on the test set a new polarization is measured. Let us visualize what a high polarization looks like compared to a low polarization before calculating the different measures of polarization,.\u003c/p\u003e \u003cp\u003eFrom Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e let us consider a large polarization between two agents on a test set for visualization purpose. Based on target values for visualization consider the first column of the graph which shows 80 test samples. The predictions are done through 2 training samples are looked by two agents are shown in two columns to the right. On the test samples there is a clear difference between the two agents\u0026rsquo; predictions. For example, with respect to large samples agent A predicts positive values whereas agent B predicts negative values. Hence, one medical practitioner believes vaccination will be risky for same set of patients. while the other medical practitioner believes vaccination will be safe.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFrom Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e the predictions are made from more training samples by two agents. The two agent\u0026rsquo;s predictions are more identical to each other and to the target when the agents observe a huge number of samples,.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFrom Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e polarization is measured using the variance that shows the expected bias and the expected variance for agents who observe a different number of training samples. The final expected variance decreases simultaneously with respect to the result as agents observe more samples. This suggests that when two people observe more information, their beliefs tend to be more aligned with each other. For example, two medical practitioners who have observed more previous patients are expected to make more similar predictions about new patients. The figure also shows that when agents observe a sufficient number of samples, the expected variance is minimized and additional samples do not reduce the variance. This suggests that once people observe enough information, their beliefs converge.\u003c/p\u003e \u003cp\u003eSimilarly, the expected bias initially decreases as agents observe more samples. However, the expected bias is notably smaller and plateaus faster than the expected variance. In our model, the variance contributes much more to the total error than the bias when agents rely on small samples. Note that the reason the agents have low bias is that linear agents are suited for solving linear regression tasks. Meaning, our model describes a setting in which people are learning about a topic that they can fully comprehend once they observe enough information. Suppose our agents are given many tasks and have to learn many target distributions, but samples are costly. If the agent\u0026rsquo;s goal is to be as accurate as possible across many tasks, but samples are costly, they may choose to only observe a few samples. This is because few samples are sufficient to minimize the expected bias. However, small samples lead to high variance, which results in greater polarization.\u003c/p\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eTo test this hypothesis, we explored a new method of modeling polarization based on an empirical analysis of machine learning. After designing a machine learning framework that emulates human learning, we used simulations to examine the impact of small sample reliance on belief polarization. Our simulation results confirmed our hypothesis that small sample reliance leads to greater polarization. The results also indicated that complex problems, high levels of noise, and low cognitive efforts also lead to greater polarization.\u003c/p\u003e \u003cp\u003eBased on our findings, we offer possible suggestions for mitigating polarization in the real world. Our simulations indicate that polarization decreases when agents observe more samples, assuming that people observe information from similar sources. However, we also found that when the problem is too complex, has high noise, or is difficult to process, increasing the sample size only reduces polarization marginally. Therefore, we suggest exposing people to more high-quality samples that can be easily digested at little cost. The first challenge is making sure that the information people observe is of high quality.\u003c/p\u003e \u003cp\u003eGiven the countless configurations involved in the study of belief formation and the social phenomenon of polarization, we believe a flexible framework like machine learning is appropriate and can provide valuable insights. For instance, our work only explores one of many possible ways in which we can instantiate the learning agent and the learning task. Based on the knowledge that human decisions tend to exhibit linear properties, we expressed the learning agent as a linear model. However, one may explore a different, perhaps non-linear, learning algorithm to depict how humans learn. Similarly, based on the knowledge that many relationships in the world are linear, we expressed the task as a linear problem. However, one may design a unique machine learning task depending on the focus of their study. In addition, our model assumed a supervised learning framework. However, different frameworks such as unsupervised learning, reinforcement learning, or even deep learning may serve as an appropriate expression of how humans learn. As the field of machine learning grows, how we can utilize machine learning to better understand cognition will be a topic of great interest.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003estudy conception and design: Dr.P.Pavankumardata collection: Dr.P.Pavankumaranalysis and interpretation of results: Dr.P.Pavankumardraft manuscript preparation: Dr.P.Pavankumar\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003e\u003cspan\u003eJoan Esteban and Gerald Schneider: Polarization and conflict: Theoretical and empirical issues: Introduction. J. Peace Res. \u003cstrong\u003e45\u003c/strong\u003e(2), 131\u0026ndash;141 (2008)\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eDavid Lazer: The rise of the social algorithm. Science. \u003cstrong\u003e348\u003c/strong\u003e(6239), 1090\u0026ndash;1091 (2015)\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMathew, D., Hardy, Bill, D., Thompson, P.M., Krafft, Thomas, L., Griffiths: Bias amplification in experimental social networks is reduced by resampling. \u003cem\u003earXiv preprint arXiv:2208.07261\u003c/em\u003e\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eYoella Bereby-Meyer and Ido Erev. On learning to become a successful loser: A comparison of alternative abstractions of learning processes in the loss domain. \u003cem\u003eJournal of mathematical psychology\u003c/em\u003e, 42(2\u0026ndash;3):266\u0026ndash;286, 1998.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eFalk Lieder and Thomas L Griffiths, \u0026ldquo;Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources\u0026rdquo;, \u003cem\u003eBehavioral and brain sciences\u003c/em\u003e, 43, 2020.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eEdward Vul, Noah Goodman, Thomas L Griffiths, and Joshua B Tenenbaum, \u0026ldquo;One and done? optimal decisions from very few samples\u0026rdquo;, \u003cem\u003eCognitive science\u003c/em\u003e, 38(4):599\u0026ndash;637, 2014.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBhattacharya, Soham, Monoranjan Guchait, and Aravind H. Vijay. \u0026quot;Boosted top quark tagging and polarization measurement using machine learning.\u0026quot; Physical Review D 105.4 (2022): 042005.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eGentzkow, Matthew, Jesse Shapiro, and Matt Taddy. \u003cem\u003eMeasuring polarization in high-\u0026nbsp;\u003c/em\u003e\u003c/span\u003e\u003cspan\u003e\u003cem\u003edimensional data: Method and application to congressional speech\u003c/em\u003e. No. id: 11114. 2016.\u003c/span\u003e\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Belief Polarization, Machine Learning, Reinforcement Learning, etc","lastPublishedDoi":"10.21203/rs.3.rs-3942048/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3942048/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSome people polarize due to heated arguments but in reality and research studies says that they even polarize with similar information also. People think irrationally and confirms their pre-existing beliefs and strive to look into that information only. With these they avoid such data that contradicts them. Polarization occurs through the bias contribute to the reinforcement of existing beliefs. The people manipulate the new evidence based on the previous beliefs and bias confirmation suggestions. In order to form beliefs, people depend on least samples rather than gathering huge data and processing is also costly. A hypothesis is generated based on the least samples collected and then belief polarization model is used to test our hypothesis. A new belief formation model is developed based on reinforcement learning. In accordance to this a new evaluation metrics is used for polarization and simulations are learned how least size affects polarization. In this work basic assumptions are done with the generated hypothesis results, such that with these least samples polarization may increase. Practical suggestions are done for obtaining polarization based on our findings.\u003c/p\u003e","manuscriptTitle":"Modeling human cognition and polarization using reinforcement Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-12 18:52:07","doi":"10.21203/rs.3.rs-3942048/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7da61615-eb96-4b40-8791-132d1722dabd","owner":[],"postedDate":"February 12th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-06-13T11:39:19+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-12 18:52:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3942048","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3942048","identity":"rs-3942048","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.