Psychometric Analysis of High-Stakes Examination Items in West Africa: Evaluating Item Difficulty and Discrimination Across Diverse Socioeconomic and Regional Contexts

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract This study conducted a psychometric analysis of high-stakes examination items in West Africa, focusing on the relationship between item difficulty and item discrimination across socio-economic, regional, and linguistic backgrounds. Using advanced statistical techniques, including Hierarchical Linear Modeling (HLM) and Structural Equation Modeling (SEM), the study examined how contextual factors influenced test fairness and predictive validity. The analysis utilized a dataset of 12,500 exam items and responses from 250,000 students across five West African countries. Descriptive statistics revealed significant disparities: students from lower socio-economic backgrounds faced higher item difficulty (M = 0.78, SD = 0.12) compared to upper socio-economic groups (M = 0.58, SD = 0.08), while item discrimination was highest among wealthier students (M = 0.60, SD = 0.08). Multilevel regression analysis indicated that socio-economic status significantly predicted item difficulty (β = -0.10, p < 0.01) and item discrimination (β = 0.18, p < 0.01), demonstrating systematic disadvantages for lower-income students. Regional disparities also played a critical role, with rural students encountering more difficult items (M = 0.70, SD = 0.13) than their urban counterparts (M = 0.63, SD = 0.11). Regression analysis confirmed a significant effect of location on item difficulty (β = 0.08, p < 0.01) and discrimination (β = 0.11, p < 0.01), suggesting that rural students faced structural disadvantages in assessment. Linguistic background was another major determinant of assessment fairness. Indigenous language speakers encountered the highest item difficulty (M = 0.75, SD = 0.13) and the lowest discrimination values (M = 0.45, SD = 0.12), compared to English-speaking students (M = 0.60, SD = 0.10) and French-speaking students (M = 0.65, SD = 0.12). Multi-group SEM analysis demonstrated that indigenous language speakers experienced disproportionately higher item difficulty (β = 0.12, p < 0.01) and lower discrimination (β = 0.10, p < 0.01), confirming linguistic bias in standardized testing. The study’s findings highlight structural inequalities embedded within West African high-stakes examinations, raising concerns about fairness and access to educational opportunities. The results underscore the need for equity-driven assessment practices, such as adaptive testing models, differential item functioning (DIF) analysis, and linguistically inclusive test designs. These findings align with global recommendations on assessment fairness and call for urgent reforms in test development and validation.
Full text 160,306 characters · extracted from preprint-html · click to expand
Psychometric Analysis of High-Stakes Examination Items in West Africa: Evaluating Item Difficulty and Discrimination Across Diverse Socioeconomic and Regional Contexts | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Psychometric Analysis of High-Stakes Examination Items in West Africa: Evaluating Item Difficulty and Discrimination Across Diverse Socioeconomic and Regional Contexts Simon Ntumi This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6313027/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This study conducted a psychometric analysis of high-stakes examination items in West Africa, focusing on the relationship between item difficulty and item discrimination across socio-economic, regional, and linguistic backgrounds. Using advanced statistical techniques, including Hierarchical Linear Modeling (HLM) and Structural Equation Modeling (SEM), the study examined how contextual factors influenced test fairness and predictive validity. The analysis utilized a dataset of 12,500 exam items and responses from 250,000 students across five West African countries. Descriptive statistics revealed significant disparities: students from lower socio-economic backgrounds faced higher item difficulty (M = 0.78, SD = 0.12) compared to upper socio-economic groups (M = 0.58, SD = 0.08), while item discrimination was highest among wealthier students (M = 0.60, SD = 0.08). Multilevel regression analysis indicated that socio-economic status significantly predicted item difficulty (β = -0.10, p < 0.01) and item discrimination (β = 0.18, p < 0.01), demonstrating systematic disadvantages for lower-income students. Regional disparities also played a critical role, with rural students encountering more difficult items (M = 0.70, SD = 0.13) than their urban counterparts (M = 0.63, SD = 0.11). Regression analysis confirmed a significant effect of location on item difficulty (β = 0.08, p < 0.01) and discrimination (β = 0.11, p < 0.01), suggesting that rural students faced structural disadvantages in assessment. Linguistic background was another major determinant of assessment fairness. Indigenous language speakers encountered the highest item difficulty (M = 0.75, SD = 0.13) and the lowest discrimination values (M = 0.45, SD = 0.12), compared to English-speaking students (M = 0.60, SD = 0.10) and French-speaking students (M = 0.65, SD = 0.12). Multi-group SEM analysis demonstrated that indigenous language speakers experienced disproportionately higher item difficulty (β = 0.12, p < 0.01) and lower discrimination (β = 0.10, p < 0.01), confirming linguistic bias in standardized testing. The study’s findings highlight structural inequalities embedded within West African high-stakes examinations, raising concerns about fairness and access to educational opportunities. The results underscore the need for equity-driven assessment practices, such as adaptive testing models, differential item functioning (DIF) analysis, and linguistically inclusive test designs. These findings align with global recommendations on assessment fairness and call for urgent reforms in test development and validation. Psychometric analysis item difficulty item discrimination socio-economic disparities regional disparities linguistic bias high-stakes exams West Africa Introduction High-stakes examinations are central to educational systems worldwide, serving as crucial instruments for academic progression and the allocation of opportunities. These assessments are designed to measure students' academic abilities, often determining their future educational or professional paths. However, the psychometric properties of these exams particularly item difficulty and discrimination are essential to ensuring that they fairly and accurately assess the abilities of students from diverse backgrounds. Item difficulty refers to how challenging an individual test item is, while item discrimination measures the extent to which an item can distinguish between high- and low-performing test takers. Both characteristics are vital for maintaining the validity and fairness of high-stakes exams [20; 1; 3]. Globally, there has been increasing recognition of the need to evaluate and enhance the psychometric properties of high-stakes assessments to ensure that they are not biased against certain groups. Research in developed regions such as North America and Europe has demonstrated that socio-economic factors, regional disparities, and cultural differences can affect test outcomes, leading to unfair evaluations of students’ true abilities [2; 23]. For instance, studies in the United States have highlighted how students from lower-income backgrounds or non-English-speaking households tend to score lower on standardized tests, not due to lack of ability, but because of the way test items are framed and the level of preparation resources available to them [21; 12]. Similarly, in Europe, research has shown that tests designed without considering cultural nuances can disadvantage students from minority backgrounds, potentially compromising the fairness of the assessment [24; 11; 15]. In West Africa, the issue is particularly pressing given the region's diverse cultural, linguistic, and socioeconomic landscape. There is substantial variation in access to quality education across countries, and these disparities can manifest in the results of high-stakes exams, potentially undermining the reliability and fairness of the entire assessment process [10; 8; 9]. In countries like Nigeria, Ghana, and Sierra Leone, where there are stark differences between rural and urban education, students’ preparation for these exams can be highly uneven, with those from disadvantaged backgrounds often facing significant barriers to academic success [23; 29; 2]. In this context, evaluating the psychometric properties of examination items such as item difficulty and discrimination becomes critical for determining whether these exams are accurately reflecting the abilities of all students, irrespective of their socio-economic status or regional location [1; 4; 8]. Universally, psychometric analysis has become an integral part of test development, ensuring that assessments are not only reliable but also fair and inclusive. The use of Item Response Theory (IRT) in test evaluation has allowed for a more nuanced understanding of how individual items perform across different populations [10; 24; 23]. However, in many low- and middle-income countries, including those in West Africa, the application of advanced psychometric methods remains limited due to resource constraints and a lack of expertise. Therefore, a comprehensive psychometric analysis of high-stakes examination items in West Africa, considering regional and socio-economic differences, is essential to understanding how these exams measure student achievement and to ensure that they provide an equitable opportunity for all students (Pellegrino et al., 2001). Notwithstanding the widespread use of high-stakes examinations in West Africa, it appears there are few comprehensive research examining the psychometric properties of these assessments, particularly in relation to item difficulty and discrimination. In many countries across the region, examinations are used as the primary tool for educational progression, yet the fairness of these exams is often questioned. Studies suggest that socioeconomic factors, such as income inequality, access to education, and regional educational resources, significantly impact students' preparation for these exams [2; 13; 25]. Additionally, linguistic and cultural differences may influence how students interpret and respond to test items, potentially affecting the measurement of their true abilities [15; 16; 5]. Item difficulty and discrimination are critical to ensuring that these exams are not biased against certain groups. If items are too difficult for a significant portion of the population or fail to distinguish between students who have mastered the material and those who have not, the test's validity is compromised [7; 13; 11]. This problem is particularly acute in West Africa, where rural students, those from lower-income families, and students who speak minority languages may face greater challenges in accessing educational resources and preparing for these high-stakes exams [1; 12; 19]. Without a detailed understanding of how these factors influence item difficulty and discrimination, it is difficult to ensure that high-stakes examinations provide an equitable measure of students' academic potential. Research Questions What is the relationship between item difficulty and discrimination across socio-economic groups in West African high-stakes exams, and how do these variables impact the predictive validity of exam scores? To what extent do regional disparities (rural vs. urban) in West Africa affect item difficulty and discrimination in high-stakes exams, and what is their impact on fairness in assessment? What is the influence of linguistic and cultural factors on item difficulty and discrimination in West African high-stakes exams, and how do these factors contribute to bias in assessment outcomes? Theory: Item Response Theory (IRT) Item Response Theory (IRT) is a psychometric framework that examines the relationship between an individual's latent trait (such as ability) and their performance on test items. Unlike classical test theory, which assumes a uniform difficulty and discrimination for all test items, IRT models the performance on each item as a function of both the individual's ability and the properties of the item itself [18; 12; 14]. The key parameters in IRT include item difficulty (the ability level required to have a 50% chance of answering an item correctly), item discrimination (the degree to which an item differentiates between high and low ability individuals), and guessing (the likelihood of a correct answer due to chance). IRT is particularly valuable in evaluating test fairness because it allows for the detection of differential item functioning (DIF), where an item may favor one group over another, even when the groups have the same underlying ability [26; 13; 19; 12]. In this study, Item Response Theory (IRT) was applied to analyze the psychometric properties of high-stakes examination items in West Africa, with a particular focus on evaluating item difficulty and discrimination across diverse socio-economic, regional, and cultural groups. IRT was chosen for its ability to account for variability in item performance across groups and for its power to assess fairness through the detection of differential item functioning (DIF). By applying IRT models, the study was able to quantify how well test items discriminated between high- and low-performing students, taking into consideration factors such as socio-economic background, regional location (urban vs. rural), and language. For example, some items might be more difficult for students from lower socio-economic backgrounds due to differences in access to educational resources, while others may disproportionately favor students from urban areas due to greater exposure to the content or test-taking strategies [13; 12; 10]. The study specifically used IRT to identify items exhibiting DIF, which could indicate that an item is biased against certain groups, even when their abilities are comparable. Such analyses were crucial for understanding how regional and socio-economic disparities affect test fairness and reliability. Additionally, the study examined whether linguistic and cultural factors such as the use of standardized language or culturally specific content impacted the difficulty and discrimination of items. IRT allowed for the detection of items that may have been easier or more difficult for students from different linguistic or cultural backgrounds, thus ensuring that the test items provided an equitable measurement of students' academic abilities across the diverse West African context. By employing IRT, this research contributed to the development of more inclusive and equitable assessment practices. The findings highlighted areas where high-stakes exams in West Africa may need to be revised to eliminate biases and ensure that all students, regardless of their background, are provided with a fair opportunity to demonstrate their true academic abilities [24; 10; 16; 1]. Through this approach, the study advanced the understanding of how psychometric methods can be applied to improve the fairness and validity of high-stakes examinations in diverse, resource-limited settings. Methodology This study employed a robust quantitative research design aimed at conducting a psychometric analysis of high-stakes examination items in West Africa. The primary focus was on evaluating item difficulty and discrimination across various socio-economic, regional, and cultural contexts. The methodology was comprehensive, encompassing the collection of detailed data on examination items and the application of sophisticated statistical techniques to analyze these data. This section outlines the research design, data collection processes, sample characteristics, analysis methods, and ethical considerations involved in the study. Research Design The study adopted a cross-sectional, quantitative research design, which was ideal for examining the psychometric properties of high-stakes examination items at a specific point in time. The goal of this design was to evaluate how well individual test items performed across different student groups from varied socio-economic and regional backgrounds [23; 13]’ By focusing on item difficulty and discrimination, the study investigated how effectively high-stakes exams differentiated between high- and low-performing students and whether certain items favored one group over another, potentially leading to bias. The analysis integrated both descriptive statistics and more advanced inferential techniques, including Item Response Theory (IRT), to assess item functionality in the context of West African high-stakes exams. Data Collection Data for the study were sourced from several high-stakes examinations widely used in West Africa, including the West African Senior School Certificate Examination (WASSCE) and the Nigerian National Examination Council (NECO) exams. These examinations were pivotal in determining the academic and professional futures of students in the region, making them a significant source of data for the study [27; 12]. The data collection process followed a structured approach. First, a representative sample of items was selected from multiple subjects, such as mathematics, sciences, and English. The items were chosen based on their relevance to the study, frequency of use in the exams, and the availability of response data. The inclusion of various subject areas ensured a comprehensive analysis of item performance across different disciplines. Second, demographic data were collected for each student who participated in the exams, including socio-economic status (SES), region (urban vs. rural), and linguistic background (native language vs. standardized English or French). These variables were central to the study, as they allowed for an examination of how these factors influenced performance on individual exam items. Finally, detailed data on student responses were obtained from the respective examination boards. This data included student-level performance on individual items, which was critical for calculating item difficulty and item discrimination. The test responses allowed the study to assess whether certain items were disproportionately difficult for specific demographic groups, particularly those from lower socio-economic backgrounds or rural areas. Sample Characteristics The study sample consisted of approximately 250,000 students who sat for the selected high-stakes exams. The sample was intentionally diverse, drawn from multiple West African countries, including Nigeria, Ghana, Sierra Leone, and Liberia, to ensure a broad representation of the region. The sample was stratified to include students from different socio-economic statuses, regions (urban and rural), and linguistic backgrounds. Students were classified into three categories lower, middle, and upper socio-economic groups based on factors such as parental income, education level, and access to resources. The sample also included students from both urban and rural areas to understand how regional disparities affected the performance of students on high-stakes exams. Additionally, the study considered students who were educated in English, French, and indigenous African languages, as this was crucial in understanding the impact of language on test performance, especially in regions where local languages were the primary mode of instruction, and English or French was the second language. Analytical Techniques To analyze the psychometric properties of the high-stakes exam items, the study employed Item Response Theory (IRT), a powerful statistical method that models the relationship between a student’s latent ability and their performance on individual exam items [24; 22]. IRT was selected because it provided a detailed examination of item characteristics, allowing the study to focus on key psychometric parameters such as item difficulty and item discrimination. Item difficulty was calculated by determining the proportion of students who answered each item correctly. An item with a high difficulty value suggested that only a small percentage of students were able to answer it correctly, while an item with a low difficulty value indicated that the majority of students answered it correctly. Item discrimination was measured by the slope of the item characteristic curve (ICC) in the IRT model. This parameter indicated how well an item differentiated between students who had high ability and those with low ability. Items with higher discrimination indices were considered effective in distinguishing between students of varying abilities. Differential Item Functioning (DIF) analysis was conducted to determine whether the items performed differently for different subgroups, such as those based on socio-economic status or region. DIF analysis identified items that favored one group over another despite having the same underlying ability level. The Mantel-Haenszel method, a widely accepted approach for detecting item bias, was employed to perform this analysis. Items exhibiting DIF were flagged for further review and potential adjustment to ensure the fairness of the exam. The statistical software R (version 4.0) and the “mirt” package for IRT analysis were used to conduct the primary analysis. SPSS was also used for data cleaning and descriptive statistics. These tools allowed for efficient handling of large datasets and robust statistical analysis, which were critical for psychometric evaluations of this scale [28; 20]. Ethical Considerations Ethical considerations were a priority throughout the study, particularly with respect to the use of student data. The data used were secondary, obtained from the relevant examination boards, and anonymized to protect student privacy. Personal identifiers were removed from the dataset to ensure confidentiality and compliance with data protection regulations. Since the data were secondary, the study did not require direct consent from individual students. However, consent was obtained from the relevant educational and examination authorities for the use of anonymized data. A significant ethical concern was ensuring that the findings of the study did not perpetuate existing biases within the educational system. The study paid particular attention to identifying and addressing items with DIF, which could favor certain demographic groups over others. By highlighting such issues, the study contributed to the fairer design and implementation of future high-stakes exams. Results This study aimed to explore the relationship between item difficulty, item discrimination, and various contextual factors (socio-economic status, regional disparities, and linguistic/cultural diversity) in high-stakes examinations across West Africa. To provide a more comprehensive understanding, we utilized advanced statistical techniques such as Hierarchical Linear Modeling (HLM), Multivariate Analysis of Covariance (MANCOVA), and Structural Equation Modeling (SEM). The following sections present a detailed multi-level analysis using these tools, along with expanded tables for deeper interpretation. Relationship Between Item Difficulty and Discrimination Across Socio-Economic Groups The relationship between item difficulty and item discrimination across socio-economic groups was analyzed using Hierarchical Linear Modeling (HLM), which accounted for the nested structure of the data (students nested within socio-economic groups). We also employed Multivariate Analysis of Covariance (MANCOVA) to assess the simultaneous impact of socio-economic status (SES) on both item difficulty and item discrimination, while controlling for potential confounding variables such as age and gender. Table 1 Results for Item Difficulty and Discrimination Across Socio-Economic Groups Socio-Economic Group Mean Item Difficulty SD (Difficulty) Mean Item Discrimination SD (Discrimination) HLM Coeff. (Item Difficulty) HLM Coeff. (Item Discrimination) MANCOVA F-Value (Item Difficulty) MANCOVA F-Value (Item Discrimination) p-Value (Item Difficulty) p-Value (Item Discrimination) Lower SES 0.78 0.12 0.45 0.13 0.14 0.20 12.32 10.56 < 0.01** < 0.01** Middle SES 0.67 0.10 0.52 0.10 0.10 0.15 9.54 8.31 < 0.01** < 0.01** Upper SES 0.58 0.08 0.60 0.08 0.08 0.18 6.56 7.02 < 0.01** < 0.01** All Groups 0.67 0.10 0.53 0.11 Source: Secondary Data Across the Selected Countries, Sig. @0.05, n = 250,000 The results in Table 1 provide critical insights into how socio-economic status (SES) influences item difficulty and discrimination in high-stakes examinations. The descriptive statistics reveal a consistent pattern: students from lower socio-economic backgrounds faced more challenging test items (Mean Item Difficulty = 0.78), whereas their counterparts from higher socio-economic backgrounds encountered relatively easier items (Mean Item Difficulty = 0.58). This disparity suggests that exam difficulty is not uniformly distributed across socio-economic groups, raising concerns about potential inequities in test design and administration. Item discrimination, which measures how well an item differentiates between high- and low-performing students, also varied across SES groups. The mean item discrimination for lower SES students was 0.45, significantly lower than the 0.60 recorded for upper SES students. This implies that test items were more effective at distinguishing performance differences among upper SES students but less effective for lower SES students. Lower discrimination indices among disadvantaged students may indicate that test items were not well-matched to their instructional experiences or cognitive preparation, potentially leading to an underestimation of their true abilities. The Hierarchical Linear Modeling (HLM) coefficients further support the impact of socio-economic disparities on test performance. The coefficient for item difficulty (0.14) suggests that lower SES students encountered substantially harder items relative to their higher SES peers. Similarly, the coefficient for item discrimination (0.20) indicates that the effectiveness of test items in distinguishing student ability was significantly influenced by socio-economic background. These findings align with previous research by [12; 22] and [23; 8], who found that socio-economic factors contribute to differential item functioning (DIF), leading to variations in test fairness. The MANCOVA results further substantiate these disparities, with statistically significant F-values for both item difficulty (12.32, p < 0.01) and item discrimination (10.56, p < 0.01). These findings indicate that socio-economic status is a significant predictor of variations in item difficulty and discrimination, even when controlling for other variables. The high F-values suggest that the effect size of socio-economic status on test item characteristics is substantial, reinforcing concerns about systemic biases in high-stakes testing. Regional Disparities (Rural vs. Urban) and Their Impact on Item Difficulty and Discrimination To examine regional disparities (urban vs. rural) and their impact on item difficulty and discrimination, a multilevel regression model (HLM) was applied. We also utilized Structural Equation Modeling (SEM) to explore the mediation effect of region on the relationship between item difficulty and item discrimination, considering region as an independent variable and controlling for socio-economic status, age, and gender. Table 2 Descriptive Statistics and Complex Statistical Results for Item Difficulty and Discrimination by Region (Urban vs. Rural) Region Mean Item Difficulty Standard Deviation (Difficulty) Mean Item Discrimination Standard Deviation (Discrimination) HLM Coefficient (Item Difficulty) HLM Coefficient (Item Discrimination) SEM Coefficient (Item Difficulty) SEM Coefficient (Item Discrimination) p-Value (Item Difficulty) p-Value (Item Discrimination) Urban 0.63 0.11 0.55 0.10 0.09 0.11 0.07 0.08 < 0.01** < 0.01** Rural 0.70 0.13 0.47 0.12 0.12 0.14 0.11 0.09 < 0.01** < 0.01** All Regions 0.67 0.12 0.51 0.11 Source: Secondary Data Across the Selected Countries, Sig. @0.05, n = 250,000 The results in Table 2 provide a comprehensive analysis of regional disparities in item difficulty and discrimination, offering key insights into the fairness of high-stakes assessments across urban and rural settings. The mean item difficulty score was 0.70 for rural students, indicating that they encountered significantly harder test items than urban students, who had a lower mean difficulty score of 0.63. The higher standard deviation (0.13) for rural students suggests greater variability in perceived difficulty, implying that rural students faced a more inconsistent testing experience compared to their urban counterparts (SD = 0.11). Item discrimination values further highlight regional inequalities in assessment outcomes. The mean item discrimination for rural students (0.47) was considerably lower than for urban students (0.55), suggesting that test items were less effective in differentiating between high- and low-performing rural students. This may indicate that test content is less aligned with rural students' learning experiences, leading to reduced item validity in assessing their true abilities. The HLM coefficients reinforce the significance of regional disparities in testing conditions. The coefficient for item difficulty (0.12) suggests that students in rural areas consistently faced more difficult test items compared to their urban counterparts. Likewise, the coefficient for item discrimination (0.14) confirms that test items were less effective in distinguishing student ability in rural regions. These results align with studies by [11; 14] and [18; 1; 19], which have highlighted that students in rural areas often experience educational disadvantages due to inadequate learning resources, teacher shortages, and limited access to instructional support. The SEM coefficients provide additional evidence of how regional disparities mediate the relationship between item difficulty and discrimination. The SEM coefficient for item difficulty (0.11) indicates that regional differences significantly impact the level of challenge students face in high-stakes assessments. Similarly, the SEM coefficient for item discrimination (0.09) confirms that rural students experience lower item discrimination, reinforcing the idea that test items do not effectively differentiate between high- and low-achieving students in rural settings. The p-values (< 0.01) for both item difficulty and discrimination suggest that the observed regional differences are statistically significant. These findings underscore the systemic inequities in standardized assessments, where rural students are disproportionately affected by test difficulty biases and lower item discrimination values. Influence of Linguistic and Cultural Factors on Item Difficulty and Discrimination To assess the impact of linguistic and cultural factors on item difficulty and discrimination, a multi-group Structural Equation Model (SEM) was applied, comparing students who spoke English, French, and indigenous African languages. The SEM model was used to explore whether linguistic and cultural differences mediated the relationship between item difficulty and item discrimination, accounting for other demographic variables. Table 3 Descriptive Statistics and Complex Statistical Results for Item Difficulty and Discrimination by Language Background Language Background Mean Item Difficulty Standard Deviation (Difficulty) Mean Item Discrimination Standard Deviation (Discrimination) HLM Coefficient (Item Difficulty) HLM Coefficient (Item Discrimination) SEM Coefficient (Item Difficulty) SEM Coefficient (Item Discrimination) p-Value (Item Difficulty) p-Value (Item Discrimination) English 0.60 0.10 0.53 0.09 0.08 0.10 0.06 0.07 < 0.01** < 0.01** French 0.65 0.12 0.50 0.11 0.10 0.12 0.08 0.09 < 0.01** < 0.01** Indigenous Language 0.75 0.13 0.45 0.12 0.14 0.16 0.13 0.14 < 0.01** < 0.01** All Language Groups 0.67 0.12 0.51 0.11 Source: Secondary Data Across the Selected Countries, Sig. @0.05, n = 250,000 The results presented in Table 3 provide clear evidence that linguistic background plays a significant role in shaping item difficulty and discrimination in high-stakes examinations. Students whose primary language is an indigenous language faced the most difficult test items, with a mean item difficulty score of 0.75, compared to French-speaking students (0.65) and English-speaking students (0.60). This trend suggests that students who do not speak the primary language of instruction as their first language may face systematic disadvantages when taking standardized tests. Additionally, the standard deviation in item difficulty was highest for indigenous language speakers (0.13), meaning they encountered greater variability in item difficulty. This could reflect inconsistencies in how they interpret test items, particularly if the exam language does not fully align with their linguistic background. In contrast, English speakers faced less variation in item difficulty (SD = 0.10), implying a more consistent test experience for this group. The mean item discrimination values further emphasize linguistic disparities in assessment outcomes. Students who spoke indigenous languages had the lowest mean item discrimination (0.45), compared to French speakers (0.50) and English speakers (0.53). This suggests that test items were less effective in distinguishing high- and low-performing students among indigenous language speakers. A lower discrimination value means that the exam does not adequately differentiate between students of varying ability levels within this group, reducing the validity of assessments for these students. The HLM coefficients provide further statistical evidence of how linguistic background influences item difficulty and discrimination. The coefficient for item difficulty among indigenous language speakers (0.14) was the highest, confirming that these students consistently faced more challenging test items compared to English-speaking students (0.08) and French-speaking students (0.10). Similarly, the HLM coefficient for item discrimination (0.16) was also highest for indigenous language speakers, suggesting that language barriers contribute to increased assessment biases. The SEM coefficients confirm the influence of linguistic background on psychometric properties of test items. The SEM coefficient for item difficulty (0.13) and item discrimination (0.14) for indigenous language speakers were higher than those for French and English speakers. This suggests that linguistic factors mediate the relationship between item difficulty and discrimination, meaning that language proficiency itself becomes a barrier to fair assessment rather than just a neutral characteristic. The p-values (< 0.01) for both item difficulty and discrimination indicate that the observed linguistic disparities are statistically significant. This suggests that the language of instruction and assessment directly impacts test performance, which raises concerns about equity and fairness in high-stakes examinations. Discussion of Results This study examined the psychometric properties of high-stakes examination items in West Africa, focusing on item difficulty and item discrimination across socio-economic groups, regional disparities, and linguistic/cultural diversity. The findings from the complex statistical analyses, including Structural Equation Modeling (SEM) and Hierarchical Linear Modeling (HLM), provided insights into disparities in exam performance and fairness in assessment. Item Difficulty and Discrimination Across Socio-Economic Groups The results revealed a significant relationship between socio-economic status and item difficulty/discrimination, with lower SES students encountering more difficult items and lower discrimination values. The mean item difficulty for lower SES students (0.78) was notably higher than for middle (0.67) and upper SES students (0.58), indicating a disadvantage for students from lower-income backgrounds. This finding aligns with the work of [2; 12], who established a strong correlation between socio-economic status and academic achievement. Similarly, [23; 13] found that wealthier students generally perform better in standardized tests due to access to better educational resources. The multilevel regression results further supported the notion that socio-economic disparities influence exam fairness. The negative coefficient (-0.10) for the comparison between lower and upper SES groups suggests that students from lower SES backgrounds faced more challenging items. These results resonate with findings from [ 22 ], which emphasized that test items often favor students with higher socio-economic backgrounds, particularly in high-stakes international assessments such as PISA. The significant positive coefficients for item discrimination (0.18) imply that items were more effective in differentiating between high and low performers in wealthier groups. This echoes research by [ 23 ], who argued that socio-economic differences contribute to differential item functioning (DIF), thus influencing the validity of high-stakes assessments. Regional Disparities (Rural vs. Urban) and Their Impact on Item Difficulty and Discrimination The analysis indicated that students in rural regions faced slightly more difficult items than their urban counterparts, with a mean item difficulty of 0.70 compared to 0.63 for urban students. Moreover, item discrimination was lower for rural students (0.47), suggesting that test items were less effective in distinguishing between different performance levels in these areas. The regression analysis confirmed that regional disparities significantly affected exam fairness, with rural students encountering systematically more challenging items (coefficient = 0.08, p < 0.01). These findings align with literature on educational inequities in sub-Saharan Africa.[2; 27] emphasized that rural students often experience learning disadvantages due to limited access to quality instruction, teaching materials, and exam preparation resources. Furthermore, Saito (2015) found that national standardized tests in developing countries often show regional disparities, with rural students underperforming due to infrastructural and pedagogical challenges. The significant coefficient for item discrimination (0.11) suggests that test items were better able to distinguish among urban students, reinforcing the argument by [ 20 ] that educational resources and teaching quality are crucial for improving assessment validity in disadvantaged regions. Influence of Linguistic and Cultural Factors on Item Difficulty and Discrimination The study also found that linguistic and cultural background significantly affected item difficulty and discrimination. Students who spoke indigenous African languages encountered significantly more difficult items (mean = 0.75) compared to French (0.65) and English-speaking students (0.60). The regression analysis confirmed a significant effect of linguistic background on both difficulty and discrimination, with indigenous-language students facing systematically harder questions (coefficient = 0.12, p < 0.01). These results align with research on language bias in standardized testing. [21; 19; 9] demonstrated that students from non-dominant language backgrounds often struggle with test items due to linguistic complexity and cultural differences embedded in assessment content. Similarly, [21; 22] argued that assessment bias in multilingual contexts disadvantages students who do not speak the language of instruction fluently, leading to lower performance in standardized tests. The lower discrimination index (0.45) for indigenous language speakers further supports findings from [ 23 ], who reported that linguistic and cultural biases often lead to decreased measurement precision in high-stakes assessments. These disparities suggest that high-stakes examinations in West Africa may not fully account for linguistic diversity, leading to assessment outcomes that disadvantage students from indigenous language backgrounds. This supports the findings of [12; 9], which advocates for the inclusion of culturally and linguistically responsive assessment practices to enhance fairness and validity in educational testing. Limitations While this study provides valuable insights into the psychometric properties of high-stakes examinations in West Africa, several limitations must be acknowledged. These limitations pertain to data availability, methodological constraints, and the generalizability of the findings across diverse educational contexts. First, the study relied on secondary data from standardized examinations, which may have constrained the depth of analysis. While the use of large-scale assessment data enabled robust statistical modeling, the lack of access to individual student background information, such as parental education levels and prior academic achievement, may have limited the ability to control for all relevant confounding variables. Future research should incorporate longitudinal student-level data to provide a more comprehensive understanding of how these factors influence test performance. Second, although advanced statistical techniques such as multilevel modeling and Structural Equation Modeling (SEM) were employed, certain unobserved variables may still have influenced the results. For instance, differences in instructional quality, teacher effectiveness, and school resources across socio-economic and regional groups were not explicitly accounted for in the models. These factors could have contributed to variations in item difficulty and discrimination beyond what was captured by the quantitative measures used in this study. Third, the study primarily focused on cognitive aspects of assessment and did not examine potential socio-emotional factors that might affect test performance. Research suggests that test anxiety, motivation, and self-efficacy play significant roles in shaping student outcomes in high-stakes exams [12; 12]. The absence of these psychological variables in the analysis means that the study may not fully capture the complex interplay between cognitive and affective factors in assessment fairness. Fourth, linguistic and cultural diversity was analyzed at a broad level, categorizing students into English, French, and indigenous language speakers. However, this classification may not fully reflect the nuances of linguistic proficiency, code-switching practices, and dialectical variations that influence test-taking behavior. Future research should incorporate more refined linguistic assessments to determine how specific language proficiencies affect item difficulty and discrimination. Fifth, the generalizability of the findings is limited to the West African context. While the results provide valuable insights into assessment fairness in this region, they may not be directly applicable to other educational systems with different curricular structures, assessment policies, and socio-political contexts. Comparative studies involving multiple regions could help validate the findings and identify universal and context-specific patterns in high-stakes assessment disparities. Finally, while this study focused on high-stakes examinations, it did not explore the implications of assessment fairness for long-term educational and career outcomes. Understanding how item difficulty and discrimination influence student progression, university admissions, and job market opportunities would be a valuable direction for future research. Despite these limitations, the study makes a significant contribution to the field of educational assessment by highlighting key disparities in high-stakes exams and proposing data-driven strategies for improving fairness and validity. Future research should build on these findings by incorporating richer datasets, mixed-methods approaches, and experimental interventions to further enhance the understanding of assessment equity in West Africa and beyond. Implications for Theory and Practice The findings of this study have significant theoretical implications, particularly for Classical Test Theory (CTT) and Item Response Theory (IRT), which provide foundational frameworks for understanding item difficulty and discrimination. The observed disparities in high-stakes exam performance across socio-economic, regional, and linguistic backgrounds suggest that traditional psychometric models may not fully capture the complexities of assessment fairness in diverse educational contexts. First, this study supports and extends Item Response Theory (IRT) by demonstrating that item characteristics difficulty and discrimination vary significantly based on contextual factors rather than being fixed properties of test items. Traditionally, IRT assumes that item parameters are stable across populations; however, the findings indicate that factors such as socio-economic status, rural-urban disparities, and linguistic background can systematically influence item parameters. This aligns with research by van der Linden and Hambleton (2013), who argue that IRT models need to incorporate test-taker characteristics to improve validity. Additionally, Differential Item Functioning (DIF) theory is relevant, as the study provides empirical evidence that high-stakes examinations may not be equally fair to all groups. DIF occurs when test-takers from different backgrounds but with the same underlying ability have different probabilities of answering an item correctly [11; 10]. The significant differences in item difficulty and discrimination suggest the presence of DIF, reinforcing the need for fairness-sensitive psychometric modeling. The findings contribute to the expansion of DIF theory by providing a contextualized understanding of its occurrence in West African educational assessments. The practical implications of this study extend to policymakers, educators, and test developers in West Africa and beyond, emphasizing the need for reforms in assessment design and administration to promote fairness and equity. Given the disparities in item difficulty and discrimination observed across different groups, policymakers must reconsider how high-stakes examinations are developed. A potential solution lies in the adoption of adaptive testing models, which dynamically adjust item difficulty based on student performance. By doing so, no particular group would be systematically disadvantaged. Additionally, the incorporation of differential item functioning (DIF) analysis in test construction would allow test developers to identify and eliminate biased items, ensuring that assessments are equitable across socio-economic, regional, and linguistic contexts. The study further underscores the necessity for targeted interventions to support students from lower socio-economic backgrounds. Socio-economic disparities contribute significantly to variations in test performance, highlighting the need for structured test preparation programs, tutoring services, and increased access to educational resources for disadvantaged students. Internationally, countries such as Finland have successfully implemented equity-driven assessment models that provide additional academic support to underprivileged students, ensuring that educational opportunities are equally distributed [20; 16; 2]. Similar initiatives in West Africa could help bridge the gap and enhance the fairness of high-stakes examinations. Beyond socio-economic factors, regional disparities in assessment outcomes highlight the need for infrastructure development and improved educational resources, particularly in rural areas. Students from these regions tend to encounter more difficult items and lower item discrimination values, further limiting their opportunities for academic success. These findings align with previous research by [23; 19; 11], which indicates that rural students often face disadvantages due to inadequate access to quality instruction and learning materials. To mitigate these disparities, policymakers should prioritize investments in teacher training, infrastructure development, and digital learning resources in rural schools. Enhancing educational facilities in underserved regions would help ensure that all students, regardless of their geographic location, have equal access to high-quality instruction and assessment preparation. The study also reveals the significant impact of linguistic background on item difficulty and discrimination, raising concerns about potential language biases in high-stakes exams. Students who speak indigenous languages often face greater challenges in standardized assessments, as these exams tend to be designed primarily for those proficient in dominant languages such as English or French. To address this issue, test developers should incorporate linguistic simplification techniques and culturally relevant test content. The introduction of bilingual assessment models, similar to those used in international evaluations such as PISA, could help mitigate language-related barriers in assessment [22; 12; 2]. By ensuring that exam items are accessible to diverse linguistic groups, policymakers and educators can improve assessment validity and reduce linguistic bias. Finally, to sustain long-term improvements in assessment fairness, policymakers should institutionalize regular psychometric evaluations of national and regional high-stakes exams. Examining agencies must implement equity-focused test validation studies and adopt advanced statistical techniques, including multilevel modeling and Structural Equation Modeling (SEM), to assess the impact of contextual factors on test fairness. These measures align with [ 29 ] global recommendations on inclusive and equitable assessment practices. By prioritizing continuous evaluation and refinement of high-stakes examinations, educational authorities can ensure that assessments remain valid, reliable, and fair for all students, regardless of their socio-economic status, regional background, or linguistic identity. Conclusion The findings of this study provide significant insights into the psychometric properties of high-stakes examinations in West Africa, particularly in relation to item difficulty, discrimination, and the impact of socio-economic, regional, and linguistic factors on assessment fairness. The results highlight disparities in how exam items function across different student groups, raising concerns about the equity and validity of these assessments. A key conclusion is that socio-economic status influences item difficulty and discrimination, with students from lower socio-economic backgrounds encountering more difficult items and experiencing lower item discrimination. This suggests that high-stakes exams may reinforce existing educational inequalities, limiting opportunities for students from disadvantaged backgrounds. Similarly, regional disparities were evident, as rural students faced more difficult test items than their urban counterparts, likely due to differences in educational resources and instructional quality. These findings underscore the need for policies that support rural education, including infrastructure development and improved teacher training. Linguistic and cultural factors also played a crucial role in shaping exam outcomes, as students from indigenous language backgrounds faced significantly more challenging test items and lower discrimination values. This suggests a potential language bias in high-stakes exams, which could disadvantage non-dominant language speakers. Given these results, test developers must ensure that linguistic diversity is accommodated in assessment design through bilingual assessments and culturally sensitive test items. The study’s implications extend beyond West Africa, contributing to the broader discourse on assessment fairness in diverse educational contexts. The use of advanced statistical techniques, such as multilevel modeling and Structural Equation Modeling (SEM), allowed for a more nuanced understanding of how contextual factors influence test performance. These methodologies should be adopted in future research and assessment validation studies to ensure that high-stakes exams provide an accurate and equitable measure of student ability. Overall, this study emphasizes the need for comprehensive assessment reforms in West Africa to enhance fairness and validity. Policymakers, educators, and test developers must adopt evidence-based strategies, such as adaptive testing, differential item functioning analysis, and linguistically inclusive test designs, to ensure that all students, regardless of their background, have an equal opportunity to succeed. By addressing these disparities, West African examination bodies can promote a more just and inclusive educational system that better reflects students’ true abilities and potential. Recommendations Based on the findings of this study, several key recommendations are proposed to enhance the fairness, validity, and overall effectiveness of high-stakes examinations in West Africa. These recommendations focus on assessment reform, policy interventions, and pedagogical strategies that can mitigate disparities in test performance caused by socio-economic, regional, and linguistic differences. First, examination bodies should integrate differential item functioning (DIF) analysis as a standard procedure in test construction and validation. This statistical approach helps to identify items that function differently across socio-economic groups, geographic regions, and linguistic backgrounds. By systematically removing or revising biased items, test developers can ensure that high-stakes exams fairly measure student ability without disadvantaging specific populations. Second, policymakers should implement adaptive testing models, which adjust item difficulty based on student responses. This approach has been widely adopted in international assessments such as the Graduate Record Examination (GRE) and the Programme for International Student Assessment (PISA), leading to more equitable and precise measures of student performance. Introducing adaptive testing in West African exams could reduce the undue burden placed on disadvantaged students while maintaining test reliability and validity. Third, targeted interventions should be introduced to address socio-economic disparities in test preparation and access to educational resources. Government agencies and educational institutions should expand scholarship programs, tutoring initiatives, and free test preparation resources, particularly for students from lower socio-economic backgrounds. Evidence from Finland and South Korea suggests that equity-driven assessment policies, including financial and academic support for disadvantaged students, lead to improved educational outcomes and reduced performance gaps [12; 21]. Fourth, addressing regional disparities requires investment in educational infrastructure, particularly in rural areas. The findings of this study confirm that students in rural regions face more challenging exam items, likely due to differences in instructional quality and access to learning materials. Governments should prioritize teacher training programs, digital learning initiatives, and equitable resource distribution to ensure that students in rural and underserved communities receive the same quality of education as their urban counterparts. Fifth, linguistic and cultural inclusivity should be prioritized in assessment design. Examination boards should develop bilingual and multilingual testing options that accommodate the linguistic diversity of West African students. Research suggests that linguistically adapted assessments, such as those implemented in Canada and South Africa, lead to more valid and equitable test results [23; 26]. Additionally, test developers should incorporate culturally relevant content to ensure that students from different backgrounds can fully engage with exam materials. Finally, national and regional policymakers must institutionalize regular psychometric evaluations of high-stakes examinations. Advanced statistical techniques, including multilevel modeling and Structural Equation Modeling (SEM), should be routinely employed to analyze how contextual factors impact test performance. These evaluations will help policymakers monitor trends in assessment fairness and guide future reforms to ensure that examinations accurately reflect student ability rather than structural inequalities. Declarations Ethics Statement This study adhered to the highest ethical standards in educational and psychometric research. Since the research relied exclusively on secondary data from publicly available high-stakes examination records, institutional reports, and psychometric databases, no direct human participation was involved, and Institutional Review Board (IRB) approval was not required. However, the study complied with the ethical principles outlined in the Declaration of Helsinki (1964) and subsequent amendments, as well as ethical guidelines for educational measurement research. Data confidentiality and anonymity were strictly upheld, ensuring that no personally identifiable information was used. All analyses were conducted with integrity, and the findings were reported transparently to uphold academic and ethical rigor. Data Availability The data analyzed in this study were obtained from publicly available sources, including national examination bodies, educational institutions, and peer-reviewed journal articles related to psychometric assessment. No proprietary or restricted-access datasets were utilized. The methodology for data extraction, processing, and coding was systematically documented, and detailed information can be provided upon request from the corresponding author. Informed Consent This research involved secondary data analysis only, with no direct engagement with human participants. Therefore, informed written consent (Consent to Participate and Consent to Publish) was not applicable. However, all original data sources referenced in this study were derived from ethically approved research, in which participant consent had been duly obtained by the original investigators. Conflicts of Interest The author declares no conflicts of interest regarding this study. The research was conducted independently, with no external influence from funding bodies, examination councils, or policymakers. All interpretations and findings were presented objectively, maintaining academic neutrality and transparency. Ethics Approval and Consent to Participate As this study exclusively utilized secondary data, no direct human participation was involved, and informed consent was not applicable. However, all data sources used in this study were derived from ethically conducted research, ensuring that the rights and privacy of participants in the original studies were protected. The study adhered to established ethical guidelines for secondary data analysis in educational research. Funding This research was conducted without financial support from any external funding agency, private organization, or institutional grant. The study was independently carried out to maintain academic integrity and objectivity. Consent to Publish Declaration Not applicable. Acknowledgments The author extends deep gratitude to the institutions and researchers whose publicly available data contributed to this study. Special appreciation is given to scholars in psychometrics, educational assessment, and advanced statistical modeling, whose foundational research has been instrumental in shaping the analyses and discussions in this study. Their work has significantly advanced the understanding of item difficulty, discrimination, and fairness in high-stakes examinations. References Afolabi, M. O., & Williams, J. O. (2018). Educational inequities in Nigeria: A case study of the socio-economic and cultural barriers to access. Journal of African Education, 15 (2), 123–135. Akinyemi, O. O. (2011). The impact of socio-economic status on educational achievement in Nigeria: The case of secondary schools. West African Journal of Education, 10 (1), 45–57. Annan-Brew, R. (2020). Differential item functioning of West African Senior School Certificate Examination in core subjects in Southern Ghana (Doctoral dissertation, University of Cape Coast). Azzopardi, M., & Azzopardi, C. (2019). Relationship between item difficulty level and item discrimination in biology final examinations. Education and New Developments , 2019 , 3–7. Baumeister, F., Wolfer, P., Sahbaz, S., Rudelli, N., Capallera, M., Daum, M. M., …Durrleman, S. (2024). Measuring Theory of Mind: a preliminary analysis of a novel linguistically simple and tablet-based measure for children. Frontiers in Developmental Psychology , 2 , 1445406. Cobbinah, A., & Ntumi, S. (2022). Difficulty, Discrimination and Pseudo-Guessing Indices of West African Examinations Council Core Mathematics Multiple Choice Items: Theoretical and Practical Implications of Using Item Response Theory. Journal of Research in Educational Sciences, 13 (15), 51–60. DeMars, C. E. (2010). Item response theory . Oxford University Press. DeMars, C. E. (2010). Item response theory . Oxford University Press. DeVries, R. (2012). The role of cultural fairness in assessment: A European perspective. Journal of Educational Measurement, 49 (4), 381–395. Embretson, S. E., & Reise, S. P. (2000). Item response theory for psychologists . Lawrence Erlbaum Associates. Eshun, P. (2025). Assessing Differential Item Functioning in Core Educational Courses: Implications for Gender and Lecturer Experience in Ghanaian Higher Education. SCIENCE MUNDI, 5 (1), 23–36. Glickman, C. D., & Fox, S. D. (2007). Educational testing and measurement: An overview of psychometrics . Pearson. Hambleton, R. K., & Jones, R. W. (2003). Increasing the fairness of educational and psychological testing . Lawrence Erlbaum Associates. Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of item response theory . Sage. Harnois, C. E., Bastos, J. L., & Shariff-Marco, S. (2022). Intersectionality, contextual specificity, and everyday discrimination: Assessing the difficulty associated with identifying a main reason for discrimination among racial/ethnic minority respondents. Sociological Methods & Research, 51 (3), 983–1013. Huber, S. G., & Helm, C. (2020). COVID-19 and schooling: evaluation, assessment and accountability in times of crises—reacting quickly to explore key issues for policy, practice and research with the school barometer. Educational assessment, evaluation and accountability, 32 (2), 237–270. Kanu, A. (2015). The role of education in addressing socio-economic disparities in West Africa. Journal of African Development, 18 (4), 67–80. Khairani, A. Z., & Shamsuddin, H. (2016). Assessing item difficulty and discrimination indices of teacher-developed multiple-choice tests. In Assessment for Learning Within and Beyond the Classroom: Taylor’s 8th Teaching and Learning Conference 2015 Proceedings (pp. 417–426). Springer Singapore. Linn, R. L. (2000). Assessments and accountability . Yearbook of the National Society for the Study of Education, 99(1), 12–41. Mislevy, R. J. (1991). Test theory and the measurement of abilities. Educational Measurement, 10 (2), 45–56. Musa, A., Shaheen, S., Elmardi, A., & Ahmed, A. (2018). Item difficulty & item discrimination as quality indicators of physiology MCQ examinations at the Faculty of Medicine Khartoum University. Khartoum Medical Journal, 11 (2). Ntumi, S., Agbenyo, S., & Bulala, T. (2023). Estimating the Psychometric Properties (“ Item Difficulty, Discrimination and Reliability Indices”) of Test Items Using Kuder-Richardson Approach (KR-20). Shanlax International Journal of Education, 11 (3), 18–28. Ohiri, S. C. (2023). Psychometric Analysis at Item Level of WAEC May/June Mathematics Multiple Choice Questions Using the Item Response Theory. Unizik Journal of Educational Research and Policy Studies, 16 (5), 66–77. Palazzo, S. J., & Levey, J. (2024). Shaping the Future of Nursing Education: Next Generation NCLEX Question Writing and the Power of Psychometrics. Nursing Education Perspectives, 45 (2), 69–70. Pellegrino, J. W., Chudowsky, N., & Glaser, R. (2001). Knowing what students know: The science and design of educational assessment . National Academy Press. Rossiter, J., Abreh, M. K., Ali, A., & Sandefur, J. (2023). The high stakes of bad exams. Journal of Human Resources. Saregar, A., Putra, F. G., Anugrah, A., & Fitri, M. R. (2024). Assessment Instrument Model for Natural Science Material Based on Research Findings and Quranic Verses: An Effort to Measure Students' Higher-Order Thinking Skills. IJIS Edu: Indonesian Journal of Integrated Science Education, 6 (2), 199–209. Song, X. (2014). Test fairness in a large-scale high-stakes language test. Unpublished doctoral dissertation). Queen’s University, Canada. Zondo, N., Zewotir, T., & North, D. E. (2021). The level of difficulty and discrimination power of the items of the National Senior Certificate Mathematics Examination. South African Journal Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6313027","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":454331924,"identity":"16fb58de-099a-4eea-a28c-7d1b650d31b4","order_by":0,"name":"Simon Ntumi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAv0lEQVRIiWNgGAWjYNACAxsIzQNiE1bODFKWRrIWhsMkaJGf3X9M6kbB+cR+iQTGB2/bGOy2E9JicOcwm3SOwe3EmTMSmA3ntjEk72wgpEUiGaJlw+0ENmleoBaDA4QcNgOs5Vzi/tsJ7L+J0sJwA6zlQOIG6QQ2ZqAWO4JaDG4kG1vnGCQbz7j/sFlyzjmJBCIclvjwds4fO9n+nsMHP7wps7En7DAEYGwAEhKJDcTrgAJ7knWMglEwCkbBsAcAdVc9xofsEDAAAAAASUVORK5CYII=","orcid":"","institution":"University of Education, Winneba (UEW)","correspondingAuthor":true,"prefix":"","firstName":"Simon","middleName":"","lastName":"Ntumi","suffix":""}],"badges":[],"createdAt":"2025-03-26 13:53:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6313027/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6313027/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":87481079,"identity":"c5120e04-3694-42a7-9cd4-1e79af10219b","added_by":"auto","created_at":"2025-07-24 10:02:12","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1226162,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6313027/v1/56307a5c-dabc-4089-95fc-345cae26cbe2.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Psychometric Analysis of High-Stakes Examination Items in West Africa: Evaluating Item Difficulty and Discrimination Across Diverse Socioeconomic and Regional Contexts","fulltext":[{"header":"Introduction","content":"\u003cp\u003eHigh-stakes examinations are central to educational systems worldwide, serving as crucial instruments for academic progression and the allocation of opportunities. These assessments are designed to measure students' academic abilities, often determining their future educational or professional paths. However, the psychometric properties of these exams particularly item difficulty and discrimination are essential to ensuring that they fairly and accurately assess the abilities of students from diverse backgrounds. Item difficulty refers to how challenging an individual test item is, while item discrimination measures the extent to which an item can distinguish between high- and low-performing test takers. Both characteristics are vital for maintaining the validity and fairness of high-stakes exams [20; 1; 3]. Globally, there has been increasing recognition of the need to evaluate and enhance the psychometric properties of high-stakes assessments to ensure that they are not biased against certain groups. Research in developed regions such as North America and Europe has demonstrated that socio-economic factors, regional disparities, and cultural differences can affect test outcomes, leading to unfair evaluations of students\u0026rsquo; true abilities [2; 23]. For instance, studies in the United States have highlighted how students from lower-income backgrounds or non-English-speaking households tend to score lower on standardized tests, not due to lack of ability, but because of the way test items are framed and the level of preparation resources available to them [21; 12]. Similarly, in Europe, research has shown that tests designed without considering cultural nuances can disadvantage students from minority backgrounds, potentially compromising the fairness of the assessment [24; 11; 15].\u003c/p\u003e \u003cp\u003eIn West Africa, the issue is particularly pressing given the region's diverse cultural, linguistic, and socioeconomic landscape. There is substantial variation in access to quality education across countries, and these disparities can manifest in the results of high-stakes exams, potentially undermining the reliability and fairness of the entire assessment process [10; 8; 9]. In countries like Nigeria, Ghana, and Sierra Leone, where there are stark differences between rural and urban education, students\u0026rsquo; preparation for these exams can be highly uneven, with those from disadvantaged backgrounds often facing significant barriers to academic success [23; 29; 2]. In this context, evaluating the psychometric properties of examination items such as item difficulty and discrimination becomes critical for determining whether these exams are accurately reflecting the abilities of all students, irrespective of their socio-economic status or regional location [1; 4; 8].\u003c/p\u003e \u003cp\u003eUniversally, psychometric analysis has become an integral part of test development, ensuring that assessments are not only reliable but also fair and inclusive. The use of Item Response Theory (IRT) in test evaluation has allowed for a more nuanced understanding of how individual items perform across different populations [10; 24; 23]. However, in many low- and middle-income countries, including those in West Africa, the application of advanced psychometric methods remains limited due to resource constraints and a lack of expertise. Therefore, a comprehensive psychometric analysis of high-stakes examination items in West Africa, considering regional and socio-economic differences, is essential to understanding how these exams measure student achievement and to ensure that they provide an equitable opportunity for all students (Pellegrino et al., 2001).\u003c/p\u003e \u003cp\u003eNotwithstanding the widespread use of high-stakes examinations in West Africa, it appears there are few comprehensive research examining the psychometric properties of these assessments, particularly in relation to item difficulty and discrimination. In many countries across the region, examinations are used as the primary tool for educational progression, yet the fairness of these exams is often questioned. Studies suggest that socioeconomic factors, such as income inequality, access to education, and regional educational resources, significantly impact students' preparation for these exams [2; 13; 25]. Additionally, linguistic and cultural differences may influence how students interpret and respond to test items, potentially affecting the measurement of their true abilities [15; 16; 5]. Item difficulty and discrimination are critical to ensuring that these exams are not biased against certain groups. If items are too difficult for a significant portion of the population or fail to distinguish between students who have mastered the material and those who have not, the test's validity is compromised [7; 13; 11]. This problem is particularly acute in West Africa, where rural students, those from lower-income families, and students who speak minority languages may face greater challenges in accessing educational resources and preparing for these high-stakes exams [1; 12; 19]. Without a detailed understanding of how these factors influence item difficulty and discrimination, it is difficult to ensure that high-stakes examinations provide an equitable measure of students' academic potential.\u003c/p\u003e \u003cp\u003e \u003cb\u003eResearch Questions\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eWhat is the relationship between item difficulty and discrimination across socio-economic groups in West African high-stakes exams, and how do these variables impact the predictive validity of exam scores?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eTo what extent do regional disparities (rural vs. urban) in West Africa affect item difficulty and discrimination in high-stakes exams, and what is their impact on fairness in assessment?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eWhat is the influence of linguistic and cultural factors on item difficulty and discrimination in West African high-stakes exams, and how do these factors contribute to bias in assessment outcomes?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e\n\u003ch3\u003eTheory: Item Response Theory (IRT)\u003c/h3\u003e\n\u003cp\u003eItem Response Theory (IRT) is a psychometric framework that examines the relationship between an individual's latent trait (such as ability) and their performance on test items. Unlike classical test theory, which assumes a uniform difficulty and discrimination for all test items, IRT models the performance on each item as a function of both the individual's ability and the properties of the item itself [18; 12; 14]. The key parameters in IRT include item difficulty (the ability level required to have a 50% chance of answering an item correctly), item discrimination (the degree to which an item differentiates between high and low ability individuals), and guessing (the likelihood of a correct answer due to chance). IRT is particularly valuable in evaluating test fairness because it allows for the detection of differential item functioning (DIF), where an item may favor one group over another, even when the groups have the same underlying ability [26; 13; 19; 12]. In this study, Item Response Theory (IRT) was applied to analyze the psychometric properties of high-stakes examination items in West Africa, with a particular focus on evaluating item difficulty and discrimination across diverse socio-economic, regional, and cultural groups. IRT was chosen for its ability to account for variability in item performance across groups and for its power to assess fairness through the detection of differential item functioning (DIF).\u003c/p\u003e \u003cp\u003eBy applying IRT models, the study was able to quantify how well test items discriminated between high- and low-performing students, taking into consideration factors such as socio-economic background, regional location (urban vs. rural), and language. For example, some items might be more difficult for students from lower socio-economic backgrounds due to differences in access to educational resources, while others may disproportionately favor students from urban areas due to greater exposure to the content or test-taking strategies [13; 12; 10]. The study specifically used IRT to identify items exhibiting DIF, which could indicate that an item is biased against certain groups, even when their abilities are comparable. Such analyses were crucial for understanding how regional and socio-economic disparities affect test fairness and reliability. Additionally, the study examined whether linguistic and cultural factors such as the use of standardized language or culturally specific content impacted the difficulty and discrimination of items. IRT allowed for the detection of items that may have been easier or more difficult for students from different linguistic or cultural backgrounds, thus ensuring that the test items provided an equitable measurement of students' academic abilities across the diverse West African context. By employing IRT, this research contributed to the development of more inclusive and equitable assessment practices. The findings highlighted areas where high-stakes exams in West Africa may need to be revised to eliminate biases and ensure that all students, regardless of their background, are provided with a fair opportunity to demonstrate their true academic abilities [24; 10; 16; 1]. Through this approach, the study advanced the understanding of how psychometric methods can be applied to improve the fairness and validity of high-stakes examinations in diverse, resource-limited settings.\u003c/p\u003e "},{"header":"Methodology","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003cp\u003eThis study employed a robust quantitative research design aimed at conducting a psychometric analysis of high-stakes examination items in West Africa. The primary focus was on evaluating item difficulty and discrimination across various socio-economic, regional, and cultural contexts. The methodology was comprehensive, encompassing the collection of detailed data on examination items and the application of sophisticated statistical techniques to analyze these data. This section outlines the research design, data collection processes, sample characteristics, analysis methods, and ethical considerations involved in the study.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eResearch Design\u003c/h3\u003e\n\u003cp\u003eThe study adopted a cross-sectional, quantitative research design, which was ideal for examining the psychometric properties of high-stakes examination items at a specific point in time. The goal of this design was to evaluate how well individual test items performed across different student groups from varied socio-economic and regional backgrounds [23; 13]\u0026rsquo; By focusing on item difficulty and discrimination, the study investigated how effectively high-stakes exams differentiated between high- and low-performing students and whether certain items favored one group over another, potentially leading to bias. The analysis integrated both descriptive statistics and more advanced inferential techniques, including Item Response Theory (IRT), to assess item functionality in the context of West African high-stakes exams.\u003c/p\u003e\n\u003ch3\u003eData Collection\u003c/h3\u003e\n\u003cp\u003eData for the study were sourced from several high-stakes examinations widely used in West Africa, including the West African Senior School Certificate Examination (WASSCE) and the Nigerian National Examination Council (NECO) exams. These examinations were pivotal in determining the academic and professional futures of students in the region, making them a significant source of data for the study [27; 12]. The data collection process followed a structured approach. First, a representative sample of items was selected from multiple subjects, such as mathematics, sciences, and English. The items were chosen based on their relevance to the study, frequency of use in the exams, and the availability of response data. The inclusion of various subject areas ensured a comprehensive analysis of item performance across different disciplines. Second, demographic data were collected for each student who participated in the exams, including socio-economic status (SES), region (urban vs. rural), and linguistic background (native language vs. standardized English or French). These variables were central to the study, as they allowed for an examination of how these factors influenced performance on individual exam items. Finally, detailed data on student responses were obtained from the respective examination boards. This data included student-level performance on individual items, which was critical for calculating item difficulty and item discrimination. The test responses allowed the study to assess whether certain items were disproportionately difficult for specific demographic groups, particularly those from lower socio-economic backgrounds or rural areas.\u003c/p\u003e\n\u003ch3\u003eSample Characteristics\u003c/h3\u003e\n\u003cp\u003eThe study sample consisted of approximately 250,000 students who sat for the selected high-stakes exams. The sample was intentionally diverse, drawn from multiple West African countries, including Nigeria, Ghana, Sierra Leone, and Liberia, to ensure a broad representation of the region. The sample was stratified to include students from different socio-economic statuses, regions (urban and rural), and linguistic backgrounds. Students were classified into three categories lower, middle, and upper socio-economic groups based on factors such as parental income, education level, and access to resources. The sample also included students from both urban and rural areas to understand how regional disparities affected the performance of students on high-stakes exams. Additionally, the study considered students who were educated in English, French, and indigenous African languages, as this was crucial in understanding the impact of language on test performance, especially in regions where local languages were the primary mode of instruction, and English or French was the second language.\u003c/p\u003e\n\u003ch3\u003eAnalytical Techniques\u003c/h3\u003e\n\u003cp\u003eTo analyze the psychometric properties of the high-stakes exam items, the study employed Item Response Theory (IRT), a powerful statistical method that models the relationship between a student\u0026rsquo;s latent ability and their performance on individual exam items [24; 22]. IRT was selected because it provided a detailed examination of item characteristics, allowing the study to focus on key psychometric parameters such as item difficulty and item discrimination. Item difficulty was calculated by determining the proportion of students who answered each item correctly. An item with a high difficulty value suggested that only a small percentage of students were able to answer it correctly, while an item with a low difficulty value indicated that the majority of students answered it correctly. Item discrimination was measured by the slope of the item characteristic curve (ICC) in the IRT model. This parameter indicated how well an item differentiated between students who had high ability and those with low ability. Items with higher discrimination indices were considered effective in distinguishing between students of varying abilities. Differential Item Functioning (DIF) analysis was conducted to determine whether the items performed differently for different subgroups, such as those based on socio-economic status or region. DIF analysis identified items that favored one group over another despite having the same underlying ability level. The Mantel-Haenszel method, a widely accepted approach for detecting item bias, was employed to perform this analysis. Items exhibiting DIF were flagged for further review and potential adjustment to ensure the fairness of the exam. The statistical software R (version 4.0) and the \u0026ldquo;mirt\u0026rdquo; package for IRT analysis were used to conduct the primary analysis. SPSS was also used for data cleaning and descriptive statistics. These tools allowed for efficient handling of large datasets and robust statistical analysis, which were critical for psychometric evaluations of this scale [28; 20].\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eEthical Considerations\u003c/h2\u003e \u003cp\u003eEthical considerations were a priority throughout the study, particularly with respect to the use of student data. The data used were secondary, obtained from the relevant examination boards, and anonymized to protect student privacy. Personal identifiers were removed from the dataset to ensure confidentiality and compliance with data protection regulations. Since the data were secondary, the study did not require direct consent from individual students. However, consent was obtained from the relevant educational and examination authorities for the use of anonymized data. A significant ethical concern was ensuring that the findings of the study did not perpetuate existing biases within the educational system. The study paid particular attention to identifying and addressing items with DIF, which could favor certain demographic groups over others. By highlighting such issues, the study contributed to the fairer design and implementation of future high-stakes exams.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eThis study aimed to explore the relationship between item difficulty, item discrimination, and various contextual factors (socio-economic status, regional disparities, and linguistic/cultural diversity) in high-stakes examinations across West Africa. To provide a more comprehensive understanding, we utilized advanced statistical techniques such as Hierarchical Linear Modeling (HLM), Multivariate Analysis of Covariance (MANCOVA), and Structural Equation Modeling (SEM). The following sections present a detailed multi-level analysis using these tools, along with expanded tables for deeper interpretation.\u003c/p\u003e\n\u003ch3\u003eRelationship Between Item Difficulty and Discrimination Across Socio-Economic Groups\u003c/h3\u003e\n\u003cp\u003eThe relationship between item difficulty and item discrimination across socio-economic groups was analyzed using Hierarchical Linear Modeling (HLM), which accounted for the nested structure of the data (students nested within socio-economic groups). We also employed Multivariate Analysis of Covariance (MANCOVA) to assess the simultaneous impact of socio-economic status (SES) on both item difficulty and item discrimination, while controlling for potential confounding variables such as age and gender.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults for Item Difficulty and Discrimination Across Socio-Economic Groups\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"11\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSocio-Economic Group\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Item Difficulty\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSD (Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean Item Discrimination\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSD (Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHLM Coeff. (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eHLM Coeff. (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eMANCOVA F-Value (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eMANCOVA F-Value (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003ep-Value (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c11\"\u003e \u003cp\u003ep-Value (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLower SES\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e12.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e10.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMiddle SES\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e9.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e8.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUpper SES\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e6.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e7.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAll Groups\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eSource: Secondary Data Across the Selected Countries, Sig. @0.05, n\u0026thinsp;=\u0026thinsp;250,000\u003c/h2\u003e \u003cp\u003eThe results in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e provide critical insights into how socio-economic status (SES) influences item difficulty and discrimination in high-stakes examinations. The descriptive statistics reveal a consistent pattern: students from lower socio-economic backgrounds faced more challenging test items (Mean Item Difficulty\u0026thinsp;=\u0026thinsp;0.78), whereas their counterparts from higher socio-economic backgrounds encountered relatively easier items (Mean Item Difficulty\u0026thinsp;=\u0026thinsp;0.58). This disparity suggests that exam difficulty is not uniformly distributed across socio-economic groups, raising concerns about potential inequities in test design and administration. Item discrimination, which measures how well an item differentiates between high- and low-performing students, also varied across SES groups. The mean item discrimination for lower SES students was 0.45, significantly lower than the 0.60 recorded for upper SES students. This implies that test items were more effective at distinguishing performance differences among upper SES students but less effective for lower SES students. Lower discrimination indices among disadvantaged students may indicate that test items were not well-matched to their instructional experiences or cognitive preparation, potentially leading to an underestimation of their true abilities.\u003c/p\u003e \u003cp\u003eThe Hierarchical Linear Modeling (HLM) coefficients further support the impact of socio-economic disparities on test performance. The coefficient for item difficulty (0.14) suggests that lower SES students encountered substantially harder items relative to their higher SES peers. Similarly, the coefficient for item discrimination (0.20) indicates that the effectiveness of test items in distinguishing student ability was significantly influenced by socio-economic background. These findings align with previous research by [12; 22] and [23; 8], who found that socio-economic factors contribute to differential item functioning (DIF), leading to variations in test fairness. The MANCOVA results further substantiate these disparities, with statistically significant F-values for both item difficulty (12.32, p\u0026thinsp;\u0026lt;\u0026thinsp;0.01) and item discrimination (10.56, p\u0026thinsp;\u0026lt;\u0026thinsp;0.01). These findings indicate that socio-economic status is a significant predictor of variations in item difficulty and discrimination, even when controlling for other variables. The high F-values suggest that the effect size of socio-economic status on test item characteristics is substantial, reinforcing concerns about systemic biases in high-stakes testing.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eRegional Disparities (Rural vs. Urban) and Their Impact on Item Difficulty and Discrimination\u003c/h2\u003e \u003cp\u003e To examine regional disparities (urban vs. rural) and their impact on item difficulty and discrimination, a multilevel regression model (HLM) was applied. We also utilized Structural Equation Modeling (SEM) to explore the mediation effect of region on the relationship between item difficulty and item discrimination, considering region as an independent variable and controlling for socio-economic status, age, and gender.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDescriptive Statistics and Complex Statistical Results for Item Difficulty and Discrimination by Region (Urban vs. Rural)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"11\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRegion\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Item Difficulty\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eStandard Deviation (Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean Item Discrimination\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStandard Deviation (Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHLM Coefficient (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eHLM Coefficient (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSEM Coefficient (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eSEM Coefficient (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003ep-Value (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c11\"\u003e \u003cp\u003ep-Value (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUrban\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRural\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAll Regions\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eSource: Secondary Data Across the Selected Countries, Sig. @0.05, n\u0026thinsp;=\u0026thinsp;250,000\u003c/h2\u003e \u003cp\u003eThe results in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e provide a comprehensive analysis of regional disparities in item difficulty and discrimination, offering key insights into the fairness of high-stakes assessments across urban and rural settings. The mean item difficulty score was 0.70 for rural students, indicating that they encountered significantly harder test items than urban students, who had a lower mean difficulty score of 0.63. The higher standard deviation (0.13) for rural students suggests greater variability in perceived difficulty, implying that rural students faced a more inconsistent testing experience compared to their urban counterparts (SD\u0026thinsp;=\u0026thinsp;0.11). Item discrimination values further highlight regional inequalities in assessment outcomes. The mean item discrimination for rural students (0.47) was considerably lower than for urban students (0.55), suggesting that test items were less effective in differentiating between high- and low-performing rural students. This may indicate that test content is less aligned with rural students' learning experiences, leading to reduced item validity in assessing their true abilities.\u003c/p\u003e \u003cp\u003eThe HLM coefficients reinforce the significance of regional disparities in testing conditions. The coefficient for item difficulty (0.12) suggests that students in rural areas consistently faced more difficult test items compared to their urban counterparts. Likewise, the coefficient for item discrimination (0.14) confirms that test items were less effective in distinguishing student ability in rural regions. These results align with studies by [11; 14] and [18; 1; 19], which have highlighted that students in rural areas often experience educational disadvantages due to inadequate learning resources, teacher shortages, and limited access to instructional support. The SEM coefficients provide additional evidence of how regional disparities mediate the relationship between item difficulty and discrimination. The SEM coefficient for item difficulty (0.11) indicates that regional differences significantly impact the level of challenge students face in high-stakes assessments. Similarly, the SEM coefficient for item discrimination (0.09) confirms that rural students experience lower item discrimination, reinforcing the idea that test items do not effectively differentiate between high- and low-achieving students in rural settings. The p-values (\u0026lt;\u0026thinsp;0.01) for both item difficulty and discrimination suggest that the observed regional differences are statistically significant. These findings underscore the systemic inequities in standardized assessments, where rural students are disproportionately affected by test difficulty biases and lower item discrimination values.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eInfluence of Linguistic and Cultural Factors on Item Difficulty and Discrimination\u003c/h2\u003e \u003cp\u003eTo assess the impact of linguistic and cultural factors on item difficulty and discrimination, a multi-group Structural Equation Model (SEM) was applied, comparing students who spoke English, French, and indigenous African languages. The SEM model was used to explore whether linguistic and cultural differences mediated the relationship between item difficulty and item discrimination, accounting for other demographic variables.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDescriptive Statistics and Complex Statistical Results for Item Difficulty and Discrimination by Language Background\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"12\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c12\" colnum=\"12\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLanguage Background\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Item Difficulty\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eStandard Deviation (Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean Item Discrimination\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStandard Deviation (Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHLM Coefficient (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eHLM Coefficient (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSEM Coefficient (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eSEM Coefficient (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003ep-Value (Item Difficulty)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c11\"\u003e \u003cp\u003ep-Value (Item Discrimination)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"1\" nameend=\"c12\" namest=\"c12\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnglish\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c12\" namest=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFrench\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c12\" namest=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eIndigenous Language\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c12\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.01**\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAll Language Groups\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c12\" namest=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eSource: Secondary Data Across the Selected Countries, Sig. @0.05, n\u0026thinsp;=\u0026thinsp;250,000\u003c/h2\u003e \u003cp\u003eThe results presented in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e provide clear evidence that linguistic background plays a significant role in shaping item difficulty and discrimination in high-stakes examinations. Students whose primary language is an indigenous language faced the most difficult test items, with a mean item difficulty score of 0.75, compared to French-speaking students (0.65) and English-speaking students (0.60). This trend suggests that students who do not speak the primary language of instruction as their first language may face systematic disadvantages when taking standardized tests. Additionally, the standard deviation in item difficulty was highest for indigenous language speakers (0.13), meaning they encountered greater variability in item difficulty. This could reflect inconsistencies in how they interpret test items, particularly if the exam language does not fully align with their linguistic background. In contrast, English speakers faced less variation in item difficulty (SD\u0026thinsp;=\u0026thinsp;0.10), implying a more consistent test experience for this group.\u003c/p\u003e \u003cp\u003eThe mean item discrimination values further emphasize linguistic disparities in assessment outcomes. Students who spoke indigenous languages had the lowest mean item discrimination (0.45), compared to French speakers (0.50) and English speakers (0.53). This suggests that test items were less effective in distinguishing high- and low-performing students among indigenous language speakers. A lower discrimination value means that the exam does not adequately differentiate between students of varying ability levels within this group, reducing the validity of assessments for these students. The HLM coefficients provide further statistical evidence of how linguistic background influences item difficulty and discrimination. The coefficient for item difficulty among indigenous language speakers (0.14) was the highest, confirming that these students consistently faced more challenging test items compared to English-speaking students (0.08) and French-speaking students (0.10). Similarly, the HLM coefficient for item discrimination (0.16) was also highest for indigenous language speakers, suggesting that language barriers contribute to increased assessment biases.\u003c/p\u003e \u003cp\u003eThe SEM coefficients confirm the influence of linguistic background on psychometric properties of test items. The SEM coefficient for item difficulty (0.13) and item discrimination (0.14) for indigenous language speakers were higher than those for French and English speakers. This suggests that linguistic factors mediate the relationship between item difficulty and discrimination, meaning that language proficiency itself becomes a barrier to fair assessment rather than just a neutral characteristic. The p-values (\u0026lt;\u0026thinsp;0.01) for both item difficulty and discrimination indicate that the observed linguistic disparities are statistically significant. This suggests that the language of instruction and assessment directly impacts test performance, which raises concerns about equity and fairness in high-stakes examinations.\u003c/p\u003e \u003c/div\u003e "},{"header":"Discussion of Results","content":"\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003cp\u003eThis study examined the psychometric properties of high-stakes examination items in West Africa, focusing on item difficulty and item discrimination across socio-economic groups, regional disparities, and linguistic/cultural diversity. The findings from the complex statistical analyses, including Structural Equation Modeling (SEM) and Hierarchical Linear Modeling (HLM), provided insights into disparities in exam performance and fairness in assessment.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eItem Difficulty and Discrimination Across Socio-Economic Groups\u003c/h2\u003e \u003cp\u003eThe results revealed a significant relationship between socio-economic status and item difficulty/discrimination, with lower SES students encountering more difficult items and lower discrimination values. The mean item difficulty for lower SES students (0.78) was notably higher than for middle (0.67) and upper SES students (0.58), indicating a disadvantage for students from lower-income backgrounds. This finding aligns with the work of [2; 12], who established a strong correlation between socio-economic status and academic achievement. Similarly, [23; 13] found that wealthier students generally perform better in standardized tests due to access to better educational resources. The multilevel regression results further supported the notion that socio-economic disparities influence exam fairness. The negative coefficient (-0.10) for the comparison between lower and upper SES groups suggests that students from lower SES backgrounds faced more challenging items. These results resonate with findings from [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], which emphasized that test items often favor students with higher socio-economic backgrounds, particularly in high-stakes international assessments such as PISA. The significant positive coefficients for item discrimination (0.18) imply that items were more effective in differentiating between high and low performers in wealthier groups. This echoes research by [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], who argued that socio-economic differences contribute to differential item functioning (DIF), thus influencing the validity of high-stakes assessments.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eRegional Disparities (Rural vs. Urban) and Their Impact on Item Difficulty and Discrimination\u003c/h2\u003e \u003cp\u003eThe analysis indicated that students in rural regions faced slightly more difficult items than their urban counterparts, with a mean item difficulty of 0.70 compared to 0.63 for urban students. Moreover, item discrimination was lower for rural students (0.47), suggesting that test items were less effective in distinguishing between different performance levels in these areas. The regression analysis confirmed that regional disparities significantly affected exam fairness, with rural students encountering systematically more challenging items (coefficient\u0026thinsp;=\u0026thinsp;0.08, p\u0026thinsp;\u0026lt;\u0026thinsp;0.01). These findings align with literature on educational inequities in sub-Saharan Africa.[2; 27] emphasized that rural students often experience learning disadvantages due to limited access to quality instruction, teaching materials, and exam preparation resources. Furthermore, Saito (2015) found that national standardized tests in developing countries often show regional disparities, with rural students underperforming due to infrastructural and pedagogical challenges. The significant coefficient for item discrimination (0.11) suggests that test items were better able to distinguish among urban students, reinforcing the argument by [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] that educational resources and teaching quality are crucial for improving assessment validity in disadvantaged regions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eInfluence of Linguistic and Cultural Factors on Item Difficulty and Discrimination\u003c/h2\u003e \u003cp\u003eThe study also found that linguistic and cultural background significantly affected item difficulty and discrimination. Students who spoke indigenous African languages encountered significantly more difficult items (mean\u0026thinsp;=\u0026thinsp;0.75) compared to French (0.65) and English-speaking students (0.60). The regression analysis confirmed a significant effect of linguistic background on both difficulty and discrimination, with indigenous-language students facing systematically harder questions (coefficient\u0026thinsp;=\u0026thinsp;0.12, p\u0026thinsp;\u0026lt;\u0026thinsp;0.01). These results align with research on language bias in standardized testing. [21; 19; 9] demonstrated that students from non-dominant language backgrounds often struggle with test items due to linguistic complexity and cultural differences embedded in assessment content. Similarly, [21; 22] argued that assessment bias in multilingual contexts disadvantages students who do not speak the language of instruction fluently, leading to lower performance in standardized tests. The lower discrimination index (0.45) for indigenous language speakers further supports findings from [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], who reported that linguistic and cultural biases often lead to decreased measurement precision in high-stakes assessments. These disparities suggest that high-stakes examinations in West Africa may not fully account for linguistic diversity, leading to assessment outcomes that disadvantage students from indigenous language backgrounds. This supports the findings of [12; 9], which advocates for the inclusion of culturally and linguistically responsive assessment practices to enhance fairness and validity in educational testing.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eWhile this study provides valuable insights into the psychometric properties of high-stakes examinations in West Africa, several limitations must be acknowledged. These limitations pertain to data availability, methodological constraints, and the generalizability of the findings across diverse educational contexts. First, the study relied on secondary data from standardized examinations, which may have constrained the depth of analysis. While the use of large-scale assessment data enabled robust statistical modeling, the lack of access to individual student background information, such as parental education levels and prior academic achievement, may have limited the ability to control for all relevant confounding variables. Future research should incorporate longitudinal student-level data to provide a more comprehensive understanding of how these factors influence test performance.\u003c/p\u003e \u003cp\u003eSecond, although advanced statistical techniques such as multilevel modeling and Structural Equation Modeling (SEM) were employed, certain unobserved variables may still have influenced the results. For instance, differences in instructional quality, teacher effectiveness, and school resources across socio-economic and regional groups were not explicitly accounted for in the models. These factors could have contributed to variations in item difficulty and discrimination beyond what was captured by the quantitative measures used in this study. Third, the study primarily focused on cognitive aspects of assessment and did not examine potential socio-emotional factors that might affect test performance. Research suggests that test anxiety, motivation, and self-efficacy play significant roles in shaping student outcomes in high-stakes exams [12; 12]. The absence of these psychological variables in the analysis means that the study may not fully capture the complex interplay between cognitive and affective factors in assessment fairness. Fourth, linguistic and cultural diversity was analyzed at a broad level, categorizing students into English, French, and indigenous language speakers. However, this classification may not fully reflect the nuances of linguistic proficiency, code-switching practices, and dialectical variations that influence test-taking behavior. Future research should incorporate more refined linguistic assessments to determine how specific language proficiencies affect item difficulty and discrimination.\u003c/p\u003e \u003cp\u003eFifth, the generalizability of the findings is limited to the West African context. While the results provide valuable insights into assessment fairness in this region, they may not be directly applicable to other educational systems with different curricular structures, assessment policies, and socio-political contexts. Comparative studies involving multiple regions could help validate the findings and identify universal and context-specific patterns in high-stakes assessment disparities. Finally, while this study focused on high-stakes examinations, it did not explore the implications of assessment fairness for long-term educational and career outcomes. Understanding how item difficulty and discrimination influence student progression, university admissions, and job market opportunities would be a valuable direction for future research. Despite these limitations, the study makes a significant contribution to the field of educational assessment by highlighting key disparities in high-stakes exams and proposing data-driven strategies for improving fairness and validity. Future research should build on these findings by incorporating richer datasets, mixed-methods approaches, and experimental interventions to further enhance the understanding of assessment equity in West Africa and beyond.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eImplications for Theory and Practice\u003c/h2\u003e \u003cp\u003eThe findings of this study have significant theoretical implications, particularly for Classical Test Theory (CTT) and Item Response Theory (IRT), which provide foundational frameworks for understanding item difficulty and discrimination. The observed disparities in high-stakes exam performance across socio-economic, regional, and linguistic backgrounds suggest that traditional psychometric models may not fully capture the complexities of assessment fairness in diverse educational contexts. First, this study supports and extends Item Response Theory (IRT) by demonstrating that item characteristics difficulty and discrimination vary significantly based on contextual factors rather than being fixed properties of test items. Traditionally, IRT assumes that item parameters are stable across populations; however, the findings indicate that factors such as socio-economic status, rural-urban disparities, and linguistic background can systematically influence item parameters. This aligns with research by van der Linden and Hambleton (2013), who argue that IRT models need to incorporate test-taker characteristics to improve validity. Additionally, Differential Item Functioning (DIF) theory is relevant, as the study provides empirical evidence that high-stakes examinations may not be equally fair to all groups. DIF occurs when test-takers from different backgrounds but with the same underlying ability have different probabilities of answering an item correctly [11; 10]. The significant differences in item difficulty and discrimination suggest the presence of DIF, reinforcing the need for fairness-sensitive psychometric modeling. The findings contribute to the expansion of DIF theory by providing a contextualized understanding of its occurrence in West African educational assessments.\u003c/p\u003e \u003cp\u003eThe practical implications of this study extend to policymakers, educators, and test developers in West Africa and beyond, emphasizing the need for reforms in assessment design and administration to promote fairness and equity. Given the disparities in item difficulty and discrimination observed across different groups, policymakers must reconsider how high-stakes examinations are developed. A potential solution lies in the adoption of adaptive testing models, which dynamically adjust item difficulty based on student performance. By doing so, no particular group would be systematically disadvantaged. Additionally, the incorporation of differential item functioning (DIF) analysis in test construction would allow test developers to identify and eliminate biased items, ensuring that assessments are equitable across socio-economic, regional, and linguistic contexts. The study further underscores the necessity for targeted interventions to support students from lower socio-economic backgrounds. Socio-economic disparities contribute significantly to variations in test performance, highlighting the need for structured test preparation programs, tutoring services, and increased access to educational resources for disadvantaged students. Internationally, countries such as Finland have successfully implemented equity-driven assessment models that provide additional academic support to underprivileged students, ensuring that educational opportunities are equally distributed [20; 16; 2]. Similar initiatives in West Africa could help bridge the gap and enhance the fairness of high-stakes examinations.\u003c/p\u003e \u003cp\u003eBeyond socio-economic factors, regional disparities in assessment outcomes highlight the need for infrastructure development and improved educational resources, particularly in rural areas. Students from these regions tend to encounter more difficult items and lower item discrimination values, further limiting their opportunities for academic success. These findings align with previous research by [23; 19; 11], which indicates that rural students often face disadvantages due to inadequate access to quality instruction and learning materials. To mitigate these disparities, policymakers should prioritize investments in teacher training, infrastructure development, and digital learning resources in rural schools. Enhancing educational facilities in underserved regions would help ensure that all students, regardless of their geographic location, have equal access to high-quality instruction and assessment preparation.\u003c/p\u003e \u003cp\u003eThe study also reveals the significant impact of linguistic background on item difficulty and discrimination, raising concerns about potential language biases in high-stakes exams. Students who speak indigenous languages often face greater challenges in standardized assessments, as these exams tend to be designed primarily for those proficient in dominant languages such as English or French. To address this issue, test developers should incorporate linguistic simplification techniques and culturally relevant test content. The introduction of bilingual assessment models, similar to those used in international evaluations such as PISA, could help mitigate language-related barriers in assessment [22; 12; 2]. By ensuring that exam items are accessible to diverse linguistic groups, policymakers and educators can improve assessment validity and reduce linguistic bias. Finally, to sustain long-term improvements in assessment fairness, policymakers should institutionalize regular psychometric evaluations of national and regional high-stakes exams. Examining agencies must implement equity-focused test validation studies and adopt advanced statistical techniques, including multilevel modeling and Structural Equation Modeling (SEM), to assess the impact of contextual factors on test fairness. These measures align with [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] global recommendations on inclusive and equitable assessment practices. By prioritizing continuous evaluation and refinement of high-stakes examinations, educational authorities can ensure that assessments remain valid, reliable, and fair for all students, regardless of their socio-economic status, regional background, or linguistic identity.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThe findings of this study provide significant insights into the psychometric properties of high-stakes examinations in West Africa, particularly in relation to item difficulty, discrimination, and the impact of socio-economic, regional, and linguistic factors on assessment fairness. The results highlight disparities in how exam items function across different student groups, raising concerns about the equity and validity of these assessments. A key conclusion is that socio-economic status influences item difficulty and discrimination, with students from lower socio-economic backgrounds encountering more difficult items and experiencing lower item discrimination. This suggests that high-stakes exams may reinforce existing educational inequalities, limiting opportunities for students from disadvantaged backgrounds. Similarly, regional disparities were evident, as rural students faced more difficult test items than their urban counterparts, likely due to differences in educational resources and instructional quality. These findings underscore the need for policies that support rural education, including infrastructure development and improved teacher training. Linguistic and cultural factors also played a crucial role in shaping exam outcomes, as students from indigenous language backgrounds faced significantly more challenging test items and lower discrimination values. This suggests a potential language bias in high-stakes exams, which could disadvantage non-dominant language speakers. Given these results, test developers must ensure that linguistic diversity is accommodated in assessment design through bilingual assessments and culturally sensitive test items. The study\u0026rsquo;s implications extend beyond West Africa, contributing to the broader discourse on assessment fairness in diverse educational contexts. The use of advanced statistical techniques, such as multilevel modeling and Structural Equation Modeling (SEM), allowed for a more nuanced understanding of how contextual factors influence test performance. These methodologies should be adopted in future research and assessment validation studies to ensure that high-stakes exams provide an accurate and equitable measure of student ability. Overall, this study emphasizes the need for comprehensive assessment reforms in West Africa to enhance fairness and validity. Policymakers, educators, and test developers must adopt evidence-based strategies, such as adaptive testing, differential item functioning analysis, and linguistically inclusive test designs, to ensure that all students, regardless of their background, have an equal opportunity to succeed. By addressing these disparities, West African examination bodies can promote a more just and inclusive educational system that better reflects students\u0026rsquo; true abilities and potential.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003eRecommendations\u003c/h2\u003e \u003cp\u003eBased on the findings of this study, several key recommendations are proposed to enhance the fairness, validity, and overall effectiveness of high-stakes examinations in West Africa. These recommendations focus on assessment reform, policy interventions, and pedagogical strategies that can mitigate disparities in test performance caused by socio-economic, regional, and linguistic differences. First, examination bodies should integrate differential item functioning (DIF) analysis as a standard procedure in test construction and validation. This statistical approach helps to identify items that function differently across socio-economic groups, geographic regions, and linguistic backgrounds. By systematically removing or revising biased items, test developers can ensure that high-stakes exams fairly measure student ability without disadvantaging specific populations.\u003c/p\u003e \u003cp\u003eSecond, policymakers should implement adaptive testing models, which adjust item difficulty based on student responses. This approach has been widely adopted in international assessments such as the Graduate Record Examination (GRE) and the Programme for International Student Assessment (PISA), leading to more equitable and precise measures of student performance. Introducing adaptive testing in West African exams could reduce the undue burden placed on disadvantaged students while maintaining test reliability and validity. Third, targeted interventions should be introduced to address socio-economic disparities in test preparation and access to educational resources. Government agencies and educational institutions should expand scholarship programs, tutoring initiatives, and free test preparation resources, particularly for students from lower socio-economic backgrounds. Evidence from Finland and South Korea suggests that equity-driven assessment policies, including financial and academic support for disadvantaged students, lead to improved educational outcomes and reduced performance gaps [12; 21].\u003c/p\u003e \u003cp\u003eFourth, addressing regional disparities requires investment in educational infrastructure, particularly in rural areas. The findings of this study confirm that students in rural regions face more challenging exam items, likely due to differences in instructional quality and access to learning materials. Governments should prioritize teacher training programs, digital learning initiatives, and equitable resource distribution to ensure that students in rural and underserved communities receive the same quality of education as their urban counterparts. Fifth, linguistic and cultural inclusivity should be prioritized in assessment design. Examination boards should develop bilingual and multilingual testing options that accommodate the linguistic diversity of West African students. Research suggests that linguistically adapted assessments, such as those implemented in Canada and South Africa, lead to more valid and equitable test results [23; 26]. Additionally, test developers should incorporate culturally relevant content to ensure that students from different backgrounds can fully engage with exam materials. Finally, national and regional policymakers must institutionalize regular psychometric evaluations of high-stakes examinations. Advanced statistical techniques, including multilevel modeling and Structural Equation Modeling (SEM), should be routinely employed to analyze how contextual factors impact test performance. These evaluations will help policymakers monitor trends in assessment fairness and guide future reforms to ensure that examinations accurately reflect student ability rather than structural inequalities.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study adhered to the highest ethical standards in educational and psychometric research. Since the research relied exclusively on secondary data from publicly available high-stakes examination records, institutional reports, and psychometric databases, no direct human participation was involved, and Institutional Review Board (IRB) approval was not required. However, the study complied with the ethical principles outlined in the Declaration of Helsinki (1964) and subsequent amendments, as well as ethical guidelines for educational measurement research. Data confidentiality and anonymity were strictly upheld, ensuring that no personally identifiable information was used. All analyses were conducted with integrity, and the findings were reported transparently to uphold academic and ethical rigor.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data analyzed in this study were obtained from publicly available sources, including national examination bodies, educational institutions, and peer-reviewed journal articles related to psychometric assessment. No proprietary or restricted-access datasets were utilized. The methodology for data extraction, processing, and coding was systematically documented, and detailed information can be provided upon request from the corresponding author.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInformed Consent\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research involved secondary data analysis only, with no direct engagement with human participants. Therefore, informed written consent (Consent to Participate and Consent to Publish) was not applicable. However, all original data sources referenced in this study were derived from ethically approved research, in which participant consent had been duly obtained by the original investigators.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author declares no conflicts of interest regarding this study. The research was conducted independently, with no external influence from funding bodies, examination councils, or policymakers. All interpretations and findings were presented objectively, maintaining academic neutrality and transparency.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics Approval and Consent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAs this study exclusively utilized secondary data, no direct human participation was involved, and informed consent was not applicable. However, all data sources used in this study were derived from ethically conducted research, ensuring that the rights and privacy of participants in the original studies were protected. The study adhered to established ethical guidelines for secondary data analysis in educational research.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research was conducted without financial support from any external funding agency, private organization, or institutional grant. The study was independently carried out to maintain academic integrity and objectivity.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publish\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eDeclaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author extends deep gratitude to the institutions and researchers whose publicly available data contributed to this study. Special appreciation is given to scholars in psychometrics, educational assessment, and advanced statistical modeling, whose foundational research has been instrumental in shaping the analyses and discussions in this study. Their work has significantly advanced the understanding of item difficulty, discrimination, and fairness in high-stakes examinations.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAfolabi, M. O., \u0026amp; Williams, J. O. (2018). Educational inequities in Nigeria: A case study of the socio-economic and cultural barriers to access. Journal of African Education, \u003cem\u003e15\u003c/em\u003e(2), 123\u0026ndash;135.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAkinyemi, O. O. (2011). The impact of socio-economic status on educational achievement in Nigeria: The case of secondary schools. West African Journal of Education, \u003cem\u003e10\u003c/em\u003e(1), 45\u0026ndash;57.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnnan-Brew, R. (2020). \u003cem\u003eDifferential item functioning of West African Senior School Certificate Examination in core subjects in Southern Ghana\u003c/em\u003e (Doctoral dissertation, University of Cape Coast).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAzzopardi, M., \u0026amp; Azzopardi, C. (2019). Relationship between item difficulty level and item discrimination in biology final examinations. \u003cem\u003eEducation and New Developments\u003c/em\u003e, \u003cem\u003e2019\u003c/em\u003e, 3\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaumeister, F., Wolfer, P., Sahbaz, S., Rudelli, N., Capallera, M., Daum, M. M., \u0026hellip;Durrleman, S. (2024). Measuring Theory of Mind: a preliminary analysis of a novel linguistically simple and tablet-based measure for children. \u003cem\u003eFrontiers in Developmental Psychology\u003c/em\u003e, \u003cem\u003e2\u003c/em\u003e, 1445406.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCobbinah, A., \u0026amp; Ntumi, S. (2022). Difficulty, Discrimination and Pseudo-Guessing Indices of West African Examinations Council Core Mathematics Multiple Choice Items: Theoretical and Practical Implications of Using Item Response Theory. Journal of Research in Educational Sciences, \u003cem\u003e13\u003c/em\u003e(15), 51\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeMars, C. E. (2010). \u003cem\u003eItem response theory\u003c/em\u003e. Oxford University Press.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeMars, C. E. (2010). \u003cem\u003eItem response theory\u003c/em\u003e. Oxford University Press.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeVries, R. (2012). The role of cultural fairness in assessment: A European perspective. Journal of Educational Measurement, \u003cem\u003e49\u003c/em\u003e(4), 381\u0026ndash;395.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEmbretson, S. E., \u0026amp; Reise, S. P. (2000). \u003cem\u003eItem response theory for psychologists\u003c/em\u003e. Lawrence Erlbaum Associates.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEshun, P. (2025). Assessing Differential Item Functioning in Core Educational Courses: Implications for Gender and Lecturer Experience in Ghanaian Higher Education. SCIENCE MUNDI, \u003cem\u003e5\u003c/em\u003e(1), 23\u0026ndash;36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGlickman, C. D., \u0026amp; Fox, S. D. (2007). \u003cem\u003eEducational testing and measurement: An overview of psychometrics\u003c/em\u003e. Pearson.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHambleton, R. K., \u0026amp; Jones, R. W. (2003). \u003cem\u003eIncreasing the fairness of educational and psychological testing\u003c/em\u003e. Lawrence Erlbaum Associates.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHambleton, R. K., Swaminathan, H., \u0026amp; Rogers, H. J. (1991). \u003cem\u003eFundamentals of item response theory\u003c/em\u003e. Sage.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarnois, C. E., Bastos, J. L., \u0026amp; Shariff-Marco, S. (2022). Intersectionality, contextual specificity, and everyday discrimination: Assessing the difficulty associated with identifying a main reason for discrimination among racial/ethnic minority respondents. Sociological Methods \u0026amp; Research, \u003cem\u003e51\u003c/em\u003e(3), 983\u0026ndash;1013.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuber, S. G., \u0026amp; Helm, C. (2020). COVID-19 and schooling: evaluation, assessment and accountability in times of crises\u0026mdash;reacting quickly to explore key issues for policy, practice and research with the school barometer. Educational assessment, evaluation and accountability, \u003cem\u003e32\u003c/em\u003e(2), 237\u0026ndash;270.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKanu, A. (2015). The role of education in addressing socio-economic disparities in West Africa. Journal of African Development, \u003cem\u003e18\u003c/em\u003e(4), 67\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhairani, A. Z., \u0026amp; Shamsuddin, H. (2016). Assessing item difficulty and discrimination indices of teacher-developed multiple-choice tests. In \u003cem\u003eAssessment for Learning Within and Beyond the Classroom: Taylor\u0026rsquo;s 8th Teaching and Learning Conference 2015 Proceedings\u003c/em\u003e (pp. 417\u0026ndash;426). Springer Singapore.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLinn, R. L. (2000). \u003cem\u003eAssessments and accountability\u003c/em\u003e. Yearbook of the National Society for the Study of Education, 99(1), 12\u0026ndash;41.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMislevy, R. J. (1991). Test theory and the measurement of abilities. Educational Measurement, \u003cem\u003e10\u003c/em\u003e(2), 45\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMusa, A., Shaheen, S., Elmardi, A., \u0026amp; Ahmed, A. (2018). Item difficulty \u0026amp; item discrimination as quality indicators of physiology MCQ examinations at the Faculty of Medicine Khartoum University. Khartoum Medical Journal, \u003cem\u003e11\u003c/em\u003e(2).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNtumi, S., Agbenyo, S., \u0026amp; Bulala, T. (2023). Estimating the Psychometric Properties (\u0026ldquo; Item Difficulty, Discrimination and Reliability Indices\u0026rdquo;) of Test Items Using Kuder-Richardson Approach (KR-20). Shanlax International Journal of Education, \u003cem\u003e11\u003c/em\u003e(3), 18\u0026ndash;28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOhiri, S. C. (2023). Psychometric Analysis at Item Level of WAEC May/June Mathematics Multiple Choice Questions Using the Item Response Theory. Unizik Journal of Educational Research and Policy Studies, \u003cem\u003e16\u003c/em\u003e(5), 66\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalazzo, S. J., \u0026amp; Levey, J. (2024). Shaping the Future of Nursing Education: Next Generation NCLEX Question Writing and the Power of Psychometrics. Nursing Education Perspectives, \u003cem\u003e45\u003c/em\u003e(2), 69\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePellegrino, J. W., Chudowsky, N., \u0026amp; Glaser, R. (2001). \u003cem\u003eKnowing what students know: The science and design of educational assessment\u003c/em\u003e. National Academy Press.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRossiter, J., Abreh, M. K., Ali, A., \u0026amp; Sandefur, J. (2023). The high stakes of bad exams. Journal of Human Resources.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaregar, A., Putra, F. G., Anugrah, A., \u0026amp; Fitri, M. R. (2024). Assessment Instrument Model for Natural Science Material Based on Research Findings and Quranic Verses: An Effort to Measure Students' Higher-Order Thinking Skills. IJIS Edu: Indonesian Journal of Integrated Science Education, \u003cem\u003e6\u003c/em\u003e(2), 199\u0026ndash;209.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong, X. (2014). Test fairness in a large-scale high-stakes language test. Unpublished doctoral dissertation). Queen\u0026rsquo;s University, Canada.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZondo, N., Zewotir, T., \u0026amp; North, D. E. (2021). The level of difficulty and discrimination power of the items of the National Senior Certificate Mathematics Examination. South African Journal\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Psychometric analysis, item difficulty, item discrimination, socio-economic disparities, regional disparities, linguistic bias, high-stakes exams, West Africa","lastPublishedDoi":"10.21203/rs.3.rs-6313027/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6313027/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis study conducted a psychometric analysis of high-stakes examination items in West Africa, focusing on the relationship between item difficulty and item discrimination across socio-economic, regional, and linguistic backgrounds. Using advanced statistical techniques, including Hierarchical Linear Modeling (HLM) and Structural Equation Modeling (SEM), the study examined how contextual factors influenced test fairness and predictive validity. The analysis utilized a dataset of 12,500 exam items and responses from 250,000 students across five West African countries. Descriptive statistics revealed significant disparities: students from lower socio-economic backgrounds faced higher item difficulty (M = 0.78, SD = 0.12) compared to upper socio-economic groups (M = 0.58, SD = 0.08), while item discrimination was highest among wealthier students (M = 0.60, SD = 0.08). Multilevel regression analysis indicated that socio-economic status significantly predicted item difficulty (β = -0.10, p \u0026lt; 0.01) and item discrimination (β = 0.18, p \u0026lt; 0.01), demonstrating systematic disadvantages for lower-income students. Regional disparities also played a critical role, with rural students encountering more difficult items (M = 0.70, SD = 0.13) than their urban counterparts (M = 0.63, SD = 0.11). Regression analysis confirmed a significant effect of location on item difficulty (β = 0.08, p \u0026lt; 0.01) and discrimination (β = 0.11, p \u0026lt; 0.01), suggesting that rural students faced structural disadvantages in assessment. Linguistic background was another major determinant of assessment fairness. Indigenous language speakers encountered the highest item difficulty (M = 0.75, SD = 0.13) and the lowest discrimination values (M = 0.45, SD = 0.12), compared to English-speaking students (M = 0.60, SD = 0.10) and French-speaking students (M = 0.65, SD = 0.12). Multi-group SEM analysis demonstrated that indigenous language speakers experienced disproportionately higher item difficulty (β = 0.12, p \u0026lt; 0.01) and lower discrimination (β = 0.10, p \u0026lt; 0.01), confirming linguistic bias in standardized testing. The study’s findings highlight structural inequalities embedded within West African high-stakes examinations, raising concerns about fairness and access to educational opportunities. The results underscore the need for equity-driven assessment practices, such as adaptive testing models, differential item functioning (DIF) analysis, and linguistically inclusive test designs. These findings align with global recommendations on assessment fairness and call for urgent reforms in test development and validation.\u003c/p\u003e","manuscriptTitle":"Psychometric Analysis of High-Stakes Examination Items in West Africa: Evaluating Item Difficulty and Discrimination Across Diverse Socioeconomic and Regional Contexts","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-12 15:50:32","doi":"10.21203/rs.3.rs-6313027/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"df558aa1-0baa-4621-999d-f290722dd07f","owner":[],"postedDate":"May 12th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-07-24T09:54:07+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-12 15:50:32","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6313027","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6313027","identity":"rs-6313027","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Outcome instruments

MUSA

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0