Predictive AI for Academic Performance: Integrating Natural Language Processing, Gamification and Self-regulated Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predictive AI for Academic Performance: Integrating Natural Language Processing, Gamification and Self-regulated Learning Laia Subirats, Beatriz Narbona, María Elena Cuenca, Sacha Gómez-Moñivas This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7801827/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Artificial Intelligence and Natural Language Processing are widely used to predict academic performance across education levels. This study examines how students’ personality and maturity relate to academic progress by comparing their actual degree program year with predictions generated through advanced Artificial Intelligence and Natural Language Processing techniques. Our predictive model achieved a mean absolute error of 0.794. Maturity and personality traits were derived from two sources: (a) Type-Token Ratio and Flesch-Kincaid Grade used to assess developmental writing features; and (b) Hexad Model player-types to classify motivational traits. Both data types were collected through post-task surveys following a gamified activity. The model was applied to students across all four years of a Tourism degree program. Results show that NLP-based linguistic features and gamification profiles were the strongest predictors. The data confirmed expected gains in lexical diversity and readability across academic years. Additionally, shifts in player types – from first to fourth year – suggest evolving motivational orientations and personality development. These findings offer valuable insights for identifying students whose developmental trajectory may not align with their academic standing. When predicted and actual program year differ, educators can use this signal to provide targeted, timely support. AI-based writing analysis fosters maturity monitoring, metacognitive awareness and self-regulated learning. Meanwhile, personality profiling enables differentiated instruction based on motivational drivers. This methodology offers a scalable tool for inclusive, personalized education – particularly in multilingual settings where accurate translation preserves students’ linguistic voice. Natural Language Processing gamification self-regulated learning machine learning academic performance higher education Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 1. Introduction Artificial intelligence (AI) and Natural Language Processing (NLP) are increasingly used for diagnosis, prognosis, prediction, and personalization in higher education, as emphasized by initiatives such as the UNESCO Beijing Consensus (UNESCO, 2019 ) and the OECD's Digital Education Outlook 2023 (OECD, 2023 ). The integration of AI into education has evolved from computer technologies (Chen et al., 2020 ; Devedžic, 2004 ) to intelligent systems (Chassignol et al., 2018 ; Kahraman et al., 2010 ; Peredo et al., 2011 ; Rus et al., 2013 ) and tools like robots and chatbots (Timms, 2016 ), now propelled by the advances in Generative AI (Pokrivcakova, 2019 ; Salas-Pilco & Yang, 2022 ). Many scholars highlight the urgency of implementing Generative AI to enhance accessibility, efficacy, and inclusivity in education (Lim et al., 2023 ). The European Union’s AI Act mandates ethical standards for AI applications in educational contexts (European Parliament, n.d.). The advantages of Generative AI in personalized learning (Barrett & Pack, 2023 ; Sajja et al., 2023 ; Samuel Girard et al., 2024), prediction, and analytics – such as forecasting student scores (Ashenafi et al., 2015 ), success rates (Ludwig et al., 2024 ), or assessing personality and maturity (Deo et al., 2020 ; Rovira et al., 2017 ) – are well-established in academic research. These benefits underpin our analysis and form the foundation of our predictive model. 2. Theoretical background 2.1. Constructivism, AI-enhanced active learning and self-regulated learning Constructivist theory, particularly as articulated by Vygotsky and Bruner, posits that learning is most effective when instruction is scaffolded to students’ current developmental stage – within their “zone of proximal development” (Vygotsky, 1978 ). Accurately predicting a student’s academic year enables educators to estimate their cognitive and linguistic maturity and design instruction that bridges prior knowledge with new content (Bruner, 1960 ; Vygotsky, 1978 ). Empirical studies confirm that academic progression is associated with qualitative shifts in writing and discourse sophistication (Crossley & Kim, 2022 ). Predicting academic year is therefore not a bureaucratic exercise, but a means to enact responsive, developmentally appropriate pedagogy (Bruner, 1960 ; Vygotsky, 1978 ). AI-based tools can facilitate this by providing interactive, learner-centered environments that promote exploration and immediate engagement with content. AI can offer tailored, interactive learning experiences that allow students to actively construct knowledge, improving learning efficiency (Holmes et al., 2019 ). Instead of passively receiving information, students using AI-driven systems (such as intelligent tutors or simulations) learn by doing – aligning with Piaget’s constructivist view that experiential interaction is crucial (Shuliang, 2025 ). Empirical evidence supports this synergy: AI-driven adaptive learning platforms continuously adjust activities to a student’s level, which “markedly enhance[s] academic performance” by keeping tasks in an optimal challenge zone ((Baker & Yacef, 2009 ) as cited in (Al Nabhani et al., 2025 )). 2.2. Self-regulated learning and individual profiles Self-regulated learning (SRL) theory emphasizes students’ ability to set goals, monitor progress, and regulate both cognition and motivation throughout the learning process (Zimmerman, 2002 ). Self-regulated learners take an active role in their education by planning, monitoring, and adjusting their cognitive and metacognitive processes to achieve their learning goals (Nicol & Macfarlane-Dick, 2006 ). Research highlights the close alignment between constructivist classrooms and SRL, with teachers scaffolding students’ ability to monitor, control, and evaluate their own learning (Nicol & Macfarlane‐Dick, 2006)). Key SRL models highlight motivational factors – such as self-efficacy, intrinsic interest, and goal orientation – as crucial to academic achievement (Pintrich & De Groot, 1990 ). Research consistently finds that differences in self-regulation and motivation explain significant variance in student outcomes (Panadero, 2017 ; Zimmerman, 2002 ). Understanding students’ motivational profiles can help educators detect learners who may lack strategic approaches or intrinsic engagement. Research suggests that low self-efficacy can hinder the use of effective learning strategies, which may be addressed through scaffolded tasks and explicit feedback (Pintrich & De Groot, 1990 ; Zimmerman, 2002 ). These insights support the tailoring of both instruction and feedback, addressing the whole learner rather than just knowledge deficits. This is particularly relevant in digital environments, where self-regulation ability predicts academic success (Broadbent & Poon, 2015 ). Data-driven models and AI can support SRL by providing timely feedback and insights that help learners manage their learning process (Afzaal et al., 2021 ). A recent systematic review found that technologies like learning analytics and AI “support SRL by providing personalized feedback and facilitating autonomous learning” (Faza & Lestari, 2025 ). 2.3. Predicting academic performance with AI for pedagogical action Formative assessment – “assessment for learning” – involves gathering evidence about students’ current understanding to inform immediate feedback and instructional adaptation (Black & Wiliam, 1998 ). Predictive diagnostics, whether through teacher judgment or algorithmic models, serve as a contemporary extension of formative assessment: they provide timely insights into students’ academic standing, linguistic development, and motivational needs. The pedagogical value of predictive analytics in education lies not in the prediction itself, but in its capacity to prompt formative feedback and inform adaptive teaching practices (Black & Wiliam, 1998 ; Nicol & Macfarlane-Dick, 2006 ). Classic mastery learning models (Bloom, 1968 ) and contemporary learning analytics both demonstrate that diagnostic insights, when coupled with targeted interventions, lead to measurable improvements in achievement and reduced failure rates (Ifenthaler & Yau, 2020 ). Thus, predictive diagnosis is pedagogically valuable not for its own sake, but for its role in activating meaningful feedback loops that promote student growth. Predicting final grades is one of the most significant applications of AI in education, as it closely ties to academic performance. AI has long supported performance monitoring and personalized content to preempt learning obstacles (Bloom, 1971 ; Bloom at al., 1981; Glaser, 1994 ; Guskey, 1985 ). Early studies on machine learning for grade prediction (Ashenafi et al., 2015 ) have focused on Massive Open Online Courses, while additional research examined early-detection systems' impact on students and universities (Raffaghelli et al., 2022 ; Subirats et al., 2023 ). AI applications also extend to supporting students with learning challenges, such as dyslexia, through methods like eye-tracking (Rello et al., 2020 ), recommending courses (Samuel Girard et al., 2024), and adaptive learning or tailored content (Kolluru et al., 2018 ; S. Liu et al., 2024 ). These applications can help students, teachers, or both (Kumar et al., 2023 ; Pankiewicz & Baker, 2023 ). Moreover, AI tools help assess problem-solving behaviors and predict success in skill simulations, using NLP techniques to analyze behavior sequences (Ludwig et al., 2024 ). As an interdisciplinary field combining AI and linguistics, NLP has proven to be among the most effective methods for personalized, student-centered learning models. AI-driven predictions are valuable only when translated into concrete educational actions. Predictive analytics can flag at-risk students based on behavioral or performance data, enabling timely interventions such as academic mentoring or adaptive scaffolding (Ifenthaler & Yau, 2020 ; Slade & Prinsloo, 2013 ). These early interventions have been shown to improve retention and learning outcomes by addressing difficulties before they escalate (Larrabee Sønderlund et al., 2019 ). Moreover, predictions of motivational or cognitive profiles allow instructors to tailor learning tasks, pacing, and feedback strategies – key components of differentiated instruction (Tomlinson, 2017 ). Instructional design thus becomes more responsive to students’ self-regulatory capacities, engagement levels, and academic trajectories. At the curriculum level, aggregate learning data can reveal patterns of difficulty across cohorts, guiding revisions that better align with learners’ developmental needs (Boroowa & Herodotou, 2022 ). In line with Nicol and Macfarlane-Dick’s ( 2006 ) model of formative assessment as a feedback loop that fosters self-regulated learning, predictive AI models can generate timely insights that inform adaptive interventions, thereby enhancing personalization and equity in learning environments. 2.4. Gamification, personality traits and academic performance Gamification’s role in education enhances personalization and student-centered learning, despite challenges to its long-term efficacy. Studies on gamification often examine user type classification (Bartle, 1996 ; Drachen et al., 2009 ; Hamari & Tuunanen, 2014 ), including models like the Brain Hex (Tondello et al., 2019 ), which align personality traits with gamified tasks. The Hexad Model has proven effective for personalizing educational activities in the teaching-learning process (Lopez & Tucker, 2019 ). Building on this, our study integrates player types as personality traits into a predictive model, based on the established connection between personality and academic performance. By leveraging survey data from gamified tasks, we integrate these personality indicators to enrich our model. Personality traits’ link to academic performance is well-established, with the Big Five model or the Eysenckian personality factors proving useful for predicting academic performance across educational contexts (Brandt et al., 2020 ; Mammadov, 2022 ; Mornar et al., 2022 ; Naqshbandi et al., 2017 ; O’Connor & Paunonen, 2007 ; Poropat, 2014 ). The role of personality in academic performance is further explored in studies that connect personality traits to course grades (Lounsbury et al., 2003 ), or examine the intersection of personality traits, motivation, and academic success, suggesting that these factors are interrelated and can contribute to predicting performance (De Feyter et al., 2012 ; Phillips et al., 2003 ). Furthermore, player types, which are associated with different levels of motivation, engagement, and academic achievement, can influence learner’s behavior in gamified environments (Lavoué et al., 2021 ), aligning with our study's exploration of player types and their role in student performance. Current trends in education emphasize (a) adopting emerging technologies like gamification, NLP, and generative AI, and (b) predicting academic performance to implement tailored interventions. Following these lines, our article proposes a methodology to predict students’ university program year or academic course (from 1st to 4th) using Hexad player types and NLP text analysis. If the predicted year is lower than the actual, it may indicate underperformance; if higher, it may suggest exceeding expectations. This assumption is based on the idea that the model draws from skills that students develop progressively through their studies. We further assess whether the predicted outcomes correlate with students’ maturity development, and the relevance of the different parameters included in such predictions. 2.5. Linguistic development and academic maturity Psycholinguistic research establishes a relationship between students’ development – whether socio-cognitive or academic – and the improvement of writing skills and linguistic maturity, with language development recognized as a lifelong process (De Bot & Schrauf, 2010 ; Rosselli et al., 2014 ). This process unfolds through different developmental stages based on age, which show changes in linguistic literacy (Ravid & Tolchinsky, 2002 o Lugo, 1996) and maturity (Ravid & Tolchinsky, 2002 o Lugo, 1996), in line with socio-cognitive evolution (Anderson, 1941 ; Bremholm et al., 2022 ; Hall-Mills & Apel, 2015 ; Wagner et al., 2011 ). Although much discourse focuses on childhood and pre-adolescence, some researchers (Berman, 2007 , 2008 ; Grimshaw et al., 1998 ; Nippold, 2000 ) emphasize that the upper grades of high school and college represent a critical period in language-specific skill development, particularly in grammatical and semantic knowledge (Berman & Nir-sagiv, 2007 ; Hansson et al., 2016 ; Karmiloff-Smith, A., 1992 ; Nippold, 2002 ; Ravid, 2005 ; Ravid et al., 2002 ; Rimmer, 2008 ; Tolchinsky & Rosado, 2005 ) and genre-dependent text organization, and in global coherence (Berman, 2008 ; Katzenberger, 2005 ). This phase is crucial, as it aligns with the consolidation of general socio-cognitive development and the enhancement of higher-order cognitive capacities (Proverbio & Zani, 2005 ), such as abstraction, perspective-taking, divergent and critical thinking, and executive control (Berman, 2017 ; Kluwe, & Logan, 2000 ; Reilly et al., 2005 ; Sato, 2022 ), alongside growing world knowledge and exposure to information during maturation and learning. The use of linguistic features such as nominal density, lexical diversity, word and sentence length, and syntactic complexity for assessing linguistic development or maturity is well-established in academic research (Crossley, 2020 ; Crossley & Kim, 2022 ; Durrant & Brenchley, 2019 ; MacArthur et al., 2019 ; Malvern et al., 2004 ; Ravid, 2006 ; Richards & Malvern, 1997 ; Shakir & Obeidat, 1991 ; Sun & Xiong, 2019 ). Our study builds on previous work that analyzes the academic writing of adolescents and young adults in expository texts (Berman & Nir, 2009 ; Berman & Nir-sagiv, 2007 ; Ragnarsdóttir et al., 2002 ; Ravid, 2005 ), “which focus on issues and ideas, and express the unfolding of claims and argumentation in causal and other logical contexts” ((Ravid, 2006 : 794). These analyses have provided evidence of the link between enhanced writing skills and the cognitive and socio-cultural development of students, demonstrating that “lexicon and syntax interact with age-related changes in overall attitudes expressed in discussing a socially relevant theme” (Berman, 2017 : 5), and the interconnectedness between language development and academic performance (Nippold, 2016 ). Building on these studies, we analyzed the writing features of university students’ responses to a questionnaire, focusing on justified-opinion answers. We employed metrics of lexical diversity and readability – excluding textual global coherence analysis due to the brevity of the expository texts – to assess students’ linguistic maturity and development within higher education environments. Since our data is in Spanish, and most NLP algorithms operate in English, we analyze the impact of translation methods to ensure result reliability. Analyzing open-ended responses introduces complexity but mitigates response bias, as students remain unaware of the specific data points under analysis. 2.6. Research questions To predict students’ academic performance, this study introduces an innovative methodological framework that combines Hexad player types, NLP techniques, and translation methods to adapt Spanish-language data for analysis with English-based NLP algorithms. Furthermore, it incorporates the examination of students justified-opinion responses to open-ended questionnaire items. This approach aims to provide insights into learners’ development by addressing the following research questions: RQ1: To what extent can the degree program year be predicted from students’ multiple-choice and open-ended questionnaire responses, along with their gamification type? RQ2: Which variables – including measures of linguistic features, Hexad player types, and responses to opinion survey items and gaming behaviors questionnaire – most strongly influence the prediction of students’ program year? RQ3: What insights into students’ maturity or personality development can be gained through text analysis parameters and player types? 3. Materials and methods This section is divided into two parts. The first one (3.1) describes the experimental context, including datasets and its exploratory analysis. The second one (3.2) describes how gamification user types, NLP, machine learning techniques and statistical analysis are used to predict the program year. 3.1 Experimental context 3.1.1 Sample Data were collected from the Degree in Tourism at XX University. A total of 189 students from the program’s four years participated in the study. The demographic characteristics of the participants are provided in Table 1: Table 1. Demographic characteristics of the participants (N=189) Age 18-19 34.91% 20-21 47.93% 22-23 13.61% 24 3.55% Program year 1st 25.39% 2nd 28.57% 3rd 31.75% 4th 14.29% Nationality Spanish 76.92% EU not Spanish 5.29% Chinese 9.47% LatAm 7.69% Gender Male 36.84% Female 61.99% Other / Prefer not to answer 1.17% 3.1.2 Procedure Data acquisition for all AI inputs was done in a single session for each class to gather the data in the most controlled environment. After scheduling each program-year group class session, students were informed one week in advance that there would be a special activity to be held in the computers lab related to heritage, gamification and topics from their degree subjects. They were also informed that the data would be used for a pilot study and that all personal data would be excluded to prevent identification. This activity is part of the Tourism Degree soft-skills acquisition tasks which contribute to fulfilling curricular competencies. These competencies are embedded within specific course syllabi and are taught through classroom activities designed to foster critical and creative thinking. By generating contexts where students express their opinions and make suggestions, these activities provide a framework for collecting open responses data. Since students are used to this format when giving their critical opinion, the controlled classroom environment improves data reliability and avoids deviation. This activity gets students answering questions linked to skills they are actively developing throughout their degree program. As these skills are an essential component of their curriculum, the activity aligns with their training and expected proficiency. This provides a scenario for the purpose of this article, as it allows for a structured comparison of differences across program years. The administration of the activity in the classroom was as follows: Attendance was controlled upon entering the lab so that the responses collection could be tracked. Students were firstly instructed on what an excavation site is, the site context and staff, and the app keyboard controls. They were then allowed to open an application with a virtual reconstruction of an excavation in Egypt and play for approximately 40 minutes. In this virtual reconstruction they can visit different areas and interact with objects that test their understanding. After that, they were told to answer the questionnaire giving feedback about the application, and later to fill in the survey about their gaming habits. Fig. 1 shows the stages of the activity, and the data obtained in each stage. As shown in the figure, all the data obtained in this activity were collected in three stages: (a) Attendance control: Demographic variables & program year. (b) A questionnaire collecting feedback about the activity: nine 1-5 points questions rating the app contents and technical features plus four open questions where they can express their opinion freely. (c) A survey about gaming habits, attitudes or personal preferences: 27 questions to be rated with 1-5 points and two open questions. The variables collected are summarized in Fig. 2 diagram (see detailed information in Appendix A). 3.1.3. Ethical issues All data collected for this study were part of an activity that was integrated into the course curriculum and mandatory for all students. Upon consultation, the Ethics Committee at Universidad XX granted approval for the study in this context. Being in a classroom setting, a strict protocol has been established for all faculty members using data collected from teaching activities to ensure that students’ personal data is not published in a way that could lead to identification. Informed consent was obtained from all participants through the following statement on the forms in which they provided the data used in this study: “During this course we are using some game-like elements. The game elements are part of a pilot activity, aiming to explore what kinds of gamification elements work in different course contexts and for different users. For this purpose, we kindly ask you to answer this brief survey. Personal data will only be used to link participants’ survey responses to respective course data. The data will be process among GDPR, and instructions of XX University (…). 1. Consent to participate. I understand the nature of the study, and ___ agree to participate in the study. ___ agree to the use of the data collected for course purposes only”. Although the activity was mandatory as part of the course requirements, students had the option to opt out of having their data used in the study. In such cases, the teacher would still use the data for course-related purposes, but the data would be excluded from the study. Not a single student opted out of having their data included in the study. The consent statement was prominently displayed at the top of the forms, ensuring that students did not submit or complete any open-ended responses, writing activities, or other types of data shown in Fig. 2 without first being informed. The generative AI tool ChatGPT 4 and the Bing Copilot of XX University have been used to perform the Generative AI section. 3.2 Machine learning prediction 3.2.1 Text data treatment The process for obtaining the machine learning’s input features is illustrated in Fig. 3. As shown, readability and lexical diversity metrics require pre-processing, as it is necessary to translate the open-ended responses from Spanish into English to utilize open-source Python libraries, which are predominantly English-specific. To minimize translation-related artifacts, we applied four translation approaches, each rigorously analyzed to ensure accuracy: Open-Source Python Libraries : Using Python open libraries for both the translation and obtaining Type-Token Ratio (TTR) and Flesch-Kincaid Grade (FKG). Python Libraries + Linguist review : Using Python open libraries for translation, reviewed later by a linguist, and Python open libraries for obtaining TTR and FKG. ChatGPT 4 + Python Libraries : Uses Generative AI (ChatGPT 4) for the translation and Python open libraries for obtaining TTR and FKG. ChatGPT 4 + Bing Copilot : Uses Generative AI both for the translation (ChatGPT 4) and obtaining TTR and FKG (Bing Copilot). 3.2.2 Natural Language Processing Among the several open-source Python libraries for Spanish-English translation, we used the free software library Translators 5.9.0 https://pypi.org/project/translators. In addition, two other translation approaches were implemented: translated text with a linguist revision, and generative AI translation using ChatGPT 4. For this last translation, careful prompting was needed to ensure that students’ writing style was not changed due to an inaccurate translation diluting their personal features, which could affect maturity or program-year predictions. As the translated text is used as input to assess lexical diversity and readability, it is relevant to preserve the original style regarding semantics, syntax and any feature that may portray students’ voice (note that readability is one of the criteria used in translation to assess quality (Nababan & Nuraeni, 2012; Wijaksono et al., 2022). Rendering stylistic characteristics and author’s personality was a problem in machine translation (Mirkin et al., 2015; Rabinovich et al., 2016), and the implementation of AI with Neural Machine Translation systems has not solved the issue. ChatGPT may produce good quality translations if suitable prompts optimize the process avoiding mistakes – poor understanding of context, biases, and flaws in detecting cultural nuances (Bender et al., 2021; Brown et al., 2020; Castilho et al., 2017; Koehn & Knowles, 2017). But Large Language Models – trained on vast datasets to create a single model – dilute individual stylistic elements and can hide the author's demographic, psychometric or personal features in unfaithful translations (Britz et al., 2017; Busch, 2024; Ji et al., 2023; Vanmassenhove, 2024). Therefore, well-crafted prompting emphasizing literal or faithful translation techniques (Newmark, P., 1988) that prevent skipping author’s stylistic characteristics (Hu et al., 2017; Rabinovich et al., 2016; Sennrich et al., 2016; Shen et al., 2017) and grant style transfer (Prabhumoye et al., 2018) was needed for our study. Thus, besides providing a description of the text’s context, specific commands were repeated urging ChatGPT to preserve the student’s degree of formality and range of vocabulary, replicate mistakes and colloquialisms, and be aware of the importance of accuracy and fidelity (see some details in Appendix B). Regarding lexical diversity and readability, the 2021-released Python open library https://github.com/WSE-research/LinguaF/tree/main (2024 update) was used, particularly its TTR and FKG measures. TTR is a measure of lexical diversity in a text. It is commonly used in linguistics and computational linguistics to quantify the variety of different words (types) relative to the total number of words (tokens) in a given text. The higher the TTR, the greater the lexical diversity in the text. A lower TTR suggests more repetition or fewer unique words relative to the total number of words. The values for TTR range typically from 0 to 1, but it can also be expressed as a percentage. The interpretation of TTR may vary depending on the context and the type of text being analyzed. Different genres, writing styles, or languages may naturally exhibit different levels of lexical diversity. TTR is a useful tool for comparing the diversity of vocabulary across different texts or for tracking changes in vocabulary richness over time. The FKG is a readability test designed to estimate the readability level of English texts. It is commonly used to assess the complexity of written material and is often applied to evaluate the difficulty of reading comprehension. The FKG is based on two factors: the average number of words per sentence and the average number of syllables per word. The formula for calculating the FKG is: FKG=0.39 (Total words / Total sentences) + 11.8 (Total Syllables / Total words) - 15.59 One of the interpretations of the resulting FKG is a correspondence with a U.S. grade level, indicating the level of education generally required to understand the text. For example, an FKG of 8.0 would suggest that the text is readable by an eighth grader. The range of FKG values typically corresponds to the 12 U.S. grade levels. Lower FKG values indicate simpler text, while higher values suggest more complex and advanced writing. The FKG is just one of many readability metrics, and its interpretation can be influenced by factors such as sentence structure and word choice. Negative values are not common and would typically occur in situations where the text is extremely short, and the formula results in a grade level under 1. 3.2.3 Gamification Hexad model of player types To integrate personality traits into our computational analysis, we applied the Hexad model (Marczewski, 2015), a validated framework in educational gamification. This model categorizes users based on distinct motivations, aligning well with our study’s focus on personalizing prediction. Data was collected via a survey completed after students engaged in a gamified virtual excavation task, which enabled us to capture diverse motivational factors in an interactive learning environment. The Hexad model, tested in educational contexts (Subirats, Nousiainen, et al., 2023), provides six non-excluding user types (Tondello et al., 2016) that reflect motivations key to predicting academic progression: Philanthropists—motivated by purpose; Socializers—motivated by relatability; Free spirits—motivated by autonomy; Achievers—motivated by competence; Players—motivated by rewards; Disruptors—motivated by change. The relationship between Hexad player types and students’ development is central to this study, as predicting the program year depends on features that may evolve as students’ progress through their studies. In the Hexad player profile, motivation levels may vary, with students often focused on broader goals in early years and increasingly career-oriented in later years as they prepare for post-graduation opportunities. 3.2.4 Machine learning. To ensure that the translation into English is not an artifact in the predictions we followed these four steps for applying AI algorithms in the prediction of the program year: Data is obtained in class (the target is the program year, with values between 1 and 4). We apply one of the translations got from each approximation. From each translation, TTR and FKG are obtained. With each TTR, FKG and other variables, we apply Random Forest for predicting the program year. Pandas library was used for processing data and implementing the Hexad framework of six user types. Three different approaches were used to impute null values: Median, k-Nearest Neighbors algorithm and Iterative imputer. After conducting comparative program-year prediction trials among the three, the iterative imputer was chosen as it showed better results for predicting the program year. The supervised learning algorithm chosen was the Random Forest regressor because it has good and robust performance when tuning the parameters accordingly, besides providing information about the most relevant parameters for the prediction. The Python open library used was https://scikit-learn.org and the Random Forest regressor was trained with 10-fold cross-validation and 80% of train data and 20% of test data. The Mean Absolute Error (MAE) and Mean Squared Error (MSE) (Karunasingha, 2022) were used as evaluating metrics to analyze the performance of the algorithm. After applying machine learning to do statistical tests on the prediction results, normality tests were conducted showing that they do not follow a normal pattern; so, the Wilcoxon signed-rank test was used (Moore et al., 2009). The Wilcoxon signed-rank test is a statistical test used to compare the means (averages) to determine if they are significantly different from each other. It helps to figure out whether the differences you observe in certain data are real or if they could have happened by chance. It was applied to compare the statistical significance of the best approximation to all the other approximation errors. Therefore, three statistical Wilcoxon tests were computed. 4. Results This section is divided into two subsections: (4.1) an exploratory data analysis performed before applying machine learning, and (4.2) the results of performing supervised learning to predict the program year of students from analyzing all the features described in Fig. 2. 4.1 Exploratory data analysis A correlation heatmap of features is shown in Fig. 4, where blank values in the correlation heatmap of “Other gender / NA” (not answered) vs. some variables mean that there is not enough data to compute these correlations due to null values (detailed correlation with numerical values provided in Appendix D). This figure shows that there are some high correlations between open-questions, and the inverse correlation of “Suggestions” readability and lexical diversity is noteworthy. High and positive correlations between students’ opinions about “Evaluation of activities” and “Technical elements of the application” are also observed. Besides, interesting contrasting correlations are seen between genders and the “Frequency of playing video games”: -0.3 between “Females” and “Frequency of playing video games”, and 0.3 between “Males” and “Frequency of playing video games”. Additionally, the correlation between gamification user types is positive too, not contradicting previous studies (Tondello et al., 2016). Detailed figures from the exploratory data analysis are provided in Appendix D, including Table A.1, which presents the mean values and the evolution of Hexad player profiles, and Fig. A.1, which offers a comprehensive view of the correlation matrix. 4.2 Prediction with supervised learning – regression Table 2 shows the results of the program year’s prediction after applying supervised learning with Random Forest regression type. To interpret its contents, the differences between MSE/MAE and Train error / Test error must be explained. The mean MSE/MAE of the training set are the errors between the predicted values and the actual values for the data the model was trained on. It provides an indication of how well the model fits the training data. A low MSE/MAE on the training set generally indicates that the model is fitting the training data well. However, if it is too low compared to the test MSE/MAE, it might indicate overfitting. The mean MSE/MAE of the test set is the average absolute error between the predicted values and the actual values for the data that was not provided to the model during training. It provides an indication of the model's generalization performance with unknown data. A low MSE/MAE on the test set suggests that the model generalizes well to new data, while a high MSE/MAE indicates poor generalization, which might be due to overfitting or underfitting. Regarding which metrics are best suited, MAE and MSE have different properties. MAE is linear, meaning all errors are weighted equally. It is more robust to outliers compared to MSE because it does not square the errors. MAE provides a clear interpretation: it gives the average error in the same units as the original data. MSE squares the errors, which means it penalizes larger errors more than smaller ones. It is sensitive to outliers due to the squaring of differences. MSE's unit is the square of the original data units, which can be less intuitive to interpret directly. Due to what has just been mentioned, MAE can be more suitable for this problem because of its better interpretability and robustness to outliers. Considering the information above, in Table 2 we can see that the minimum values of the test’s errors are obtained when the translation is reviewed by a linguist, closely followed by the errors obtained by ChatGPT 4 translation. Table 2. Errors of the program-year prediction from students’ open data using Random Forest. Approximation 1: Python open libraries Approximation 2: Python open libraries + linguist-reviewed translation Approximation 3: ChatGPT4 translation + Python open libraries computing TTR & FKG Approximation 4: ChatGPT4 translation + Bing Copilot computing TTR & FKG Mean MSE 0.842 0.840 0.820 0.754 SD MSE 0.162 0.169 0.149 0.152 Test MSE 1.004 0.922 1.016 0.984 Test MAE 0.856 0.794 0.856 0.809 Lowest test-set error values in bold. By applying the Shapiro-Wilk test, we have concluded that the data can significantly deviate from a normal distribution. Then we apply the Wilcoxon test from the list of errors obtained from the 4 approximations to compare the best one (second) to the others. We have obtained that, in all cases, we cannot find statistically significant differences. The difference that is closer to being significant is between approximations 1 and 2, with statistic=236.5 and p-value=0.08. Thanks to the fact that the Random Forest is explainable and can provide information about the relevance of each feature, the most important attributes of the prediction could be identified. Fig. 5 shows a lollipop chart of the features’ importance. 5. Discussion and implications This study is anchored in constructivist and self-regulated learning theories. Predicting students’ program year through AI-supported models enables more tailored educational scaffolding, aligning instruction with students’ developmental stages (Bruner, 1960; Crossley & Kim, 2022; Vygotsky, 1978). Methodologically, combining open- and closed-question survey data with advanced machine learning represents a step forward for educational analytics. Integrating gamification user types responds to calls for personalized learning environments where motivation and individual profiles influence academic success (Broadbent & Poon, 2015; Panadero, 2017; Pintrich & De Groot, 1990). 5.1. Research questions discussion Our application of lexical diversity and readability analysis via NLP, together with gamification user types, aimed to predict students’ academic maturity (i.e., program year) based on questionnaire responses. This approach addressed three research questions, discussed below based on our findings. RQ1: To what extent can the degree program year be predicted from students’ multiple-choice and open-ended questionnaire responses, along with gamification types? As shown in Table 2, predictions of program year were reasonably accurate, with errors 1 in all cases – except when using unedited free translation tools. The lowest prediction error, under 0.8, occurred when human post-editing was applied. Although differences among translation methods were not statistically significant, human revision consistently improved accuracy. This highlights the continued importance of human oversight in AI-assisted translations, especially when high precision is needed. However, when handling large datasets, human intervention may be impractical due to time and resource demands. In such cases, automatic translation tools remain acceptable for large-scale analysis (Castilho et al., 2017). Where possible, responses in student’s native languages should be prioritized, especially when they are not native speakers of English or Russian – the languages supported by the LinguaF library. Using a second language may distort readability and lexical diversity metrics, as non-native writing may not fully reflect students’ academic development, personality traits, or natural writing style. Therefore, collecting data in students’ first language preserves the authenticity of linguistic features used in our analysis. RQ2: Which variables – including measures of linguistic features, Hexad player types, and responses to opinion survey items and gaming behaviors questionnaire – most strongly influence the prediction of students’ program year? As Fig. 5 illustrates, lexical diversity and readability metrics were the most predictive features, outperforming demographic and attitudinal variables. This aligns with prior research showing that writing sophistication increases throughout adolescence and early adulthood (Berman & Nir-sagiv, 2007; Crossley, 2020; Nippold, 2016; Ravid, 2005). Only one Likert-scale item, evaluating “Technical elements of the application 4” ranked among the top ten predictors – specifically, in tenth place. Among the Hexad player types, “Free Spirit,” “Socializer,” and “Philanthropist,” showed higher predictive relevance, suggesting that motivational orientations evolve across program years (Santos et al., 2023; Tondello et al., 2016). While these features are key to the model’s performance, Table A.1 and the correlation matrix revealed no linear correlations between them and program year. This supports the use of non-linear AI models to capture more complex relationships (Deo et al., 2020; Rovira et al., 2017). The correlation matrix also showed positive associations among parameters from the same source. Likert-scale responses tended to correlate positively with one another, though without revealing distinct trends across program years. This suggests that students who rate one application feature highly often rate others similarly. However, these ratings – along with demographic data like gender – had minimal predictive value and appeared at the bottom of the feature importance rankings in Fig. 5. These findings indicate that prediction of program year is driven primarily by user type and NLP-based features. These are likely to reflect skills acquired during the degree program or personality traits that evolve over time. The comparison between predicted and actual program year can thus offer meaningful insights into students’ academic maturity and potential risk factors. RQ3: What insights into students’ maturity or personality development can be gained through text analysis parameters and player types? Identifying student maturity or personality development requires: 1) a well-defined link between personality traits and the parameters used; and 2) evidence of correlations between these traits and program year. Our analysis confirmed the relevance of lexical diversity and readability in predicting academic progress, These NLP-derived features reflect writing ability development consistent with psycholinguistic theories that describe gradual semantic and syntactic growth through adolescence and early adulthood (Berman, 2007, 2008; Berman & Nir-sagiv, 2007; Grimshaw et al., 1998; Hansson et al., 2016; Karmiloff-Smith, 1992; Nippold, 2000, 2002, 2016; Ravid & Tolchinsky, 2002; Ravid, 2005; Ravid et al., 2002; Rimmer, 2008). No single indicator serves as a definite maturity marker, but the aggregate output from NLP offers a holistic view of development (De Bot & Schrauf, 2010). Open-ended questions encouraging critical reflection were particularly effective at eliciting complex language and metacognitive awareness (Aminah, 2024; Sato, 2022). One such item, addressing “Time consumption readability,” proved especially predictive. It required students to justify their choice between virtual reconstructions and traditional teaching methods – blending expository and argumentative writing with introspection, practical reasoning and creative thinking (Aminah, 2024; Berman & Nir-sagiv, 2007; Dimitrova, 2024; Hidayat et al., 2018; Ragnarsdóttir et al., 2002; Rahayuni̇Ngsi̇H et al., 2021; Ravid, 2005, 2006; Sato, 2022; Zhang et al., 2024). Similar prompts have been used to assess personality traits and maturity through problem-solving and critical thinking (Arumningsih et al., 2023; Dai et al., 2022; Y. Liu et al., 2022). The cognitive and linguistic complexity of such questions makes them useful not only for predictive modeling but also for identifying growth in maturity and writing skills. Variability in the linguistic sophistication of students’ responses to these tasks may reveal differences in developmental stage, making such items valuable for current prediction and for future longitudinal tracking. Regarding Hexad user types, no strong direct correlations were found with program year. However, "Socializer," "Philanthropist" and "Free Spirit" types emerged as slightly more predictive than others, suggesting motivational shifts across academic years (Brandt et al., 2020). Prior studies indicate that altruistic traits, as seen in “Philanthropists,” tend to increase with age (Tondello et al., 2019), aligning with our data showing an association between program year and age. This supports the interpretation of motivational evolution as part of student maturation. The "Achiever" type may reflect growing awareness of academic performance and career planning, echoing findings from the Job Outlook 2020 report by the National Association of Colleges and Employers (National Association of Colleges and Employers, 2019). Competitive traits may be present early in studies but often intensify as graduation approaches, with a sharper focus on employability. Similarly, the "Player" profile may evolve, with motivations shifting as students advance in their academic journey (Santos et al., 2023). Gender was not a significant predictor of program year, suggesting that the maturity and personality-related parameters captured by the model do not require gender-based differentiation. This indicates similar patterns of development, concerns and motivations across male and female students. 5.2. Educational Implications The findings offer multiple implications for teaching and learning. First, integrating AI-driven predictive diagnostics into formative assessment can support personalized, developmentally appropriate instruction (Black & Wiliam, 1998; Ifenthaler & Yau, 2020; Nicol & Macfarlane‐Dick, 2006). Such tools can help identify students at risk and inform differentiated instruction or peer-based strategies (Tempelaar et al., 2017). Feedback on linguistic and motivational development may also promote self-awareness, self-regulation, and intrinsic motivation (Panadero, 2017; Zimmerman, 2002). However, maturity analysis must be applied with care. Awareness of academic underperformance could negatively affect student confidence. Used strategically, though, these insights can guide instructional measures that anticipate students’ development risks and improve learning outcomes. This model can contribute to students’ soft skills and academic growth in four key areas: Self-awareness: By understanding their learning styles and gamification profile, students can better recognize and manage their strengths and weaknesses. For instance, a student not yet motivated by long-term academic or career goals might adjust their strategies after reflecting on their “Player” versus “Achiever” profile. Teachers can support this by highlighting overlooked challenges and achievements. Self-discipline: The analysis encourages better time management and structured learning. Sophisticated writing, associated with higher maturity, requires planning, critical thinking and self-monitoring. This evolution also enhances learning efficiency, as tasks completed without critical engagement may be superficial and ineffective in fostering real academic growth. Therefore, upon detecting a writing skill weakness, teachers can design writing tasks to target underdeveloped writing skills and promote deeper engagement and critical thinking. Motivation: Understanding one’s own developmental stage may foster intrinsic motivation, reducing reliance on external rewards. As motivation develops longitudinally and personal growth is often imperceptible over short periods, students may not be aware of their progress. Teachers can support this by explicitly recognizing student progress, helping learners appreciate their academic evolution. This can be a useful method to enhance motivation in students who have difficulty in perceiving their own improvements. Collaboration: Awareness of peers’ diverse maturity levels can improve group work. A comparative analysis of students at different levels of academic maturity can help identify successful learning strategies. By sharing approaches used by peers with higher academic development, students can adopt more effective learning strategies to achieve their goals. In this line, our predictive methodology can help create structured collaborative activities that pair students with differing developmental profiles. This may generate mutually beneficial learning dynamics. These insights allow instructors and academic advisors to better align teaching strategies with students’ evolving traits, supporting more effective and inclusive environments. 6. Conclusions This study highlights the effectiveness of artificial intelligence in analyzing students’ open-ended responses to predict academic progression and identifying those at risk of underperformance. Our results demonstrate that the lexical diversity and readability features, along with gamification profiles, are key predictors of students’ program year, with models reaching acceptable accuracy. As such, the system can help detect students falling behind – i.e., when the predicted year is lower than the actual one – enabling timely feedback and intervention. Beyond predictive accuracy, the model offers pedagogical value by estimating students’ cognitive and motivational maturity. The ability to infer program year from textual and gamification data supports the design of instruction that aligns with students’ developmental stages and allows for early, personalized support. These findings are consistent with constructivist and self-regulated learning theories, which emphasize adapting instruction to students’ evolving competencies. Although lexical and readability features proved central to the prediction model, and gamification profiles improved prediction performance, no single parameter showed a clear linear relationship with academic stage. This suggests motivational changes as students progress and supports the model’s developmental relevance. Besides, it reflects the complex and multifaceted nature of maturity, which often eludes traditional linear metrics. AI’s ability to model these non-linear dynamics strengthens its relevance for educational applications. The model’s strong performance with AI-assisted translations also broadens its applicability across multilingual contexts and allows responses written in students’ native languages. When combined with AI tools capable of reliably assessing lexical and readability metrics, this approach can be applied to diverse educational settings. The insights offered by this model extend beyond grade prediction. They support formative feedback loops and developmentally aligned instruction, allowing educators to recognize and assist underperforming students. Looking ahead, several future research directions emerge. Longitudinal studies could track changes in students’ linguistic and motivational profiles throughout their academic programs, clarifying how these features evolve over time. Cross-disciplinary applications may allow the model to be trained for fields where writing tasks are less central. Translation sensitivity could be further explored, especially the impact of translation accuracy on key metrics like TTR and FKG. Additionally, comparative studies with different AI tools and multilingual datasets could improve robustness and generalizability. Abbreviations AI – Artificial Intelligence FKG – Flesch Kincaid Grade MAE – Mean Absolute Error MSE – Mean Squared Eror NLP – Natural Language Processing SRL – Self-Regulated Learning TTR – Type Token Ratio Declarations Author Contribution L. S.: Data curation, Investigation, Methodology, Resources, Software, Visualization, Writing – original draft. B. N.: Data curation, Investigation, Resources, Writing – original draft & review and editing. M. E. C.: Investigation, Writing – original draft & review and editing. S. G. M.: Conceptualization, Formal Analysis, Investigation, Methodology, Project administration, Supervision, Writing – review and editing. Data Availability The datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request. References Afzaal, M., Nouri, J., Zia, A., Papapetrou, P., Fors, U., Wu, Y., Li, X., & Weegar, R. (2021). Explainable AI for Data-Driven Feedback and Intelligent Action Recommendations to Support Students Self-Regulation. Frontiers in Artificial Intelligence , 4 , 723447. https://doi.org/10.3389/frai.2021.723447 Al Nabhani, F., Hamzah, M. B., & Abuhassna, H. (2025). The role of artificial intelligence in personalizing educational content: Enhancing the learning experience and developing the teacher’s role in an integrated educational environment. Contemporary Educational Technology , 17 (2), ep573. https://doi.org/10.30935/cedtech/16089 Aminah, M. (2024). Project-Based Learning: Optimizing Students’ Critical Thinking Skills. Media Bina Ilmiah , 18 (12), 3177–3184. Anderson, J. E. (1941). Principles of growth and maturity in language. The Elementary English Review , 18 (7), 250–277. Arumningsih, E., Setyawati, R., & Murtianto, Y.H. (2023). Students’ Creative Thinking Ability in Solving Open-Ended Problems Based on Personality Type. Hipotenusa Journal of Mathematical Society , 5 , 121–131. https://doi.org/10.18326/hipotenusa.v5i2.280 Ashenafi, M. M., Riccardi, G., & Ronchetti, M. (2015). Predicting students’ final exam scores from their course activities. 2015 IEEE Frontiers in Education Conference (FIE) , 1–9. https://doi.org/10.1109/FIE.2015.7344081 Baker, R. S. J. d., & Yacef, K. (2009). The State of Educational Data Mining in 2009: A Review and Future Visions . https://doi.org/10.5281/ZENODO.3554657 Barrett, A., & Pack, A. (2023). Not quite eye to A.I.: Student and teacher perspectives on the use of generative artificial intelligence in the writing process. International Journal of Educational Technology in Higher Education , 20 (1), 59. https://doi.org/10.1186/s41239-023-00427-0 Bartle, R. (1996). Hearts, clubs, diamonds, spades: Players who suit MUDs. MUD Research , 1 , 19–20. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , 610–623. https://doi.org/10.1145/3442188.3445922 Berman, R. A. (2007). Developing Linguistic Knowledge and Language Use Across Adolescence. In E. Hoff & M. Shatz (Eds.), Blackwell Handbook of Language Development (pp. 347–367). Blackwell Publishing Ltd. https://doi.org/10.1002/9780470757833.ch17 Berman, R. A. (2008). The psycholinguistics of developing text construction. Journal of Child Language , 35 (4), 735–771. https://doi.org/10.1017/S0305000908008787 Berman, R. A. (2017). Language Development and Literacy. In R. J. R. Levesque (Ed.), Encyclopedia of Adolescence (pp. 1–11). Springer International Publishing. https://doi.org/10.1007/978-3-319-32132-5_19-2 Berman, R. A., & Nir, B. (2009). Cognitive and linguistic factors in evaluating text quality: Global versus local? In V. Evans & S. Pourcel (Eds.), Human Cognitive Processing (Vol. 24, pp. 421–440). John Benjamins Publishing Company. https://doi.org/10.1075/hcp.24.26ber Berman, R. A., & Nir-sagiv, B. (2007). Comparing Narrative and Expository Text Construction Across Adolescence: A Developmental Paradox. Discourse Processes , 43 (2), 79–120. https://doi.org/10.1080/01638530709336894 Black, P., & Wiliam, D. (1998). Assessment and Classroom Learning. Assessment in Education: Principles, Policy & Practice , 5 (1), 7–74. https://doi.org/10.1080/0969595980050102 Bloom, B. S. (1968). Learning for mastery. Instruction and curriculum. Regional education laboratory for the Carolinas and Virginia, topical papers and reprints, number 1. Evaluation Comment, 1(2) . https://eric.ed.gov/?id=eD053419 Bloom, B. S. (1971). Mastery learning. In J. H. Block (Ed.), Mastery learning, theory and practice (pp. 47-63). New York: Holt, Rinehart, and Winston. Bloom, B. S., Madaus, G. F., & Hastings, J. T. (1981). Evaluation to Improve Learning. New York: McGraw-Hill. Boroowa, A., & Herodotou, C. (2022). Learning Analytics in Open and Distance Higher Education: The Case of the Open University UK. In P. Prinsloo, S. Slade, & M. Khalil (Eds.), Learning Analytics in Open and Distributed Learning (pp. 47–62). Springer Nature Singapore. https://doi.org/10.1007/978-981-19-0786-9_4 Brandt, N. D., Lechner, C. M., Tetzner, J., & Rammstedt, B. (2020). Personality, cognitive ability, and academic performance: Differential associations across school subjects and school tracks. Journal of Personality , 88 (2), 249–265. https://doi.org/10.1111/jopy.12482 Bremholm, J., Bundsgaard, J., & Kabel, K. (2022). Proficiency scales for early writing development. Writing & Pedagogy , 13 (1–3), 121–154. https://doi.org/10.1558/wap.21490 Britz, D., Goldie, A., Luong, M.-T., & Le, Q. (2017). Massive Exploration of Neural Machine Translation Architectures. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , 1442–1451. https://doi.org/10.18653/v1/D17-1151 Broadbent, J., & Poon, W. L. (2015). Self-regulated learning strategies & academic achievement in online higher education learning environments: A systematic review. The Internet and Higher Education , 27 , 1–13. https://doi.org/10.1016/j.iheduc.2015.04.007 Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners . arXiv. https://doi.org/10.48550/ARXIV.2005.14165 Bruner, J. S. (1960). The culture of education . Harvard University Press. Busch, D. (2024). AI translation and intercultural communication: New questions for a new field of research . https://doi.org/10.31235/osf.io/r3zdx Castilho, S., Moorkens, J., Gaspari, F., Calixto, I., Tinsley, J., & Way, A. (2017). Is Neural Machine Translation the New State of the Art? The Prague Bulletin of Mathematical Linguistics , 108 (1), 109–120. https://doi.org/10.1515/pralin-2017-0013 Chassignol, M., Khoroshavin, A., Klimova, A., & Bilyatdinova, A. (2018). Artificial Intelligence trends in education: A narrative overview. Procedia Computer Science , 136 , 16–24. https://doi.org/10.1016/j.procs.2018.08.233 Chen, L., Chen, P., & Lin, Z. (2020). Artificial Intelligence in Education: A Review. IEEE Access , 8 , 75264–75278. https://doi.org/10.1109/ACCESS.2020.2988510 Crossley, S. (2020). Linguistic features in writing quality and development: An overview. Journal of Writing Research , 11 (vol. 11 issue 3), 415–443. https://doi.org/10.17239/jowr-2020.11.03.01 Crossley, S. A., & Kim, M. (2022). Linguistic Features of Writing Quality and Development: A Longitudinal Approach. The Journal of Writing Analytics , 6 (1), 59–93. https://doi.org/10.37514/JWA-J.2022.6.1.04 Dai, Y., Jayaratne, M., & Jayatilleke, B. (2022). Explainable Personality Prediction Using Answers to Open-Ended Interview Questions. Frontiers in Psychology , 13 , 865841. https://doi.org/10.3389/fpsyg.2022.865841 De Bot, K., & Schrauf, R. W. (2010). Language Development Over the Lifespan (0 ed.). Routledge. https://doi.org/10.4324/9780203880937 De Feyter, T., Caers, R., Vigna, C., & Berings, D. (2012). Unraveling the impact of the Big Five personality traits on academic performance: The moderating and mediating effects of self-efficacy and academic motivation. Learning and Individual Differences , 22 (4), 439–448. https://doi.org/10.1016/j.lindif.2012.03.013 Deo, R. C., Yaseen, Z. M., Al-Ansari, N., Nguyen-Huy, T., Langlands, T. A. M., & Galligan, L. (2020). Modern Artificial Intelligence Model Development for Undergraduate Student Performance Prediction: An Investigation on Engineering Mathematics Courses. IEEE Access , 8 , 136697–136724. https://doi.org/10.1109/ACCESS.2020.3010938 Devedžic, V. (2004). Web Intelligence and Artificial Intelligence in Education. Journal of Educational Technology and Society , 7 , 29–39. Dimitrova, K. (2024). UTILIZATION OF STEM APPROACH IN PRIMARY EDUCATION – THEORETICAL FOUNDATIONS AND METHODOLOGICAL SOLUTIONS . 8448–8458. https://doi.org/10.21125/edulearn.2024.2011 Drachen, A., Canossa, A., & Yannakakis, G., N. (2009). Player Modeling using Hamari Self-Organization in Tomb Raider: Underworld . In Proceedings of the IEEE Symposium on Computational Intelligence and Games, 7-10. Milan, Italy, September. 10.1109/CIG.2009.5286500 Durrant, P., & Brenchley, M. (2019). Development of vocabulary sophistication across genres in English children’s writing. Reading and Writing , 32 (8), 1927–1953. https://doi.org/10.1007/s11145-018-9932-8 European Parliament. (n.d.). Artificial Intelligence Act, Corrigendum, 19 April 2024. Interinstitutional File: 2021/0106(COD) . Retrieved September 1, 2024, from https://www.europarl.europa.eu/doceo/document/TA-9-2024-0138-FNL-COR01_EN.pdf Faza, A., & Lestari, I. A. (2025). Self-Regulated Learning in the Digital Age: A Systematic Review of Strategies, Technologies, Benefits, and Challenges. The International Review of Research in Open and Distributed Learning , 26 (2), 23–58. https://doi.org/10.19173/irrodl.v26i2.8119 Glaser, R. (1994). Instructional Technology and the Measwement of Learning Outcomes: Some Questions 1 . Educational Measurement: Issues and Practice , 13 (4), 6–8. https://doi.org/10.1111/j.1745-3992.1994.tb00561.x Grimshaw, G. M., Adelstein, A., Bryden, M. P., & MacKinnon, G. E. (1998). First-Language Acquisition in Adolescence: Evidence for a Critical Period for Verbal Language Development. Brain and Language , 63 (2), 237–255. https://doi.org/10.1006/brln.1997.1943 Guskey, T. R. (1985). Implementing Mastery Learning . Hall-Mills, S., & Apel, K. (2015). Linguistic Feature Development Across Grades and Genre in Elementary Writing. Language, Speech, and Hearing Services in Schools , 46 (3), 242–255. https://doi.org/10.1044/2015_LSHSS-14-0043 Hamari, J., & Tuunanen, J. (2014). Player Types: A Meta-synthesis. Transactions of the Digital Games Research Association , 1 (2). https://doi.org/10.26503/todigra.v1i2.13 Hansson, K., Bååth, R., Löhndorf, S., Sahlén, B., & Sikström, S. (2016). Quantifying Semantic Linguistic Maturity in Children. Journal of Psycholinguistic Research , 45 (5), 1183–1199. https://doi.org/10.1007/s10936-015-9398-7 Hidayat, T., Susilaningsih, E., & Kurniawan, C. (2018). The effectiveness of enrichment test instruments design to measure students’ creative thinking skills and problem-solving. Thinking Skills and Creativity , 29 , 161–169. https://doi.org/10.1016/j.tsc.2018.02.011 Holmes, W., Bialik, M., & Fadel, C. (2019). Artificial intelligence in education: Promises and implications for teaching and learning . Center for Curriculum Redesign. Hu, Z., Yang, Z., Liang, X., Salakhutdinov, R., & Xing, E. P. (2017). Toward Controlled Generation of Text . arXiv. https://doi.org/10.48550/ARXIV.1703.00955 Ifenthaler, D., & Yau, J. Y.-K. (2020). Utilising learning analytics to support study success in higher education: A systematic review. Educational Technology Research and Development , 68 (4), 1961–1990. https://doi.org/10.1007/s11423-020-09788-z Ji, M., Bouillon, P., & Seligman, M. (2023). Translation Technology in Accessible Health Communication (1st ed.). Cambridge University Press. https://doi.org/10.1017/9781108938976 Kahraman, H. T., Sagiroglu, S., & Colak, I. (2010). Development of adaptive and intelligent web-based educational systems. 2010 4th International Conference on Application of Information and Communication Technologies , 1–5. https://doi.org/10.1109/ICAICT.2010.5612054 Karmiloff-Smith, A. (1992). Beyond modularity: A developmental perspective on cognitive science . Cambridge: MIT Press. Karunasingha, D. S. K. (2022). Root mean square error or mean absolute error? Use their ratio as well. Information Sciences , 585 , 609–629. https://doi.org/10.1016/j.ins.2021.11.036 Katzenberger, I. (2005). The Super-Structure of Written Expository Texts—A Developmental Perspective. In D. D. Ravid & H. B.-Z. Shyldkrot (Eds.), Perspectives on Language and Language Development (pp. 327–336). Springer US. https://doi.org/10.1007/1-4020-7911-7_24 Kluwe, R., & Logan, G. D. (2000). Executive control . Psychological Research, 63 [Special issue], 3-4. Koehn, P., & Knowles, R. (2017). Six Challenges for Neural Machine Translation . arXiv. https://doi.org/10.48550/ARXIV.1706.03872 Kolluru, V., Mungara, S., & Chintakunta, A. N. (2018). Adaptive Learning Systems: Harnessing AI for Customized Educational Experiences. International Journal of Computational Science and Information Technology , 6 (3), 13–26. https://doi.org/10.5121/ijcsity.2018.6302 Kumar, H., Musabirov, I., Reza, M., Shi, J., Wang, X., Williams, J. J., Kuzminykh, A., & Liut, M. (2023). Impact of Guidance and Interaction Strategies for LLM Use on Learner Performance and Perception . https://doi.org/10.48550/ARXIV.2310.13712 Larrabee Sønderlund, A., Hughes, E., & Smith, J. (2019). The efficacy of learning analytics interventions in higher education: A systematic review. British Journal of Educational Technology , 50 (5), 2594–2618. https://doi.org/10.1111/bjet.12720 Lavoué, É., Ju, Q., Hallifax, S., & Serna, A. (2021). Analyzing the relationships between learners’ motivation and observable engaged behaviors in a gamified learning environment. International Journal of Human-Computer Studies , 154 , 102670. https://doi.org/10.1016/j.ijhcs.2021.102670 Lim, W. M., Gunasekara, A., Pallant, J. L., Pallant, J. I., & Pechenkina, E. (2023). Generative AI and the future of education: Ragnarök or reformation? A paradoxical perspective from management educators. The International Journal of Management Education , 21 (2), 100790. https://doi.org/10.1016/j.ijme.2023.100790 Liu, S., Yu, Z., Huang, F., Bulbulia, Y., Bergen, A., & Liut, M. (2024). Can Small Language Models With Retrieval-Augmented Generation Replace Large Language Models When Learning Computer Science? Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 , 388–393. https://doi.org/10.1145/3649217.3653554 Liu, Y., Wang, T., Bo, H., & Zhang, N. (2022). The Influence of Personality on Epistemic Network in an Open-Ended Question Discussion Scenario. 2022 4th International Conference on Computer Science and Technologies in Education (CSTE) , 215–220. https://doi.org/10.1109/CSTE55932.2022.00046 Lopez, C. E., & Tucker, C. S. (2019). The effects of player type on performance: A gamification case study. Computers in Human Behavior , 91 , 333–345. https://doi.org/10.1016/j.chb.2018.10.005 Lounsbury, J. W., Sundstrom, E., Loveland, J. M., & Gibson, L. W. (2003). Intelligence, “Big Five” personality traits, and work drive as predictors of course grade. Personality and Individual Differences , 35 (6), 1231–1239. https://doi.org/10.1016/S0191-8869(02)00330-6 Ludwig, S., Rausch, A., Deutscher, V., & Seifried, J. (2024). Predicting problem-solving success in an office simulation applying N-grams and a random forest to behavioral process data. Computers & Education , 218 , 105093. https://doi.org/10.1016/j.compedu.2024.105093 MacArthur, C. A., Jennings, A., & Philippakos, Z. A. (2019). Which linguistic features predict quality of argumentative writing for college basic writers, and how do those features change with instruction? Reading and Writing , 32 (6), 1553–1574. https://doi.org/10.1007/s11145-018-9853-6 Malvern, D., Richards, B., Chipere, N., & Durán, P. (2004). Lexical Diversity and Language Development . Palgrave Macmillan UK. https://doi.org/10.1057/9780230511804 Mammadov, S. (2022). Big Five personality traits and academic performance: A meta‐analysis. Journal of Personality , 90 (2), 222–255. https://doi.org/10.1111/jopy.12663 Marczewski, A. (2015). User Types . In Even Ninja Monkeys Like to Play: Gamification, Game Thinking and Motivational Design (1st ed.); CreateSpace Independent Publishing Platform: Scotts Valley, CA, USA; Volume 65–80, 240–290. https://www.gamified.uk/user-types (October 2018 update). Mirkin, S., Nowson, S., Brun, C., & Perez, J. (2015). Motivating Personality-aware Machine Translation. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , 1102–1108. https://doi.org/10.18653/v1/D15-1130 Moore, D. S., McCabe, G. P., & Craig, B. A. (2009). Introduction to the Practice of Statistics . (Vol. 4). New York: WH Freeman. Mornar, M., Marušić, I., & Šabić, J. (2022). Academic self-efficacy and learning strategies as mediators of the relation between personality and elementary school students’ achievement. European Journal of Psychology of Education , 37 (4), 1237–1254. https://doi.org/10.1007/s10212-021-00576-8 Nababan, M., & Nuraeni, A. (2012). Pengembangan model penilaian kualitas terjemahan . Surakarta: Kajian Linguistik dan Sastra. 24(1), 39-57. Naqshbandi, M. M., Ainin, S., Jaafar, N. I., & Mohd Shuib, N. L. (2017). To Facebook or to Face Book? An investigation of how academic performance of different personalities is affected through the intervention of Facebook usage. Computers in Human Behavior , 75 , 167–176. https://doi.org/10.1016/j.chb.2017.05.012 National Association of Colleges and Employers. (2019). Job Outlook 2020. Arkansas: Bethlehem . https://in.nau.edu/wp-content/uploads/sites/204/2020-nace-job-outlook.pdf Newmark, P. (1988). A Textbook on Translation . New York: Prentice-Hall International. Nicol, D. J., & Macfarlane‐Dick, D. (2006). Formative assessment and self‐regulated learning: A model and seven principles of good feedback practice. Studies in Higher Education , 31 (2), 199–218. https://doi.org/10.1080/03075070600572090 Nippold, M. A. (2000). Language Development during the Adolescent Years: Aspects of Pragmatics, Syntax, and Semantics. Topics in Language Disorders , 20 (2), 15–28. https://doi.org/10.1097/00011363-200020020-00004 Nippold, M. A. (2002). Lexical learning in school-age children, adolescents, and adults: A process where language and literacy converge. Journal of Child Language , 29 (2), 449–488. https://doi.org/10.1017/S0305000902275340 Nippold, M. A. (2016). Later language development: School-age children, adolescents, and young adults . Austin, TX: PRO-ED. O’Connor, M. C., & Paunonen, S. V. (2007). Big Five personality predictors of post-secondary academic performance. Personality and Individual Differences, 43(5), 971-990. OECD. (2023). Digital Education Outlook 2023 . https://www.oecd.org/en/publications/oecd-digital-education-outlook-2023_c74f03de-en.html Panadero, E. (2017). A Review of Self-regulated Learning: Six Models and Four Directions for Research. Frontiers in Psychology , 8 , 422. https://doi.org/10.3389/fpsyg.2017.00422 Pankiewicz, M., & Baker, R. S. (2023). Large Language Models (GPT) for automating feedback on programming assignments . arXiv. https://doi.org/10.48550/ARXIV.2307.00150 Peredo, R., Canales, A., Menchaca, A., & Peredo, I. (2011). Intelligent Web-based education system for adaptive learning. Expert Systems with Applications , 38 (12), 14690–14702. https://doi.org/10.1016/j.eswa.2011.05.013 Phillips, P., Abraham, C., & Bond, R. (2003). Personality, cognition, and university students’ examination performance. European Journal of Personality , 17 (6), 435–448. https://doi.org/10.1002/per.488 Pintrich, P. R., & De Groot, E. V. (1990). Motivational and self-regulated learning components of classroom academic performance. Journal of Educational Psychology , 82 (1), 33. Pokrivcakova, S. (2019). Preparing teachers for the application of AI-powered technologies in foreign language education. Journal of Language and Cultural Education , 7 (3), 135–153. https://doi.org/10.2478/jolace-2019-0025 Poropat, A. E. (2014). A meta‐analysis of adult‐rated child personality and academic performance in primary education. British Journal of Educational Psychology , 84 (2), 239–252. https://doi.org/10.1111/bjep.12019 Prabhumoye, S., Tsvetkov, Y., Black, A. W., & Salakhutdinov, R. (2018). Style Transfer Through Multilingual and Feedback-Based Back-Translation . arXiv. https://doi.org/10.48550/ARXIV.1809.06284 Proverbio, A. M., & Zani, A. (2005). Developmental changes in the linguistic brain after puberty. Trends in Cognitive Sciences , 9 (4), 164–167. https://doi.org/10.1016/j.tics.2005.02.001 Rabinovich, E., Mirkin, S., Patel, R. N., Specia, L., & Wintner, S. (2016). Personalized Machine Translation: Preserving Original Author Traits . arXiv. https://doi.org/10.48550/ARXIV.1610.05461 Raffaghelli, J. E., Rodríguez, M. E., Guerrero-Roldán, A.-E., & Bañeres, D. (2022). Applying the UTAUT model to explain the students’ acceptance of an early warning system in Higher Education. Computers & Education , 182 , 104468. https://doi.org/10.1016/j.compedu.2022.104468 Ragnarsdóttir, H., Aparici, M., Cahana-Amitay, D., Van Hell, J. G., & Viguié-Simon, A. (2002). Verbal structure and content in written discourse: Expository and narrative texts. Written Language & Literacy , 5 (1), 95–126. https://doi.org/10.1075/wll.5.1.05rag Rahayuni̇Ngsi̇H, S., Si̇Rajuddi̇N, S., & Ikram, M. (2021). Using Open-ended Problem-solving Tests to Identify Students’ Mathematical Creative Thinking Ability. Participatory Educational Research , 8 (3), 285–299. https://doi.org/10.17275/per.21.66.8.3 Ravid, D. (2005). Emergence of Linguistic Complexity in Later Language Development: Evidence from Expository Text Construction. In D. D. Ravid & H. B.-Z. Shyldkrot (Eds.), Perspectives on Language and Language Development (pp. 337–355). Springer US. https://doi.org/10.1007/1-4020-7911-7_25 Ravid, D. (2006). Semantic development in textual contexts during the school years: Noun Scale analyses. Journal of Child Language , 33 (4), 791–821. https://doi.org/10.1017/S0305000906007586 Ravid, D., & Tolchinsky, L. (2002). Developing linguistic literacy: A comprehensive model. Journal of Child Language, 29, 417–447 . Ravid, D., Van Hell, J. G., Rosado, E., & Zamora, A. (2002). Subject NP patterning in the development of text production: Speech and writing. Written Language & Literacy , 5 (1), 69–93. https://doi.org/10.1075/wll.5.1.04rav Reilly, J., Zamora, A., & Mcgivern, R. (2005). Acquiring perspective in English: The development of stance. Journal of Pragmatics , 37 (2), 185–208. https://doi.org/10.1016/S0378-2166(04)00191-2 Rello, L., Baeza-Yates, R., Ali, A., Bigham, J. P., & Serra, M. (2020). Predicting risk of dyslexia with an online gamified test. PLOS ONE , 15 (12), e0241687. https://doi.org/10.1371/journal.pone.0241687 Richards, B. J. & Malvern, D. D. (1997). The new Bulmershe papers. Quantifying lexical diversity in the study of language development . Reading: The University of Reading. Rimmer, W. (2008). Putting grammatical complexity in context. Literacy , 42 (1), 29–35. https://doi.org/10.1111/j.1467-9345.2008.00478.x Río Lugo, N. D. (1996). David R. Olson, The world on paper. The conceptual and cognitive implications of writing and reading. Cambridge University Press, Cambridge, 1994; 318 pp. Nueva Revista de Filología Hispánica (NRFH) , 44 (1), 197–200. https://doi.org/10.24201/nrfh.v44i1.1919 Rosselli, M., Ardila, A., Matute, E., & Vélez-Uribe, I. (2014). Language Development across the Life Span: A Neuropsychological/Neuroimaging Perspective. Neuroscience Journal , 2014 , 1–21. https://doi.org/10.1155/2014/585237 Rovira, S., Puertas, E., & Igual, L. (2017). Data-driven system to predict academic grades and dropout. PLOS ONE , 12 (2), e0171207. https://doi.org/10.1371/journal.pone.0171207 Rus, V., D’Mello, S., Hu, X., & Graesser, A. C. (2013). Recent Advances in Conversational Intelligent Tutoring Systems. AI Magazine , 34 (3), 42–54. https://doi.org/10.1609/aimag.v34i3.2485 Sajja, R., Sermet, Y., Cwiertny, D., & Demir, I. (2023). Platform-independent and curriculum-oriented intelligent assistant for higher education. International Journal of Educational Technology in Higher Education , 20 (1), 42. https://doi.org/10.1186/s41239-023-00412-7 Salas-Pilco, S. Z., & Yang, Y. (2022). Artificial intelligence applications in Latin American higher education: A systematic review. International Journal of Educational Technology in Higher Education , 19 (1), 21. https://doi.org/10.1186/s41239-022-00326-w Samuel Girard, Jill-Jênn Vie, Françoise Tort, & Amel Bouzeghoub. (2024). Optimizing Human Learning using Reinforcement Learning . https://doi.org/10.5281/ZENODO.12730016 Santos, A. C. G., Oliveira, W., Hamari, J., Joaquim, S., & Isotani, S. (2023). The Consistency of Gamification User Types: A Study on the Change of Preferences over Time. Proceedings of the ACM on Human-Computer Interaction , 7 (CHI PLAY), 1253–1281. https://doi.org/10.1145/3611068 Sato, T. (2022). Assessing critical thinking through L2 argumentative essays: An investigation of relevant and salient criteria from raters’ perspectives. Language Testing in Asia , 12 (1), 9. https://doi.org/10.1186/s40468-022-00159-4 Sennrich, R., Haddow, B., & Birch, A. (2016). Controlling Politeness in Neural Machine Translation via Side Constraints. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 35–40. https://doi.org/10.18653/v1/N16-1005 Shakir, A., & Obeidat, H. (1991). Maturity in AFL Student-Written Texts: A case study. Al-’Arabiyya, 24, 65–81. Http://Www.Jstor.Org/Stable/43192653 . Shen, T., Lei, T., Barzilay, R., & Jaakkola, T. (2017). Style Transfer from Non-Parallel Text by Cross-Alignment . arXiv. https://doi.org/10.48550/ARXIV.1705.09655 Shuliang, M. (2025). Developmental constructivism . Springer Nature Singapore. Slade, S., & Prinsloo, P. (2013). Learning Analytics: Ethical Issues and Dilemmas. American Behavioral Scientist , 57 (10), 1510–1529. https://doi.org/10.1177/0002764213479366 Subirats, L., Nousiainen, T., Hooda, A., Rubio-Andrada, L., Fort, S., Vesisenaho, M., & Sacha, G. M. (2023). Gamification Based on User Types: When and Where It Is Worth Applying. Applied Sciences , 13 (4), 2269. https://doi.org/10.3390/app13042269 Subirats, L., Palacios Corral, A., Pérez-Ruiz, S., Fort, S., & Sacha, G.-M. (2023). Temporal analysis of academic performance in higher education before, during and after COVID-19 confinement using artificial intelligence. PLOS ONE , 18 (2), e0282306. https://doi.org/10.1371/journal.pone.0282306 Sun, K., & Xiong, W. (2019). A computational model for measuring discourse complexity. Discourse Studies , 21 (6), 690–712. https://doi.org/10.1177/1461445619866985 Tempelaar, D. T., Rienties, B., & Nguyen, Q. (2017). Towards Actionable Learning Analytics Using Dispositions. IEEE Transactions on Learning Technologies , 10 (1), 6–16. https://doi.org/10.1109/TLT.2017.2662679 Timms, M. J. (2016). Letting Artificial Intelligence in Education Out of the Box: Educational Cobots and Smart Classrooms. International Journal of Artificial Intelligence in Education , 26 (2), 701–712. https://doi.org/10.1007/s40593-016-0095-y Tolchinsky, L., & Rosado, E. (2005). The effect of literacy, text type, and modality on the use of grammatical means for agency alternation in Spanish. Journal of Pragmatics , 37 (2), 209–237. https://doi.org/10.1016/S0378-2166(04)00195-X Tomlinson, C. A. (2017). How to differentiate instruction in academically diverse classrooms (3rd ed.). ASCD. Tondello, G. F., Mora, A., Marczewski, A., & Nacke, L. E. (2019). Empirical validation of the Gamification User Types Hexad scale in English and Spanish. International Journal of Human-Computer Studies , 127 , 95–111. https://doi.org/10.1016/j.ijhcs.2018.10.002 Tondello, G. F., Wehbe, R. R., Diamond, L., Busch, M., Marczewski, A., & Nacke, L. E. (2016). The Gamification User Types Hexad Scale. Proceedings of the 2016 Annual Symposium on Computer-Human Interaction in Play , 229–243. https://doi.org/10.1145/2967934.2968082 UNESCO. (2019). Beijing Consensus on Artificial Intelligence and Education . https://unesdoc.unesco.org/ark:/48223/pf0000368303 Vanmassenhove, E. (2024). Gender Bias in Machine Translation and The Era of Large Language Models . arXiv. https://doi.org/10.48550/ARXIV.2401.10016 Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes . Harvard University Press. Wagner, R. K., Puranik, C. S., Foorman, B., Foster, E., Wilson, L. G., Tschinkel, E., & Kantor, P. T. (2011). Modeling the development of written language. Reading and Writing , 24 (2), 203–220. https://doi.org/10.1007/s11145-010-9266-7 Wijaksono, R. N., Hilman, E. H., & Mustolih, A. (2022). TRANSLATION METHODS AND QUALITY OF IDIOMATIC EXPRESSION IN MY SISTER’S KEEPER MOVIE. JURNAL BASIS , 9 (1), 73–84. https://doi.org/10.33884/basisupb.v9i1.5428 Zhang, Z., Zhang, E., Liu, H., & Han, S. (2024). Examining the association between discussion strategies and learners’ critical thinking in asynchronous online discussion. Thinking Skills and Creativity , 53 , 101588. https://doi.org/10.1016/j.tsc.2024.101588 Zimmerman, B. J. (2002). Becoming a Self-Regulated Learner: An Overview. Theory Into Practice , 41 (2), 64–70. https://doi.org/10.1207/s15430421tip4102_2 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7801827","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":545758129,"identity":"b443cb44-407f-4e4f-aad9-343a4ab88b20","order_by":0,"name":"Laia Subirats","email":"","orcid":"","institution":"Open University of Catalonia","correspondingAuthor":false,"prefix":"","firstName":"Laia","middleName":"","lastName":"Subirats","suffix":""},{"id":545758130,"identity":"0611683a-9cd8-4f3a-9e94-031faff5461d","order_by":1,"name":"Beatriz Narbona","email":"","orcid":"","institution":"Autonomous University of Madrid","correspondingAuthor":false,"prefix":"","firstName":"Beatriz","middleName":"","lastName":"Narbona","suffix":""},{"id":545758131,"identity":"50ed88fd-d833-45b7-b7e0-99c02ed0996f","order_by":2,"name":"María Elena Cuenca","email":"","orcid":"","institution":"Autonomous University of Madrid","correspondingAuthor":false,"prefix":"","firstName":"María","middleName":"Elena","lastName":"Cuenca","suffix":""},{"id":545758132,"identity":"db16a8c4-6e01-46ce-b70e-704989901e98","order_by":3,"name":"Sacha Gómez-Moñivas","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA0klEQVRIiWNgGAWjYHACNhiD8QHJWpgNSNbCJkGUevn2w88e/GColdNtb39WzVOzjYGf/wB+LQZn0swNexiOG5udOZB2m+fYbQbJGQkEtDDksEnwMBxL3HYj4dhtHrbbDAY3CDms/w2b5B+QlvsP24p5/t1msD9PwGEMN3LYpHkYaoC2MLMx87YBbWEg5LAbz8ykZQwOAP2Sxiw5t+82j8QNAlrk+5OfSb6pqJMzO3784Yc3327L8fcTchjErsNwJg8x6kGgjliFo2AUjIJRMBIBACrhQeQ2pzwSAAAAAElFTkSuQmCC","orcid":"","institution":"Autonomous University of Madrid","correspondingAuthor":true,"prefix":"","firstName":"Sacha","middleName":"","lastName":"Gómez-Moñivas","suffix":""}],"badges":[],"createdAt":"2025-10-07 17:53:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7801827/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7801827/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96437742,"identity":"5d2a50f5-62e6-444b-a2ed-ff1c5a2fc976","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":766592,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/04fe5bc8c9977fc8b71da53f.docx"},{"id":96437741,"identity":"601af6ae-fbca-4201-acc8-9012c9823ebd","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6330,"visible":true,"origin":"","legend":"","description":"","filename":"508fe5dd00b1497d93f44bf767f9be60.json","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/e80f10c9ffa982e4b0843bd2.json"},{"id":96454866,"identity":"ac85d577-83ca-4066-89ec-654bab9a34ed","added_by":"auto","created_at":"2025-11-21 10:03:13","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":287862,"visible":true,"origin":"","legend":"","description":"","filename":"508fe5dd00b1497d93f44bf767f9be601enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/740269a5bd31f387d3570c97.xml"},{"id":96437747,"identity":"50749f5a-2aac-4a60-8b8a-969723584562","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80529,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/f7a05b9d1d663997cfe4a44a.jpeg"},{"id":96437753,"identity":"5f29bfd1-d32e-4538-a724-ef051be07385","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":101784,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/829c98e5cc9f6e66f4d603a1.jpeg"},{"id":96455430,"identity":"ef734767-0efc-4f9f-a384-657adcbb1990","added_by":"auto","created_at":"2025-11-21 10:04:08","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":60881,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/7e9a66cafbbfe5898cbfe9d6.jpeg"},{"id":96437750,"identity":"cf474e76-1ed2-4b62-93d6-0a562e458ab0","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":102511,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/814a29495d769d5d9dcb607b.jpeg"},{"id":96437744,"identity":"2fdac2ea-4216-419a-949f-36cca54c4705","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":55676,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/14a0eb5314cd6a6064d75fe6.jpeg"},{"id":96437752,"identity":"a84eca34-ff30-4c7d-ab06-d848d1a62a61","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":282875,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/fa78a985e43422aae9d19d6e.jpeg"},{"id":96437754,"identity":"8420c2ee-9d96-46f8-b4c5-25c41487b57d","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":50094,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/35601fbc310c39107b3b5b9a.png"},{"id":96455039,"identity":"0353d3f5-29b9-4b19-a98e-0a299d7ad05a","added_by":"auto","created_at":"2025-11-21 10:03:27","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":61657,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/37d363845be76cff761594fa.png"},{"id":96437756,"identity":"cb3cb9db-6bd3-4143-9fe3-0a99f67efd17","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":25164,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/902dda2f1b9a8317202339b3.png"},{"id":96455459,"identity":"3b2bb357-8ed1-4aae-b890-36396eb45173","added_by":"auto","created_at":"2025-11-21 10:04:09","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":57449,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/c87ef2cf6c72c2ecedcb47cc.png"},{"id":96455426,"identity":"1747bd5c-d897-4615-89a2-f16a285599ca","added_by":"auto","created_at":"2025-11-21 10:04:08","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":32636,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/b7592a954120f9166ded2491.png"},{"id":96437755,"identity":"d934f34e-04a7-4eb8-aa7b-0a4c0047a079","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":297479,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/c9446447f259ac3f9f5c4c88.png"},{"id":96437759,"identity":"533a71e2-e293-4eb8-b241-51100dbd3f86","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"xml","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":287756,"visible":true,"origin":"","legend":"","description":"","filename":"508fe5dd00b1497d93f44bf767f9be601structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/b0d42e0c66f663b362904740.xml"},{"id":96437758,"identity":"55865aa0-12c6-4f97-bfc9-4368519b5a35","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"html","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":308820,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/d3809eeb9dff488c6d8d34e1.html"},{"id":96454814,"identity":"aae7b15a-a655-4b12-af44-199d5b2e9018","added_by":"auto","created_at":"2025-11-21 10:03:09","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":80529,"visible":true,"origin":"","legend":"\u003cp\u003eSchematic representation of the classroom activity and data obtained.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/084cb9b24cf39bc08f394202.jpeg"},{"id":96437738,"identity":"13f7cf66-d5be-4f45-9222-6f49236f0c3e","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":101784,"visible":true,"origin":"","legend":"\u003cp\u003eVariables collected in data.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/5266bdb1af3380a40df564ef.jpeg"},{"id":96437740,"identity":"ecbe5049-b766-4801-b63e-207627c37350","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":60881,"visible":true,"origin":"","legend":"\u003cp\u003eMethodology for machine learning inputs and text translation for TTR and FKG analysis.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/484c704795c157db0226e6cf.jpeg"},{"id":96437743,"identity":"009fd81d-094c-4bdf-8abf-0a3b6f9fdf4d","added_by":"auto","created_at":"2025-11-21 06:09:25","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":102511,"visible":true,"origin":"","legend":"\u003cp\u003eExploratory data analysis: correlation matrix of the variables of the program year (course) before the prediction not computing null values.\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/00d2261c3f58a86fcc5912ff.jpeg"},{"id":96454318,"identity":"46da4367-6271-407e-bac0-fed2abd4523d","added_by":"auto","created_at":"2025-11-21 10:02:36","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":55676,"visible":true,"origin":"","legend":"\u003cp\u003eLollipop chart of the features’ importance of the Random Forest Regressor algorithm.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/3753d6fa9f3c0b8b9ca476a3.jpeg"},{"id":103816243,"identity":"f0316c85-9fd5-48c3-8a78-44d3b445e7a5","added_by":"auto","created_at":"2026-03-03 09:13:18","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1418023,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7801827/v1/86333af1-3f94-4078-9e19-05d94f294aaa.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predictive AI for Academic Performance: Integrating Natural Language Processing, Gamification and Self-regulated Learning","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eArtificial intelligence (AI) and Natural Language Processing (NLP) are increasingly used for diagnosis, prognosis, prediction, and personalization in higher education, as emphasized by initiatives such as the UNESCO Beijing Consensus (UNESCO, \u003cspan citationid=\"CR131\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) and the OECD's \u003cem\u003eDigital Education Outlook 2023\u003c/em\u003e (OECD, \u003cspan citationid=\"CR86\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The integration of AI into education has evolved from computer technologies (Chen et al., \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Devedžic, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2004\u003c/span\u003e) to intelligent systems (Chassignol et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Kahraman et al., \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2010\u003c/span\u003e; Peredo et al., \u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e2011\u003c/span\u003e; Rus et al., \u003cspan citationid=\"CR111\" class=\"CitationRef\"\u003e2013\u003c/span\u003e) and tools like robots and chatbots (Timms, \u003cspan citationid=\"CR126\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), now propelled by the advances in Generative AI (Pokrivcakova, \u003cspan citationid=\"CR92\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Salas-Pilco \u0026amp; Yang, \u003cspan citationid=\"CR113\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Many scholars highlight the urgency of implementing Generative AI to enhance accessibility, efficacy, and inclusivity in education (Lim et al., \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The European Union\u0026rsquo;s AI Act mandates ethical standards for AI applications in educational contexts (European Parliament, n.d.). The advantages of Generative AI in personalized learning (Barrett \u0026amp; Pack, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Sajja et al., \u003cspan citationid=\"CR112\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Samuel Girard et al., 2024), prediction, and analytics \u0026ndash; such as forecasting student scores (Ashenafi et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2015\u003c/span\u003e), success rates (Ludwig et al., \u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), or assessing personality and maturity (Deo et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Rovira et al., \u003cspan citationid=\"CR110\" class=\"CitationRef\"\u003e2017\u003c/span\u003e) \u0026ndash; are well-established in academic research. These benefits underpin our analysis and form the foundation of our predictive model.\u003c/p\u003e"},{"header":"2. Theoretical background","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1. Constructivism, AI-enhanced active learning and self-regulated learning\u003c/h2\u003e\u003cp\u003eConstructivist theory, particularly as articulated by Vygotsky and Bruner, posits that learning is most effective when instruction is scaffolded to students\u0026rsquo; current developmental stage \u0026ndash; within their \u0026ldquo;zone of proximal development\u0026rdquo; (Vygotsky, \u003cspan citationid=\"CR133\" class=\"CitationRef\"\u003e1978\u003c/span\u003e). Accurately predicting a student\u0026rsquo;s academic year enables educators to estimate their cognitive and linguistic maturity and design instruction that bridges prior knowledge with new content (Bruner, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e1960\u003c/span\u003e; Vygotsky, \u003cspan citationid=\"CR133\" class=\"CitationRef\"\u003e1978\u003c/span\u003e). Empirical studies confirm that academic progression is associated with qualitative shifts in writing and discourse sophistication (Crossley \u0026amp; Kim, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Predicting academic year is therefore not a bureaucratic exercise, but a means to enact responsive, developmentally appropriate pedagogy (Bruner, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e1960\u003c/span\u003e; Vygotsky, \u003cspan citationid=\"CR133\" class=\"CitationRef\"\u003e1978\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eAI-based tools can facilitate this by providing interactive, learner-centered environments that promote exploration and immediate engagement with content. AI can offer tailored, interactive learning experiences that allow students to actively construct knowledge, improving learning efficiency (Holmes et al., \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Instead of passively receiving information, students using AI-driven systems (such as intelligent tutors or simulations) learn by doing \u0026ndash; aligning with Piaget\u0026rsquo;s constructivist view that experiential interaction is crucial (Shuliang, \u003cspan citationid=\"CR120\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Empirical evidence supports this synergy: AI-driven adaptive learning platforms continuously adjust activities to a student\u0026rsquo;s level, which \u0026ldquo;markedly enhance[s] academic performance\u0026rdquo; by keeping tasks in an optimal challenge zone ((Baker \u0026amp; Yacef, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) as cited in (Al Nabhani et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2025\u003c/span\u003e)).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2. Self-regulated learning and individual profiles\u003c/h2\u003e\u003cp\u003eSelf-regulated learning (SRL) theory emphasizes students\u0026rsquo; ability to set goals, monitor progress, and regulate both cognition and motivation throughout the learning process (Zimmerman, \u003cspan citationid=\"CR137\" class=\"CitationRef\"\u003e2002\u003c/span\u003e). Self-regulated learners take an active role in their education by planning, monitoring, and adjusting their cognitive and metacognitive processes to achieve their learning goals (Nicol \u0026amp; Macfarlane-Dick, \u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e2006\u003c/span\u003e). Research highlights the close alignment between constructivist classrooms and SRL, with teachers scaffolding students\u0026rsquo; ability to monitor, control, and evaluate their own learning (Nicol \u0026amp; Macfarlane‐Dick, 2006)). Key SRL models highlight motivational factors \u0026ndash; such as self-efficacy, intrinsic interest, and goal orientation \u0026ndash; as crucial to academic achievement (Pintrich \u0026amp; De Groot, \u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e1990\u003c/span\u003e). Research consistently finds that differences in self-regulation and motivation explain significant variance in student outcomes (Panadero, \u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Zimmerman, \u003cspan citationid=\"CR137\" class=\"CitationRef\"\u003e2002\u003c/span\u003e). Understanding students\u0026rsquo; motivational profiles can help educators detect learners who may lack strategic approaches or intrinsic engagement. Research suggests that low self-efficacy can hinder the use of effective learning strategies, which may be addressed through scaffolded tasks and explicit feedback (Pintrich \u0026amp; De Groot, \u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e1990\u003c/span\u003e; Zimmerman, \u003cspan citationid=\"CR137\" class=\"CitationRef\"\u003e2002\u003c/span\u003e). These insights support the tailoring of both instruction and feedback, addressing the whole learner rather than just knowledge deficits.\u003c/p\u003e\u003cp\u003eThis is particularly relevant in digital environments, where self-regulation ability predicts academic success (Broadbent \u0026amp; Poon, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). Data-driven models and AI can support SRL by providing timely feedback and insights that help learners manage their learning process (Afzaal et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). A recent systematic review found that technologies like learning analytics and AI \u0026ldquo;support SRL by providing personalized feedback and facilitating autonomous learning\u0026rdquo; (Faza \u0026amp; Lestari, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3. Predicting academic performance with AI for pedagogical action\u003c/h2\u003e\u003cp\u003eFormative assessment \u0026ndash; \u0026ldquo;assessment for learning\u0026rdquo; \u0026ndash; involves gathering evidence about students\u0026rsquo; current understanding to inform immediate feedback and instructional adaptation (Black \u0026amp; Wiliam, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e1998\u003c/span\u003e). Predictive diagnostics, whether through teacher judgment or algorithmic models, serve as a contemporary extension of formative assessment: they provide timely insights into students\u0026rsquo; academic standing, linguistic development, and motivational needs. The pedagogical value of predictive analytics in education lies not in the prediction itself, but in its capacity to prompt formative feedback and inform adaptive teaching practices (Black \u0026amp; Wiliam, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e1998\u003c/span\u003e; Nicol \u0026amp; Macfarlane-Dick, \u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e2006\u003c/span\u003e). Classic mastery learning models (Bloom, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e1968\u003c/span\u003e) and contemporary learning analytics both demonstrate that diagnostic insights, when coupled with targeted interventions, lead to measurable improvements in achievement and reduced failure rates (Ifenthaler \u0026amp; Yau, \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Thus, predictive diagnosis is pedagogically valuable not for its own sake, but for its role in activating meaningful feedback loops that promote student growth.\u003c/p\u003e\u003cp\u003ePredicting final grades is one of the most significant applications of AI in education, as it closely ties to academic performance. AI has long supported performance monitoring and personalized content to preempt learning obstacles (Bloom, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e1971\u003c/span\u003e; Bloom at al., 1981; Glaser, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e1994\u003c/span\u003e; Guskey, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e1985\u003c/span\u003e). Early studies on machine learning for grade prediction (Ashenafi et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2015\u003c/span\u003e) have focused on Massive Open Online Courses, while additional research examined early-detection systems' impact on students and universities (Raffaghelli et al., \u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Subirats et al., \u003cspan citationid=\"CR122\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). AI applications also extend to supporting students with learning challenges, such as dyslexia, through methods like eye-tracking (Rello et al., \u003cspan citationid=\"CR105\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), recommending courses (Samuel Girard et al., 2024), and adaptive learning or tailored content (Kolluru et al., \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; S. Liu et al., \u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). These applications can help students, teachers, or both (Kumar et al., \u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Pankiewicz \u0026amp; Baker, \u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Moreover, AI tools help assess problem-solving behaviors and predict success in skill simulations, using NLP techniques to analyze behavior sequences (Ludwig et al., \u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). As an interdisciplinary field combining AI and linguistics, NLP has proven to be among the most effective methods for personalized, student-centered learning models.\u003c/p\u003e\u003cp\u003eAI-driven predictions are valuable only when translated into concrete educational actions. Predictive analytics can flag at-risk students based on behavioral or performance data, enabling timely interventions such as academic mentoring or adaptive scaffolding (Ifenthaler \u0026amp; Yau, \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Slade \u0026amp; Prinsloo, \u003cspan citationid=\"CR121\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). These early interventions have been shown to improve retention and learning outcomes by addressing difficulties before they escalate (Larrabee S\u0026oslash;nderlund et al., \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Moreover, predictions of motivational or cognitive profiles allow instructors to tailor learning tasks, pacing, and feedback strategies \u0026ndash; key components of differentiated instruction (Tomlinson, \u003cspan citationid=\"CR128\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Instructional design thus becomes more responsive to students\u0026rsquo; self-regulatory capacities, engagement levels, and academic trajectories. At the curriculum level, aggregate learning data can reveal patterns of difficulty across cohorts, guiding revisions that better align with learners\u0026rsquo; developmental needs (Boroowa \u0026amp; Herodotou, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). In line with Nicol and Macfarlane-Dick\u0026rsquo;s (\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e2006\u003c/span\u003e) model of formative assessment as a feedback loop that fosters self-regulated learning, predictive AI models can generate timely insights that inform adaptive interventions, thereby enhancing personalization and equity in learning environments.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4. Gamification, personality traits and academic performance\u003c/h2\u003e\u003cp\u003eGamification\u0026rsquo;s role in education enhances personalization and student-centered learning, despite challenges to its long-term efficacy. Studies on gamification often examine user type classification (Bartle, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e1996\u003c/span\u003e; Drachen et al., \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2009\u003c/span\u003e; Hamari \u0026amp; Tuunanen, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2014\u003c/span\u003e), including models like the Brain Hex (Tondello et al., \u003cspan citationid=\"CR129\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), which align personality traits with gamified tasks. The Hexad Model has proven effective for personalizing educational activities in the teaching-learning process (Lopez \u0026amp; Tucker, \u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Building on this, our study integrates player types as personality traits into a predictive model, based on the established connection between personality and academic performance. By leveraging survey data from gamified tasks, we integrate these personality indicators to enrich our model.\u003c/p\u003e\u003cp\u003ePersonality traits\u0026rsquo; link to academic performance is well-established, with the Big Five model or the Eysenckian personality factors proving useful for predicting academic performance across educational contexts (Brandt et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Mammadov, \u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Mornar et al., \u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Naqshbandi et al., \u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; O\u0026rsquo;Connor \u0026amp; Paunonen, \u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e2007\u003c/span\u003e; Poropat, \u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). The role of personality in academic performance is further explored in studies that connect personality traits to course grades (Lounsbury et al., \u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e2003\u003c/span\u003e), or examine the intersection of personality traits, motivation, and academic success, suggesting that these factors are interrelated and can contribute to predicting performance (De Feyter et al., \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2012\u003c/span\u003e; Phillips et al., \u003cspan citationid=\"CR90\" class=\"CitationRef\"\u003e2003\u003c/span\u003e). Furthermore, player types, which are associated with different levels of motivation, engagement, and academic achievement, can influence learner\u0026rsquo;s behavior in gamified environments (Lavou\u0026eacute; et al., \u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), aligning with our study's exploration of player types and their role in student performance.\u003c/p\u003e\u003cp\u003eCurrent trends in education emphasize (a) adopting emerging technologies like gamification, NLP, and generative AI, and (b) predicting academic performance to implement tailored interventions. Following these lines, our article proposes a methodology to predict students\u0026rsquo; university program year or academic course (from 1st to 4th) using Hexad player types and NLP text analysis. If the predicted year is lower than the actual, it may indicate underperformance; if higher, it may suggest exceeding expectations. This assumption is based on the idea that the model draws from skills that students develop progressively through their studies. We further assess whether the predicted outcomes correlate with students\u0026rsquo; maturity development, and the relevance of the different parameters included in such predictions.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e2.5. Linguistic development and academic maturity\u003c/h2\u003e\u003cp\u003ePsycholinguistic research establishes a relationship between students\u0026rsquo; development \u0026ndash; whether socio-cognitive or academic \u0026ndash; and the improvement of writing skills and linguistic maturity, with language development recognized as a lifelong process (De Bot \u0026amp; Schrauf, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2010\u003c/span\u003e; Rosselli et al., \u003cspan citationid=\"CR109\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). This process unfolds through different developmental stages based on age, which show changes in linguistic literacy (Ravid \u0026amp; Tolchinsky, \u003cspan citationid=\"CR102\" class=\"CitationRef\"\u003e2002\u003c/span\u003eo Lugo, 1996) and maturity (Ravid \u0026amp; Tolchinsky, \u003cspan citationid=\"CR102\" class=\"CitationRef\"\u003e2002\u003c/span\u003eo Lugo, 1996), in line with socio-cognitive evolution (Anderson, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e1941\u003c/span\u003e; Bremholm et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Hall-Mills \u0026amp; Apel, \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Wagner et al., \u003cspan citationid=\"CR134\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). Although much discourse focuses on childhood and pre-adolescence, some researchers (Berman, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2007\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Grimshaw et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e1998\u003c/span\u003e; Nippold, \u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e2000\u003c/span\u003e) emphasize that the upper grades of high school and college represent a critical period in language-specific skill development, particularly in grammatical and semantic knowledge (Berman \u0026amp; Nir-sagiv, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2007\u003c/span\u003e; Hansson et al., \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Karmiloff-Smith, A., \u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e1992\u003c/span\u003e; Nippold, \u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e2002\u003c/span\u003e; Ravid, \u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e2005\u003c/span\u003e; Ravid et al., \u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e2002\u003c/span\u003e; Rimmer, \u003cspan citationid=\"CR107\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Tolchinsky \u0026amp; Rosado, \u003cspan citationid=\"CR127\" class=\"CitationRef\"\u003e2005\u003c/span\u003e) and genre-dependent text organization, and in global coherence (Berman, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Katzenberger, \u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e2005\u003c/span\u003e). This phase is crucial, as it aligns with the consolidation of general socio-cognitive development and the enhancement of higher-order cognitive capacities (Proverbio \u0026amp; Zani, \u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e2005\u003c/span\u003e), such as abstraction, perspective-taking, divergent and critical thinking, and executive control (Berman, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Kluwe, \u0026amp; Logan, \u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e2000\u003c/span\u003e; Reilly et al., \u003cspan citationid=\"CR104\" class=\"CitationRef\"\u003e2005\u003c/span\u003e; Sato, \u003cspan citationid=\"CR116\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), alongside growing world knowledge and exposure to information during maturation and learning.\u003c/p\u003e\u003cp\u003eThe use of linguistic features such as nominal density, lexical diversity, word and sentence length, and syntactic complexity for assessing linguistic development or maturity is well-established in academic research (Crossley, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Crossley \u0026amp; Kim, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Durrant \u0026amp; Brenchley, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; MacArthur et al., \u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Malvern et al., \u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e2004\u003c/span\u003e; Ravid, \u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e2006\u003c/span\u003e; Richards \u0026amp; Malvern, \u003cspan citationid=\"CR106\" class=\"CitationRef\"\u003e1997\u003c/span\u003e; Shakir \u0026amp; Obeidat, \u003cspan citationid=\"CR118\" class=\"CitationRef\"\u003e1991\u003c/span\u003e; Sun \u0026amp; Xiong, \u003cspan citationid=\"CR124\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Our study builds on previous work that analyzes the academic writing of adolescents and young adults in expository texts (Berman \u0026amp; Nir, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2009\u003c/span\u003e; Berman \u0026amp; Nir-sagiv, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2007\u003c/span\u003e; Ragnarsd\u0026oacute;ttir et al., \u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e2002\u003c/span\u003e; Ravid, \u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e2005\u003c/span\u003e), \u0026ldquo;which focus on issues and ideas, and express the unfolding of claims and argumentation in causal and other logical contexts\u0026rdquo; ((Ravid, \u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e2006\u003c/span\u003e: 794). These analyses have provided evidence of the link between enhanced writing skills and the cognitive and socio-cultural development of students, demonstrating that \u0026ldquo;lexicon and syntax interact with age-related changes in overall attitudes expressed in discussing a socially relevant theme\u0026rdquo; (Berman, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2017\u003c/span\u003e: 5), and the interconnectedness between language development and academic performance (Nippold, \u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e2016\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eBuilding on these studies, we analyzed the writing features of university students\u0026rsquo; responses to a questionnaire, focusing on justified-opinion answers. We employed metrics of lexical diversity and readability \u0026ndash; excluding textual global coherence analysis due to the brevity of the expository texts \u0026ndash; to assess students\u0026rsquo; linguistic maturity and development within higher education environments. Since our data is in Spanish, and most NLP algorithms operate in English, we analyze the impact of translation methods to ensure result reliability. Analyzing open-ended responses introduces complexity but mitigates response bias, as students remain unaware of the specific data points under analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e2.6. Research questions\u003c/h2\u003e\u003cp\u003eTo predict students\u0026rsquo; academic performance, this study introduces an innovative methodological framework that combines Hexad player types, NLP techniques, and translation methods to adapt Spanish-language data for analysis with English-based NLP algorithms. Furthermore, it incorporates the examination of students justified-opinion responses to open-ended questionnaire items. This approach aims to provide insights into learners\u0026rsquo; development by addressing the following research questions:\u003c/p\u003e\u003cp\u003eRQ1: To what extent can the degree program year be predicted from students\u0026rsquo; multiple-choice and open-ended questionnaire responses, along with their gamification type?\u003c/p\u003e\u003cp\u003eRQ2: Which variables \u0026ndash; including measures of linguistic features, Hexad player types, and responses to opinion survey items and gaming behaviors questionnaire \u0026ndash; most strongly influence the prediction of students\u0026rsquo; program year?\u003c/p\u003e\u003cp\u003eRQ3: What insights into students\u0026rsquo; maturity or personality development can be gained through text analysis parameters and player types?\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Materials and methods","content":"\u003cp\u003eThis section is divided into two parts. The first one (3.1) describes the experimental context, including datasets and its exploratory analysis. The second one (3.2) describes how gamification user types, NLP, machine learning techniques and statistical analysis are used to predict the program year.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e3.1 Experimental context\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e3.1.1 Sample\u003c/p\u003e\n\u003cp\u003eData were collected from the Degree in Tourism at XX University. A total of 189 students from the program\u0026rsquo;s four years participated in the study. The demographic characteristics of the participants are provided in Table 1:\u003c/p\u003e\n\u003cp\u003eTable 1. Demographic characteristics of the participants (N=189)\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"510\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 94px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAge\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e18-19\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e34.91%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e20-21\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e47.93%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e22-23\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e13.61%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e24\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e3.55%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 94px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eProgram year\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1st\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e25.39%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e2nd\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e28.57%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e3rd\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e31.75%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e4th\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e14.29%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 94px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNationality\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSpanish\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e76.92%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEU not Spanish\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e5.29%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eChinese\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e9.47%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLatAm\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e7.69%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 94px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eGender\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 85px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMale\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e36.84%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFemale\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e61.99%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOther / Prefer not to answer\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e1.17%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e3.1.2 Procedure\u003c/p\u003e\n\u003cp\u003eData acquisition for all AI inputs was done in a single session for each class to gather the data in the most controlled environment. After scheduling each program-year group class session, students were informed one week in advance that there would be a special activity to be held in the computers lab related to heritage, gamification and topics from their degree subjects. They were also informed that the data would be used for a pilot study and that all personal data would be excluded to prevent identification.\u003c/p\u003e\n\u003cp\u003eThis activity is part of the Tourism Degree soft-skills acquisition tasks which contribute to fulfilling curricular competencies. These competencies are embedded within specific course syllabi and are taught through classroom activities designed to foster critical and creative thinking. By generating contexts where students express their opinions and make suggestions, these activities provide a framework for collecting open responses data. Since students are used to this format when giving their critical opinion, the controlled classroom environment improves data reliability and avoids deviation. This activity gets students answering questions linked to skills they are actively developing throughout their degree program. As these skills are an essential component of their curriculum, the activity aligns with their training and expected proficiency. \u0026nbsp;This provides a scenario for the purpose of this article, as it allows for a structured comparison of differences across program years.\u003c/p\u003e\n\u003cp\u003eThe administration of the activity in the classroom was as follows: Attendance was controlled upon entering the lab so that the responses collection could be tracked. Students were firstly instructed on what an excavation site is, the site context and staff, and the app keyboard controls. They were then allowed to open an application with a virtual reconstruction of an excavation in Egypt and play for approximately 40 minutes. In this virtual reconstruction they can visit different areas and interact with objects that test their understanding. After that, they were told to answer the questionnaire giving feedback about the application, and later to fill in the survey about their gaming habits. Fig. 1 shows the stages of the activity, and the data obtained in each stage.\u003c/p\u003e\n\u003cp\u003eAs shown in the figure, all the data obtained in this activity were collected in three stages:\u003c/p\u003e\n\u003cp\u003e(a) Attendance control: Demographic variables \u0026amp; program year.\u003c/p\u003e\n\u003cp\u003e(b) A questionnaire collecting feedback about the activity: nine 1-5 points questions rating the app contents and technical features plus four open questions where they can express their opinion freely.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e(c) A survey about gaming habits, attitudes or personal preferences: 27 questions to be rated with 1-5 points and two open questions.\u003c/p\u003e\n\u003cp\u003eThe variables collected are summarized in Fig. 2 diagram (see detailed information in Appendix A).\u003c/p\u003e\n\u003cp\u003e3.1.3. Ethical issues\u003c/p\u003e\n\u003cp\u003eAll data collected for this study were part of an activity that was integrated into the course curriculum and mandatory for all students. Upon consultation, the Ethics Committee at Universidad XX granted approval for the study in this context. Being in a classroom setting, a strict protocol has been established for all faculty members using data collected from teaching activities to ensure that students\u0026rsquo; personal data is not published in a way that could lead to identification. Informed consent was obtained from all participants through the following statement on the forms in which they provided the data used in this study:\u003c/p\u003e\n\u003cp\u003e\u0026ldquo;During this course we are using some game-like elements. The game elements are part of a pilot activity, aiming to explore what kinds of gamification elements work in different course contexts and for different users. For this purpose, we kindly ask you to answer this brief survey. Personal data will only be used to link participants\u0026rsquo; survey responses to respective course data. The data will be process among GDPR, and instructions of XX University (\u0026hellip;).\u003c/p\u003e\n\u003cp\u003e1. Consent to participate. I understand the nature of the study, and\u003c/p\u003e\n\u003cp\u003e___ agree to participate in the study.\u003c/p\u003e\n\u003cp\u003e___ agree to the use of the data collected for course purposes only\u0026rdquo;.\u003c/p\u003e\n\u003cp\u003eAlthough the activity was mandatory as part of the course requirements, students had the option to opt out of having their data used in the study. In such cases, the teacher would still use the data for course-related purposes, but the data would be excluded from the study. Not a single student opted out of having their data included in the study.\u003c/p\u003e\n\u003cp\u003eThe consent statement was prominently displayed at the top of the forms, ensuring that students did not submit or complete any open-ended responses, writing activities, or other types of data shown in Fig. 2 without first being informed.\u003c/p\u003e\n\u003cp\u003eThe generative AI tool ChatGPT 4 and the Bing Copilot of XX University have been used to perform the Generative AI section.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e3.2 Machine learning prediction\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e3.2.1 Text data treatment\u003c/p\u003e\n\u003cp\u003eThe process for obtaining the machine learning\u0026rsquo;s input features is illustrated in Fig. 3. As shown, readability and lexical diversity metrics require pre-processing, as it is necessary to translate the open-ended responses from Spanish into English to utilize open-source Python libraries, which are predominantly English-specific. To minimize translation-related artifacts, we applied four translation approaches, each rigorously analyzed to ensure accuracy:\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003e\u003cstrong\u003e\u003cem\u003eOpen-Source Python Libraries\u003c/em\u003e\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eUsing Python open libraries for both the translation and obtaining Type-Token Ratio (TTR) and Flesch-Kincaid Grade (FKG).\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003e\u003cem\u003ePython Libraries + Linguist review\u003c/em\u003e\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eUsing Python open libraries for translation, reviewed later by a linguist, and Python open libraries for obtaining TTR and FKG.\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003e\u003cem\u003eChatGPT 4 + Python Libraries\u003c/em\u003e\u003c/strong\u003e\u003cem\u003e:\u0026nbsp;\u003c/em\u003eUses Generative AI (ChatGPT 4) for the translation and Python open libraries for obtaining TTR and FKG.\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003e\u003cem\u003eChatGPT 4 + Bing Copilot\u003c/em\u003e\u003c/strong\u003e: Uses Generative AI both for the translation (ChatGPT 4) and obtaining TTR and FKG (Bing Copilot).\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003e3.2.2 Natural Language Processing\u003c/p\u003e\n\u003cp\u003eAmong the several open-source Python libraries for Spanish-English translation, we used the free software library Translators 5.9.0 https://pypi.org/project/translators. In addition, two other translation approaches were implemented: translated text with a linguist revision, and generative AI translation using ChatGPT 4. For this last translation, careful prompting was needed to ensure that students\u0026rsquo; writing style was not changed due to an inaccurate translation diluting their personal features, which could affect maturity or program-year predictions.\u003c/p\u003e\n\u003cp\u003eAs the translated text is used as input to assess lexical diversity and readability, it is relevant to preserve the original style regarding semantics, syntax and any feature that may portray students\u0026rsquo; voice (note that readability is one of the criteria used in translation to assess quality (Nababan \u0026amp; Nuraeni, 2012; Wijaksono et al., 2022). Rendering stylistic characteristics and author\u0026rsquo;s personality was a problem in machine translation (Mirkin et al., 2015; Rabinovich et al., 2016), and the implementation of AI with Neural Machine Translation systems has not solved the issue. ChatGPT may produce good quality translations if suitable prompts optimize the process avoiding mistakes \u0026ndash; poor understanding of context, biases, and flaws in detecting cultural nuances (Bender et al., 2021; Brown et al., 2020; Castilho et al., 2017; Koehn \u0026amp; Knowles, 2017). But Large Language Models \u0026ndash; trained on vast datasets to create a single model \u0026ndash; dilute individual stylistic elements and can hide the author\u0026apos;s demographic, psychometric or personal features in unfaithful translations (Britz et al., 2017; Busch, 2024; Ji et al., 2023; Vanmassenhove, 2024). Therefore, well-crafted prompting emphasizing literal or faithful translation techniques (Newmark, P., 1988) that prevent skipping author\u0026rsquo;s stylistic characteristics (Hu et al., 2017; Rabinovich et al., 2016; Sennrich et al., 2016; Shen et al., 2017) and grant style transfer (Prabhumoye et al., 2018) was needed for our study. Thus, besides providing a description of the text\u0026rsquo;s context, specific commands were repeated urging ChatGPT to preserve the student\u0026rsquo;s degree of formality and range of vocabulary, replicate mistakes and colloquialisms, and be aware of the importance of accuracy and fidelity (see some details in Appendix B).\u003c/p\u003e\n\u003cp\u003eRegarding lexical diversity and readability, the 2021-released Python open library \u0026nbsp;https://github.com/WSE-research/LinguaF/tree/main (2024 update) was used, particularly its TTR and FKG measures.\u003c/p\u003e\n\u003cp\u003eTTR is a measure of lexical diversity in a text. It is commonly used in linguistics and computational linguistics to quantify the variety of different words (types) relative to the total number of words (tokens) in a given text. The higher the TTR, the greater the lexical diversity in the text. A lower TTR suggests more repetition or fewer unique words relative to the total number of words. The values for TTR range typically from 0 to 1, but it can also be expressed as a percentage. \u0026nbsp;The interpretation of TTR may vary depending on the context and the type of text being analyzed. Different genres, writing styles, or languages may naturally exhibit different levels of lexical diversity. TTR is a useful tool for comparing the diversity of vocabulary across different texts or for tracking changes in vocabulary richness over time.\u003c/p\u003e\n\u003cp\u003eThe FKG is a readability test designed to estimate the readability level of English texts. It is commonly used to assess the complexity of written material and is often applied to evaluate the difficulty of reading comprehension. The FKG is based on two factors: the average number of words per sentence and the average number of syllables per word. The formula for calculating the FKG is: \u003cem\u003eFKG=0.39 (Total words / Total sentences) + 11.8 (Total Syllables / Total words) - 15.59\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eOne of the interpretations of the resulting FKG is a correspondence with a U.S. grade level, indicating the level of education generally required to understand the text. For example, an FKG of 8.0 would suggest that the text is readable by an eighth grader. The range of FKG values typically corresponds to the 12 U.S. grade levels. Lower FKG values indicate simpler text, while higher values suggest more complex and advanced writing. The FKG is just one of many readability metrics, and its interpretation can be influenced by factors such as sentence structure and word choice. Negative values are not common and would typically occur in situations where the text is extremely short, and the formula results in a grade level under 1.\u003c/p\u003e\n\u003cp\u003e3.2.3 Gamification Hexad model of player types\u003c/p\u003e\n\u003cp\u003eTo integrate personality traits into our computational analysis, we applied the Hexad model (Marczewski, 2015), a validated framework in educational gamification. This model categorizes users based on distinct motivations, aligning well with our study\u0026rsquo;s focus on personalizing prediction. Data was collected via a survey completed after students engaged in a gamified virtual excavation task, which enabled us to capture diverse motivational factors in an interactive learning environment. The Hexad model, tested in educational contexts (Subirats, Nousiainen, et al., 2023), provides six non-excluding user types (Tondello et al., 2016) that reflect motivations key to predicting academic progression:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003ePhilanthropists\u0026mdash;motivated by purpose;\u003c/li\u003e\n \u003cli\u003eSocializers\u0026mdash;motivated by relatability;\u003c/li\u003e\n \u003cli\u003eFree spirits\u0026mdash;motivated by autonomy;\u003c/li\u003e\n \u003cli\u003eAchievers\u0026mdash;motivated by competence;\u003c/li\u003e\n \u003cli\u003ePlayers\u0026mdash;motivated by rewards;\u003c/li\u003e\n \u003cli\u003eDisruptors\u0026mdash;motivated by change.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThe relationship between Hexad player types and students\u0026rsquo; development is central to this study, as predicting the program year depends on features that may evolve as students\u0026rsquo; progress through their studies. In the Hexad player profile, motivation levels may vary, with students often focused on broader goals in early years and increasingly career-oriented in later years as they prepare for post-graduation opportunities.\u003c/p\u003e\n\u003cp\u003e3.2.4 Machine learning.\u003c/p\u003e\n\u003cp\u003eTo ensure that the translation into English is not an artifact in the predictions we followed these four steps for applying AI algorithms in the prediction of the program year:\u003c/p\u003e\n\u003col start=\"1\" type=\"1\"\u003e\n \u003cli\u003eData is obtained in class (the target is the program year, with values between 1 and 4).\u003c/li\u003e\n \u003cli\u003eWe apply one of the translations got from each approximation.\u003c/li\u003e\n \u003cli\u003eFrom each translation, TTR and FKG are obtained.\u003c/li\u003e\n \u003cli\u003eWith each TTR, FKG and other variables, we apply Random Forest for predicting the program year.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003ePandas library was used for processing data and implementing the Hexad framework of six user types. Three different approaches were used to impute null values: Median, k-Nearest Neighbors algorithm and Iterative imputer. After conducting comparative program-year prediction trials among the three, the iterative imputer was chosen as it showed better results for predicting the program year.\u003c/p\u003e\n\u003cp\u003eThe supervised learning algorithm chosen was the Random Forest regressor because it has good and robust performance when tuning the parameters accordingly, besides providing information about the most relevant parameters for the prediction. The Python open library used was https://scikit-learn.org and the Random Forest regressor was trained with 10-fold cross-validation and 80% of train data and 20% of test data. The Mean Absolute Error (MAE) and Mean Squared Error (MSE) (Karunasingha, 2022) were used as evaluating metrics to analyze the performance of the algorithm.\u003c/p\u003e\n\u003cp\u003eAfter applying machine learning to do statistical tests on the prediction results, normality tests were conducted showing that they do not follow a normal pattern; so, the Wilcoxon signed-rank test was used (Moore et al., 2009). The Wilcoxon signed-rank test is a statistical test used to compare the means (averages) to determine if they are significantly different from each other. It helps to figure out whether the differences you observe in certain data are real or if they could have happened by chance. It was applied to compare the statistical significance of the best approximation to all the other approximation errors. Therefore, three statistical Wilcoxon tests were computed.\u003c/p\u003e"},{"header":"4.\tResults","content":"\u003cp\u003eThis section is divided into two subsections: (4.1) an exploratory data analysis performed before applying machine learning, and (4.2) the results of performing supervised learning to predict the program year of students from analyzing all the features described in Fig. 2.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e4.1 Exploratory data analysis\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eA correlation heatmap of features is shown in Fig. 4, where blank values in the correlation heatmap of \u0026ldquo;Other gender / NA\u0026rdquo; (not answered) vs. some variables mean that there is not enough data to compute these correlations due to null values (detailed correlation with numerical values provided in Appendix D). This figure shows that there are some high correlations between open-questions, and the inverse correlation of \u0026ldquo;Suggestions\u0026rdquo; readability and lexical diversity is noteworthy. High and positive correlations between students\u0026rsquo; opinions about \u0026ldquo;Evaluation of activities\u0026rdquo; and \u0026ldquo;Technical elements of the application\u0026rdquo; are also observed. Besides, interesting contrasting correlations are seen between genders and the \u0026ldquo;Frequency of playing video games\u0026rdquo;: -0.3 between \u0026ldquo;Females\u0026rdquo; and \u0026ldquo;Frequency of playing video games\u0026rdquo;, and 0.3 between \u0026ldquo;Males\u0026rdquo; and \u0026ldquo;Frequency of playing video games\u0026rdquo;. Additionally, the correlation between gamification user types is positive too, not contradicting previous studies (Tondello et al., 2016).\u003c/p\u003e\n\u003cp\u003eDetailed figures from the exploratory data analysis are provided in Appendix D, including Table A.1, which presents the mean values and the evolution of Hexad player profiles, and Fig. A.1, which offers a comprehensive view of the correlation matrix.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e4.2 Prediction with supervised learning \u0026ndash; regression\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eTable 2 shows the results of the program year\u0026rsquo;s prediction after applying supervised learning with Random Forest regression type. To interpret its contents, the differences between MSE/MAE and Train error / Test error must be explained.\u003c/p\u003e\n\u003cp\u003eThe mean MSE/MAE of the training set are the errors between the predicted values and the actual values for the data the model was trained on. It provides an indication of how well the model fits the training data. A low MSE/MAE on the training set generally indicates that the model is fitting the training data well. However, if it is too low compared to the test MSE/MAE, it might indicate overfitting.\u003c/p\u003e\n\u003cp\u003eThe mean MSE/MAE of the test set is the average absolute error between the predicted values and the actual values for the data that was not provided to the model during training. It provides an indication of the model\u0026apos;s generalization performance with unknown data. A low MSE/MAE on the test set suggests that the model generalizes well to new data, while a high MSE/MAE indicates poor generalization, which might be due to overfitting or underfitting.\u003c/p\u003e\n\u003cp\u003eRegarding which metrics are best suited, MAE and MSE have different properties. MAE is linear, meaning all errors are weighted equally. It is more robust to outliers compared to MSE because it does not square the errors. MAE provides a clear interpretation: it gives the average error in the same units as the original data. MSE squares the errors, which means it penalizes larger errors more than smaller ones. It is sensitive to outliers due to the squaring of differences. MSE\u0026apos;s unit is the square of the original data units, which can be less intuitive to interpret directly. Due to what has just been mentioned, MAE can be more suitable for this problem because of its better interpretability and robustness to outliers.\u003c/p\u003e\n\u003cp\u003eConsidering the information above, in Table 2 we can see that the minimum values of the test\u0026rsquo;s errors are obtained when the translation is reviewed by a linguist, closely followed by the errors obtained by ChatGPT 4 translation.\u003c/p\u003e\n\u003cp\u003eTable 2. Errors of the program-year prediction from students\u0026rsquo; open data using Random Forest.\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eApproximation 1: Python open libraries\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eApproximation 2: Python open libraries + linguist-reviewed translation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eApproximation 3: ChatGPT4 translation + Python open libraries computing TTR \u0026amp; FKG\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eApproximation 4: ChatGPT4 translation + Bing Copilot computing TTR \u0026amp; FKG\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eMean MSE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.842\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.840\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.820\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.754\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eSD MSE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.162\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.169\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.149\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.152\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eTest MSE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e1.004\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.922\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e1.016\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.984\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 125px;\"\u003e\n \u003cp\u003eTest MAE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.856\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.794\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.856\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e0.809\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eLowest test-set error values in bold.\u003c/p\u003e\n\u003cp\u003eBy applying the Shapiro-Wilk test, we have concluded that the data can significantly deviate from a normal distribution. Then we apply the Wilcoxon test from the list of errors obtained from the 4 approximations to compare the best one (second) to the others. We have obtained that, in all cases, we cannot find statistically significant differences. The difference that is closer to being significant is between approximations 1 and 2, with statistic=236.5 and p-value=0.08.\u003c/p\u003e\n\u003cp\u003eThanks to the fact that the Random Forest is explainable and can provide information about the relevance of each feature, the most important attributes of the prediction could be identified. Fig. 5 shows a lollipop chart of the features\u0026rsquo; importance.\u003c/p\u003e"},{"header":"5.\tDiscussion and implications","content":"\u003cp\u003eThis study is anchored in constructivist and self-regulated learning theories. Predicting students\u0026rsquo; program year through AI-supported models enables more tailored educational scaffolding, aligning instruction with students\u0026rsquo; developmental stages (Bruner, 1960; Crossley \u0026amp; Kim, 2022; Vygotsky, 1978). Methodologically, combining open- and closed-question survey data with advanced machine learning represents a step forward for educational analytics. Integrating gamification user types responds to calls for personalized learning environments where motivation and individual profiles influence academic success (Broadbent \u0026amp; Poon, 2015; Panadero, 2017; Pintrich \u0026amp; De Groot, 1990). \u003c/p\u003e\n\u003cp\u003e\u003cem\u003e5.1. Research questions discussion\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eOur application of lexical diversity and readability analysis via NLP, together with gamification user types, aimed to predict students\u0026rsquo; academic maturity (i.e., program year) based on questionnaire responses. This approach addressed three research questions, discussed below based on our findings.\u003c/p\u003e\n\u003cp\u003eRQ1: To what extent can the degree program year be predicted from students\u0026rsquo; multiple-choice and open-ended questionnaire responses, along with gamification types? \u003c/p\u003e\n\u003cp\u003eAs shown in Table 2, predictions of program year were reasonably accurate, with errors 1 in all cases \u0026ndash; except when using unedited free translation tools. The lowest prediction error, under 0.8, occurred when human post-editing was applied. Although differences among translation methods were not statistically significant, human revision consistently improved accuracy. This highlights the continued importance of human oversight in AI-assisted translations, especially when high precision is needed. However, when handling large datasets, human intervention may be impractical due to time and resource demands. In such cases, automatic translation tools remain acceptable for large-scale analysis (Castilho et al., 2017).\u003c/p\u003e\n\u003cp\u003eWhere possible, responses in student\u0026rsquo;s native languages should be prioritized, especially when they are not native speakers of English or Russian \u0026ndash; the languages supported by the LinguaF library. Using a second language may distort readability and lexical diversity metrics, as non-native writing may not fully reflect students\u0026rsquo; academic development, personality traits, or natural writing style. Therefore, collecting data in students\u0026rsquo; first language preserves the authenticity of linguistic features used in our analysis.\u003c/p\u003e\n\u003cp\u003eRQ2: Which variables \u0026ndash; including measures of linguistic features, Hexad player types, and responses to opinion survey items and gaming behaviors questionnaire \u0026ndash; most strongly influence the prediction of students\u0026rsquo; program year?\u003c/p\u003e\n\u003cp\u003eAs Fig. 5 illustrates, lexical diversity and readability metrics were the most predictive features, outperforming demographic and attitudinal variables. This aligns with prior research showing that writing sophistication increases throughout adolescence and early adulthood (Berman \u0026amp; Nir-sagiv, 2007; Crossley, 2020; Nippold, 2016; Ravid, 2005). Only one Likert-scale item, evaluating \u0026ldquo;Technical elements of the application 4\u0026rdquo; ranked among the top ten predictors \u0026ndash; specifically, in tenth place. Among the Hexad player types, \u0026ldquo;Free Spirit,\u0026rdquo; \u0026ldquo;Socializer,\u0026rdquo; and \u0026ldquo;Philanthropist,\u0026rdquo; showed higher predictive relevance, suggesting that motivational orientations evolve across program years (Santos et al., 2023; Tondello et al., 2016). While these features are key to the model\u0026rsquo;s performance, Table A.1 and the correlation matrix revealed no linear correlations between them and program year. This supports the use of non-linear AI models to capture more complex relationships (Deo et al., 2020; Rovira et al., 2017).\u003c/p\u003e\n\u003cp\u003eThe correlation matrix also showed positive associations among parameters from the same source. Likert-scale responses tended to correlate positively with one another, though without revealing distinct trends across program years. This suggests that students who rate one application feature highly often rate others similarly. However, these ratings \u0026ndash; along with demographic data like gender \u0026ndash; had minimal predictive value and appeared at the bottom of the feature importance rankings in Fig. 5.\u003c/p\u003e\n\u003cp\u003eThese findings indicate that prediction of program year is driven primarily by user type and NLP-based features. These are likely to reflect skills acquired during the degree program or personality traits that evolve over time. The comparison between predicted and actual program year can thus offer meaningful insights into students\u0026rsquo; academic maturity and potential risk factors.\u003c/p\u003e\n\u003cp\u003eRQ3: What insights into students\u0026rsquo; maturity or personality development can be gained through text analysis parameters and player types?\u003c/p\u003e\n\u003cp\u003eIdentifying student maturity or personality development requires: 1) a well-defined link between personality traits and the parameters used; and 2) evidence of correlations between these traits and program year. Our analysis confirmed the relevance of lexical diversity and readability in predicting academic progress, These NLP-derived features reflect writing ability development consistent with psycholinguistic theories that describe gradual semantic and syntactic growth through adolescence and early adulthood (Berman, 2007, 2008; Berman \u0026amp; Nir-sagiv, 2007; Grimshaw et al., 1998; Hansson et al., 2016; Karmiloff-Smith, 1992; Nippold, 2000, 2002, 2016; Ravid \u0026amp; Tolchinsky, 2002; Ravid, 2005; Ravid et al., 2002; Rimmer, 2008). No single indicator serves as a definite maturity marker, but the aggregate output from NLP offers a holistic view of development (De Bot \u0026amp; Schrauf, 2010).\u003c/p\u003e\n\u003cp\u003eOpen-ended questions encouraging critical reflection were particularly effective at eliciting complex language and metacognitive awareness (Aminah, 2024; Sato, 2022). One such item, addressing \u0026ldquo;Time consumption readability,\u0026rdquo; proved especially predictive. It required students to justify their choice between virtual reconstructions and traditional teaching methods \u0026ndash; blending expository and argumentative writing with introspection, practical reasoning and creative thinking (Aminah, 2024; Berman \u0026amp; Nir-sagiv, 2007; Dimitrova, 2024; Hidayat et al., 2018; Ragnarsd\u0026oacute;ttir et al., 2002; Rahayuni̇Ngsi̇H et al., 2021; Ravid, 2005, 2006; Sato, 2022; Zhang et al., 2024). Similar prompts have been used to assess personality traits and maturity through problem-solving and critical thinking (Arumningsih et al., 2023; Dai et al., 2022; Y. Liu et al., 2022). The cognitive and linguistic complexity of such questions makes them useful not only for predictive modeling but also for identifying growth in maturity and writing skills. Variability in the linguistic sophistication of students\u0026rsquo; responses to these tasks may reveal differences in developmental stage, making such items valuable for current prediction and for future longitudinal tracking.\u003c/p\u003e\n\u003cp\u003eRegarding Hexad user types, no strong direct correlations were found with program year. However, \u0026quot;Socializer,\u0026quot; \u0026quot;Philanthropist\u0026quot; and \u0026quot;Free Spirit\u0026quot; types emerged as slightly more predictive than others, suggesting motivational shifts across academic years (Brandt et al., 2020). Prior studies indicate that altruistic traits, as seen in \u0026ldquo;Philanthropists,\u0026rdquo; tend to increase with age (Tondello et al., 2019), aligning with our data showing an association between program year and age. This supports the interpretation of motivational evolution as part of student maturation. The \u0026quot;Achiever\u0026quot; type may reflect growing awareness of academic performance and career planning, echoing findings from the \u003cem\u003eJob Outlook 2020\u003c/em\u003e report by the National Association of Colleges and Employers (National Association of Colleges and Employers, 2019). Competitive traits may be present early in studies but often intensify as graduation approaches, with a sharper focus on employability. Similarly, the \u0026quot;Player\u0026quot; profile may evolve, with motivations shifting as students advance in their academic journey (Santos et al., 2023).\u003c/p\u003e\n\u003cp\u003eGender was not a significant predictor of program year, suggesting that the maturity and personality-related parameters captured by the model do not require gender-based differentiation. This indicates similar patterns of development, concerns and motivations across male and female students.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e5.2. Educational Implications\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe findings offer multiple implications for teaching and learning. First, integrating AI-driven predictive diagnostics into formative assessment can support personalized, developmentally appropriate instruction (Black \u0026amp; Wiliam, 1998; Ifenthaler \u0026amp; Yau, 2020; Nicol \u0026amp; Macfarlane‐Dick, 2006). Such tools can help identify students at risk and inform differentiated instruction or peer-based strategies (Tempelaar et al., 2017). Feedback on linguistic and motivational development may also promote self-awareness, self-regulation, and intrinsic motivation (Panadero, 2017; Zimmerman, 2002).\u003c/p\u003e\n\u003cp\u003eHowever, maturity analysis must be applied with care. Awareness of academic underperformance could negatively affect student confidence. Used strategically, though, these insights can guide instructional measures that anticipate students\u0026rsquo; development risks and improve learning outcomes. This model can contribute to students\u0026rsquo; soft skills and academic growth in four key areas:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eSelf-awareness: By understanding their learning styles and gamification profile, students can better recognize and manage their strengths and weaknesses. For instance, a student not yet motivated by long-term academic or career goals might adjust their strategies after reflecting on their \u0026ldquo;Player\u0026rdquo; versus \u0026ldquo;Achiever\u0026rdquo; profile. Teachers can support this by highlighting overlooked challenges and achievements.\u003c/li\u003e\n\u003cli\u003eSelf-discipline: The analysis encourages better time management and structured learning. Sophisticated writing, associated with higher maturity, requires planning, critical thinking and self-monitoring. This evolution also enhances learning efficiency, as tasks completed without critical engagement may be superficial and ineffective in fostering real academic growth. Therefore, upon detecting a writing skill weakness, teachers can design writing tasks to target underdeveloped writing skills and promote deeper engagement and critical thinking.\u003c/li\u003e\n\u003cli\u003eMotivation: Understanding one\u0026rsquo;s own developmental stage may foster intrinsic motivation, reducing reliance on external rewards. As motivation develops longitudinally and personal growth is often imperceptible over short periods, students may not be aware of their progress. Teachers can support this by explicitly recognizing student progress, helping learners appreciate their academic evolution. This can be a useful method to enhance motivation in students who have difficulty in perceiving their own improvements.\u003c/li\u003e\n\u003cli\u003eCollaboration: Awareness of peers\u0026rsquo; diverse maturity levels can improve group work. A comparative analysis of students at different levels of academic maturity can help identify successful learning strategies. By sharing approaches used by peers with higher academic development, students can adopt more effective learning strategies to achieve their goals. In this line, our predictive methodology can help create structured collaborative activities that pair students with differing developmental profiles. This may generate mutually beneficial learning dynamics.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThese insights allow instructors and academic advisors to better align teaching strategies with students\u0026rsquo; evolving traits, supporting more effective and inclusive environments.\u003c/p\u003e"},{"header":"6.\tConclusions","content":"\u003cp\u003eThis study highlights the effectiveness of artificial intelligence in analyzing students\u0026rsquo; open-ended responses to predict academic progression and identifying those at risk of underperformance. Our results demonstrate that the lexical diversity and readability features, along with gamification profiles, are key predictors of students\u0026rsquo; program year, with models reaching acceptable accuracy. As such, the system can help detect students falling behind \u0026ndash; i.e., when the predicted year is lower than the actual one \u0026ndash; enabling timely feedback and intervention.\u003c/p\u003e\n\u003cp\u003eBeyond predictive accuracy, the model offers pedagogical value by estimating students\u0026rsquo; cognitive and motivational maturity. The ability to infer program year from textual and gamification data supports the design of instruction that aligns with students\u0026rsquo; developmental stages and allows for early, personalized support. These findings are consistent with constructivist and self-regulated learning theories, which emphasize adapting instruction to students\u0026rsquo; evolving competencies. Although lexical and readability features proved central to the prediction model, and gamification profiles improved prediction performance, no single parameter showed a clear linear relationship with academic stage. This suggests motivational changes as students progress and supports the model\u0026rsquo;s developmental relevance. Besides, it reflects the complex and multifaceted nature of maturity, which often eludes traditional linear metrics. AI\u0026rsquo;s ability to model these non-linear dynamics strengthens its relevance for educational applications.\u003c/p\u003e\n\u003cp\u003eThe model\u0026rsquo;s strong performance with AI-assisted translations also broadens its applicability across multilingual contexts and allows responses written in students\u0026rsquo; native languages. When combined with AI tools capable of reliably assessing lexical and readability metrics, this approach can be applied to diverse educational settings.\u003c/p\u003e\n\u003cp\u003eThe insights offered by this model extend beyond grade prediction. They support formative feedback loops and developmentally aligned instruction, allowing educators to recognize and assist underperforming students.\u003c/p\u003e\n\u003cp\u003eLooking ahead, several future research directions emerge. Longitudinal studies could track changes in students\u0026rsquo; linguistic and motivational profiles throughout their academic programs, clarifying how these features evolve over time. Cross-disciplinary applications may allow the model to be trained for fields where writing tasks are less central. Translation sensitivity could be further explored, especially the impact of translation accuracy on key metrics like TTR and FKG. Additionally, comparative studies with different AI tools and multilingual datasets could improve robustness and generalizability.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAI \u0026ndash; Artificial Intelligence\u003c/p\u003e\n\u003cp\u003eFKG \u0026ndash; Flesch Kincaid Grade\u003c/p\u003e\n\u003cp\u003eMAE \u0026ndash; Mean Absolute Error\u003c/p\u003e\n\u003cp\u003eMSE \u0026ndash; Mean Squared Eror\u003c/p\u003e\n\u003cp\u003eNLP \u0026ndash; Natural Language Processing\u003c/p\u003e\n\u003cp\u003eSRL \u0026ndash; Self-Regulated Learning\u003c/p\u003e\n\u003cp\u003eTTR \u0026ndash; Type Token Ratio\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eL. S.: Data curation, Investigation, Methodology, Resources, Software, Visualization, Writing \u0026ndash; original draft. B. N.: Data curation, Investigation, Resources, Writing \u0026ndash; original draft \u0026amp; review and editing. M. E. C.: Investigation, Writing \u0026ndash; original draft \u0026amp; review and editing. S. G. M.: Conceptualization, Formal Analysis, Investigation, Methodology, Project administration, Supervision, Writing \u0026ndash; review and editing.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eAfzaal, M., Nouri, J., Zia, A., Papapetrou, P., Fors, U., Wu, Y., Li, X., \u0026amp; Weegar, R. (2021). Explainable AI for Data-Driven Feedback and Intelligent Action Recommendations to Support Students Self-Regulation. \u003cem\u003eFrontiers in Artificial Intelligence\u003c/em\u003e, \u003cem\u003e4\u003c/em\u003e, 723447. https://doi.org/10.3389/frai.2021.723447\u003c/li\u003e\n \u003cli\u003eAl Nabhani, F., Hamzah, M. B., \u0026amp; Abuhassna, H. (2025). The role of artificial intelligence in personalizing educational content: Enhancing the learning experience and developing the teacher\u0026rsquo;s role in an integrated educational environment. \u003cem\u003eContemporary Educational Technology\u003c/em\u003e, \u003cem\u003e17\u003c/em\u003e(2), ep573. https://doi.org/10.30935/cedtech/16089\u003c/li\u003e\n \u003cli\u003eAminah, M. (2024). Project-Based Learning: Optimizing Students\u0026rsquo; Critical Thinking Skills. \u003cem\u003eMedia Bina Ilmiah\u003c/em\u003e, \u003cem\u003e18\u003c/em\u003e(12), 3177\u0026ndash;3184.\u003c/li\u003e\n \u003cli\u003eAnderson, J. E. (1941). Principles of growth and maturity in language. \u003cem\u003eThe Elementary English Review\u003c/em\u003e, \u003cem\u003e18\u003c/em\u003e(7), 250\u0026ndash;277.\u003c/li\u003e\n \u003cli\u003eArumningsih, E., Setyawati, R., \u0026amp; Murtianto, Y.H. (2023). Students\u0026rsquo; Creative Thinking Ability in Solving Open-Ended Problems Based on Personality Type. \u003cem\u003eHipotenusa Journal of Mathematical Society\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e, 121\u0026ndash;131. https://doi.org/10.18326/hipotenusa.v5i2.280\u003c/li\u003e\n \u003cli\u003eAshenafi, M. M., Riccardi, G., \u0026amp; Ronchetti, M. (2015). Predicting students\u0026rsquo; final exam scores from their course activities. \u003cem\u003e2015 IEEE Frontiers in Education Conference (FIE)\u003c/em\u003e, 1\u0026ndash;9. https://doi.org/10.1109/FIE.2015.7344081\u003c/li\u003e\n \u003cli\u003eBaker, R. S. J. d., \u0026amp; Yacef, K. (2009). \u003cem\u003eThe State of Educational Data Mining in 2009: A Review and Future Visions\u003c/em\u003e. https://doi.org/10.5281/ZENODO.3554657\u003c/li\u003e\n \u003cli\u003eBarrett, A., \u0026amp; Pack, A. (2023). Not quite eye to A.I.: Student and teacher perspectives on the use of generative artificial intelligence in the writing process. \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e(1), 59. https://doi.org/10.1186/s41239-023-00427-0\u003c/li\u003e\n \u003cli\u003eBartle, R. (1996). Hearts, clubs, diamonds, spades: Players who suit MUDs. \u003cem\u003eMUD Research\u003c/em\u003e, \u003cem\u003e1\u003c/em\u003e, 19\u0026ndash;20.\u003c/li\u003e\n \u003cli\u003eBender, E. M., Gebru, T., McMillan-Major, A., \u0026amp; Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜. \u003cem\u003eProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency\u003c/em\u003e, 610\u0026ndash;623. https://doi.org/10.1145/3442188.3445922\u003c/li\u003e\n \u003cli\u003eBerman, R. A. (2007). Developing Linguistic Knowledge and Language Use Across Adolescence. In E. Hoff \u0026amp; M. Shatz (Eds.), \u003cem\u003eBlackwell Handbook of Language Development\u003c/em\u003e (pp. 347\u0026ndash;367). Blackwell Publishing Ltd. https://doi.org/10.1002/9780470757833.ch17\u003c/li\u003e\n \u003cli\u003eBerman, R. A. (2008). The psycholinguistics of developing text construction. \u003cem\u003eJournal of Child Language\u003c/em\u003e, \u003cem\u003e35\u003c/em\u003e(4), 735\u0026ndash;771. https://doi.org/10.1017/S0305000908008787\u003c/li\u003e\n \u003cli\u003eBerman, R. A. (2017). Language Development and Literacy. In R. J. R. Levesque (Ed.), \u003cem\u003eEncyclopedia of Adolescence\u003c/em\u003e (pp. 1\u0026ndash;11). Springer International Publishing. https://doi.org/10.1007/978-3-319-32132-5_19-2\u003c/li\u003e\n \u003cli\u003eBerman, R. A., \u0026amp; Nir, B. (2009). Cognitive and linguistic factors in evaluating text quality: Global versus local? In V. Evans \u0026amp; S. Pourcel (Eds.), \u003cem\u003eHuman Cognitive Processing\u003c/em\u003e (Vol. 24, pp. 421\u0026ndash;440). John Benjamins Publishing Company. https://doi.org/10.1075/hcp.24.26ber\u003c/li\u003e\n \u003cli\u003eBerman, R. A., \u0026amp; Nir-sagiv, B. (2007). Comparing Narrative and Expository Text Construction Across Adolescence: A Developmental Paradox. \u003cem\u003eDiscourse Processes\u003c/em\u003e, \u003cem\u003e43\u003c/em\u003e(2), 79\u0026ndash;120. https://doi.org/10.1080/01638530709336894\u003c/li\u003e\n \u003cli\u003eBlack, P., \u0026amp; Wiliam, D. (1998). Assessment and Classroom Learning. \u003cem\u003eAssessment in Education: Principles, Policy \u0026amp; Practice\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e(1), 7\u0026ndash;74. https://doi.org/10.1080/0969595980050102\u003c/li\u003e\n \u003cli\u003eBloom, B. S. (1968). \u003cem\u003eLearning for mastery. Instruction and curriculum. Regional education laboratory for the Carolinas and Virginia, topical papers and reprints, number 1. Evaluation Comment, 1(2)\u003c/em\u003e. https://eric.ed.gov/?id=eD053419\u003c/li\u003e\n \u003cli\u003eBloom, B. S. (1971). \u003cem\u003eMastery learning. In J. H. Block (Ed.), Mastery learning, theory and practice (pp. 47-63). New York: Holt, Rinehart, and Winston.\u003c/em\u003e\u003c/li\u003e\n \u003cli\u003eBloom, B. S., Madaus, G. F., \u0026amp; Hastings, J. T. (1981). \u003cem\u003eEvaluation to Improve Learning. New York: McGraw-Hill.\u003c/em\u003e\u003c/li\u003e\n \u003cli\u003eBoroowa, A., \u0026amp; Herodotou, C. (2022). Learning Analytics in Open and Distance Higher Education: The Case of the Open University UK. In P. Prinsloo, S. Slade, \u0026amp; M. Khalil (Eds.), \u003cem\u003eLearning Analytics in Open and Distributed Learning\u003c/em\u003e (pp. 47\u0026ndash;62). Springer Nature Singapore. https://doi.org/10.1007/978-981-19-0786-9_4\u003c/li\u003e\n \u003cli\u003eBrandt, N. D., Lechner, C. M., Tetzner, J., \u0026amp; Rammstedt, B. (2020). Personality, cognitive ability, and academic performance: Differential associations across school subjects and school tracks. \u003cem\u003eJournal of Personality\u003c/em\u003e, \u003cem\u003e88\u003c/em\u003e(2), 249\u0026ndash;265. https://doi.org/10.1111/jopy.12482\u003c/li\u003e\n \u003cli\u003eBremholm, J., Bundsgaard, J., \u0026amp; Kabel, K. (2022). Proficiency scales for early writing development. \u003cem\u003eWriting \u0026amp; Pedagogy\u003c/em\u003e, \u003cem\u003e13\u003c/em\u003e(1\u0026ndash;3), 121\u0026ndash;154. https://doi.org/10.1558/wap.21490\u003c/li\u003e\n \u003cli\u003eBritz, D., Goldie, A., Luong, M.-T., \u0026amp; Le, Q. (2017). Massive Exploration of Neural Machine Translation Architectures. \u003cem\u003eProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing\u003c/em\u003e, 1442\u0026ndash;1451. https://doi.org/10.18653/v1/D17-1151\u003c/li\u003e\n \u003cli\u003eBroadbent, J., \u0026amp; Poon, W. L. (2015). Self-regulated learning strategies \u0026amp; academic achievement in online higher education learning environments: A systematic review. \u003cem\u003eThe Internet and Higher Education\u003c/em\u003e, \u003cem\u003e27\u003c/em\u003e, 1\u0026ndash;13. https://doi.org/10.1016/j.iheduc.2015.04.007\u003c/li\u003e\n \u003cli\u003eBrown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., \u0026hellip; Amodei, D. (2020). \u003cem\u003eLanguage Models are Few-Shot Learners\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.2005.14165\u003c/li\u003e\n \u003cli\u003eBruner, J. S. (1960). \u003cem\u003eThe culture of education\u003c/em\u003e. Harvard University Press.\u003c/li\u003e\n \u003cli\u003eBusch, D. (2024). \u003cem\u003eAI translation and intercultural communication: New questions for a new field of research\u003c/em\u003e. https://doi.org/10.31235/osf.io/r3zdx\u003c/li\u003e\n \u003cli\u003eCastilho, S., Moorkens, J., Gaspari, F., Calixto, I., Tinsley, J., \u0026amp; Way, A. (2017). Is Neural Machine Translation the New State of the Art? \u003cem\u003eThe Prague Bulletin of Mathematical Linguistics\u003c/em\u003e, \u003cem\u003e108\u003c/em\u003e(1), 109\u0026ndash;120. https://doi.org/10.1515/pralin-2017-0013\u003c/li\u003e\n \u003cli\u003eChassignol, M., Khoroshavin, A., Klimova, A., \u0026amp; Bilyatdinova, A. (2018). Artificial Intelligence trends in education: A narrative overview. \u003cem\u003eProcedia Computer Science\u003c/em\u003e, \u003cem\u003e136\u003c/em\u003e, 16\u0026ndash;24. https://doi.org/10.1016/j.procs.2018.08.233\u003c/li\u003e\n \u003cli\u003eChen, L., Chen, P., \u0026amp; Lin, Z. (2020). Artificial Intelligence in Education: A Review. \u003cem\u003eIEEE Access\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e, 75264\u0026ndash;75278. https://doi.org/10.1109/ACCESS.2020.2988510\u003c/li\u003e\n \u003cli\u003eCrossley, S. (2020). Linguistic features in writing quality and development: An overview. \u003cem\u003eJournal of Writing Research\u003c/em\u003e, \u003cem\u003e11\u003c/em\u003e(vol. 11 issue 3), 415\u0026ndash;443. https://doi.org/10.17239/jowr-2020.11.03.01\u003c/li\u003e\n \u003cli\u003eCrossley, S. A., \u0026amp; Kim, M. (2022). Linguistic Features of Writing Quality and Development: A Longitudinal Approach. \u003cem\u003eThe Journal of Writing Analytics\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(1), 59\u0026ndash;93. https://doi.org/10.37514/JWA-J.2022.6.1.04\u003c/li\u003e\n \u003cli\u003eDai, Y., Jayaratne, M., \u0026amp; Jayatilleke, B. (2022). Explainable Personality Prediction Using Answers to Open-Ended Interview Questions. \u003cem\u003eFrontiers in Psychology\u003c/em\u003e, \u003cem\u003e13\u003c/em\u003e, 865841. https://doi.org/10.3389/fpsyg.2022.865841\u003c/li\u003e\n \u003cli\u003eDe Bot, K., \u0026amp; Schrauf, R. W. (2010). \u003cem\u003eLanguage Development Over the Lifespan\u003c/em\u003e (0 ed.). Routledge. https://doi.org/10.4324/9780203880937\u003c/li\u003e\n \u003cli\u003eDe Feyter, T., Caers, R., Vigna, C., \u0026amp; Berings, D. (2012). Unraveling the impact of the Big Five personality traits on academic performance: The moderating and mediating effects of self-efficacy and academic motivation. \u003cem\u003eLearning and Individual Differences\u003c/em\u003e, \u003cem\u003e22\u003c/em\u003e(4), 439\u0026ndash;448. https://doi.org/10.1016/j.lindif.2012.03.013\u003c/li\u003e\n \u003cli\u003eDeo, R. C., Yaseen, Z. M., Al-Ansari, N., Nguyen-Huy, T., Langlands, T. A. M., \u0026amp; Galligan, L. (2020). Modern Artificial Intelligence Model Development for Undergraduate Student Performance Prediction: An Investigation on Engineering Mathematics Courses. \u003cem\u003eIEEE Access\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e, 136697\u0026ndash;136724. https://doi.org/10.1109/ACCESS.2020.3010938\u003c/li\u003e\n \u003cli\u003eDevedžic, V. (2004). Web Intelligence and Artificial Intelligence in Education. \u003cem\u003eJournal of Educational Technology and Society\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e, 29\u0026ndash;39.\u003c/li\u003e\n \u003cli\u003eDimitrova, K. (2024). \u003cem\u003eUTILIZATION OF STEM APPROACH IN PRIMARY EDUCATION \u0026ndash; THEORETICAL FOUNDATIONS AND METHODOLOGICAL SOLUTIONS\u003c/em\u003e. 8448\u0026ndash;8458. https://doi.org/10.21125/edulearn.2024.2011\u003c/li\u003e\n \u003cli\u003eDrachen, A., Canossa, A., \u0026amp; Yannakakis, G., N. (2009). \u003cem\u003ePlayer Modeling using Hamari Self-Organization in Tomb Raider: Underworld\u003c/em\u003e. In Proceedings of the IEEE Symposium on Computational Intelligence and Games, 7-10. Milan, Italy, September. 10.1109/CIG.2009.5286500\u003c/li\u003e\n \u003cli\u003eDurrant, P., \u0026amp; Brenchley, M. (2019). Development of vocabulary sophistication across genres in English children\u0026rsquo;s writing. \u003cem\u003eReading and Writing\u003c/em\u003e, \u003cem\u003e32\u003c/em\u003e(8), 1927\u0026ndash;1953. https://doi.org/10.1007/s11145-018-9932-8\u003c/li\u003e\n \u003cli\u003eEuropean Parliament. (n.d.). \u003cem\u003eArtificial Intelligence Act, Corrigendum, 19 April 2024. Interinstitutional File: 2021/0106(COD)\u003c/em\u003e. Retrieved September 1, 2024, from https://www.europarl.europa.eu/doceo/document/TA-9-2024-0138-FNL-COR01_EN.pdf\u003c/li\u003e\n \u003cli\u003eFaza, A., \u0026amp; Lestari, I. A. (2025). Self-Regulated Learning in the Digital Age: A Systematic Review of Strategies, Technologies, Benefits, and Challenges. \u003cem\u003eThe International Review of Research in Open and Distributed Learning\u003c/em\u003e, \u003cem\u003e26\u003c/em\u003e(2), 23\u0026ndash;58. https://doi.org/10.19173/irrodl.v26i2.8119\u003c/li\u003e\n \u003cli\u003eGlaser, R. (1994). Instructional Technology and the Measwement of Learning Outcomes: Some Questions\u003csup\u003e1\u003c/sup\u003e. \u003cem\u003eEducational Measurement: Issues and Practice\u003c/em\u003e, \u003cem\u003e13\u003c/em\u003e(4), 6\u0026ndash;8. https://doi.org/10.1111/j.1745-3992.1994.tb00561.x\u003c/li\u003e\n \u003cli\u003eGrimshaw, G. M., Adelstein, A., Bryden, M. P., \u0026amp; MacKinnon, G. E. (1998). First-Language Acquisition in Adolescence: Evidence for a Critical Period for Verbal Language Development. \u003cem\u003eBrain and Language\u003c/em\u003e, \u003cem\u003e63\u003c/em\u003e(2), 237\u0026ndash;255. https://doi.org/10.1006/brln.1997.1943\u003c/li\u003e\n \u003cli\u003eGuskey, T. R. (1985). \u003cem\u003eImplementing Mastery Learning\u003c/em\u003e.\u003c/li\u003e\n \u003cli\u003eHall-Mills, S., \u0026amp; Apel, K. (2015). Linguistic Feature Development Across Grades and Genre in Elementary Writing. \u003cem\u003eLanguage, Speech, and Hearing Services in Schools\u003c/em\u003e, \u003cem\u003e46\u003c/em\u003e(3), 242\u0026ndash;255. https://doi.org/10.1044/2015_LSHSS-14-0043\u003c/li\u003e\n \u003cli\u003eHamari, J., \u0026amp; Tuunanen, J. (2014). Player Types: A Meta-synthesis. \u003cem\u003eTransactions of the Digital Games Research Association\u003c/em\u003e, \u003cem\u003e1\u003c/em\u003e(2). https://doi.org/10.26503/todigra.v1i2.13\u003c/li\u003e\n \u003cli\u003eHansson, K., B\u0026aring;\u0026aring;th, R., L\u0026ouml;hndorf, S., Sahl\u0026eacute;n, B., \u0026amp; Sikstr\u0026ouml;m, S. (2016). Quantifying Semantic Linguistic Maturity in Children. \u003cem\u003eJournal of Psycholinguistic Research\u003c/em\u003e, \u003cem\u003e45\u003c/em\u003e(5), 1183\u0026ndash;1199. https://doi.org/10.1007/s10936-015-9398-7\u003c/li\u003e\n \u003cli\u003eHidayat, T., Susilaningsih, E., \u0026amp; Kurniawan, C. (2018). The effectiveness of enrichment test instruments design to measure students\u0026rsquo; creative thinking skills and problem-solving. \u003cem\u003eThinking Skills and Creativity\u003c/em\u003e, \u003cem\u003e29\u003c/em\u003e, 161\u0026ndash;169. https://doi.org/10.1016/j.tsc.2018.02.011\u003c/li\u003e\n \u003cli\u003eHolmes, W., Bialik, M., \u0026amp; Fadel, C. (2019). \u003cem\u003eArtificial intelligence in education: Promises and implications for teaching and learning\u003c/em\u003e. Center for Curriculum Redesign.\u003c/li\u003e\n \u003cli\u003eHu, Z., Yang, Z., Liang, X., Salakhutdinov, R., \u0026amp; Xing, E. P. (2017). \u003cem\u003eToward Controlled Generation of Text\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.1703.00955\u003c/li\u003e\n \u003cli\u003eIfenthaler, D., \u0026amp; Yau, J. Y.-K. (2020). Utilising learning analytics to support study success in higher education: A systematic review. \u003cem\u003eEducational Technology Research and Development\u003c/em\u003e, \u003cem\u003e68\u003c/em\u003e(4), 1961\u0026ndash;1990. https://doi.org/10.1007/s11423-020-09788-z\u003c/li\u003e\n \u003cli\u003eJi, M., Bouillon, P., \u0026amp; Seligman, M. (2023). \u003cem\u003eTranslation Technology in Accessible Health Communication\u003c/em\u003e (1st ed.). Cambridge University Press. https://doi.org/10.1017/9781108938976\u003c/li\u003e\n \u003cli\u003eKahraman, H. T., Sagiroglu, S., \u0026amp; Colak, I. (2010). Development of adaptive and intelligent web-based educational systems. \u003cem\u003e2010 4th International Conference on Application of Information and Communication Technologies\u003c/em\u003e, 1\u0026ndash;5. https://doi.org/10.1109/ICAICT.2010.5612054\u003c/li\u003e\n \u003cli\u003eKarmiloff-Smith, A. (1992). \u003cem\u003eBeyond modularity: A developmental perspective on cognitive science\u003c/em\u003e. Cambridge: MIT Press.\u003c/li\u003e\n \u003cli\u003eKarunasingha, D. S. K. (2022). Root mean square error or mean absolute error? Use their ratio as well. \u003cem\u003eInformation Sciences\u003c/em\u003e, \u003cem\u003e585\u003c/em\u003e, 609\u0026ndash;629. https://doi.org/10.1016/j.ins.2021.11.036\u003c/li\u003e\n \u003cli\u003eKatzenberger, I. (2005). The Super-Structure of Written Expository Texts\u0026mdash;A Developmental Perspective. In D. D. Ravid \u0026amp; H. B.-Z. Shyldkrot (Eds.), \u003cem\u003ePerspectives on Language and Language Development\u003c/em\u003e (pp. 327\u0026ndash;336). Springer US. https://doi.org/10.1007/1-4020-7911-7_24\u003c/li\u003e\n \u003cli\u003eKluwe, R., \u0026amp; Logan, G. D. (2000). \u003cem\u003eExecutive control\u003c/em\u003e. Psychological Research, 63 [Special issue], 3-4.\u003c/li\u003e\n \u003cli\u003eKoehn, P., \u0026amp; Knowles, R. (2017). \u003cem\u003eSix Challenges for Neural Machine Translation\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.1706.03872\u003c/li\u003e\n \u003cli\u003eKolluru, V., Mungara, S., \u0026amp; Chintakunta, A. N. (2018). Adaptive Learning Systems: Harnessing AI for Customized Educational Experiences. \u003cem\u003eInternational Journal of Computational Science and Information Technology\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(3), 13\u0026ndash;26. https://doi.org/10.5121/ijcsity.2018.6302\u003c/li\u003e\n \u003cli\u003eKumar, H., Musabirov, I., Reza, M., Shi, J., Wang, X., Williams, J. J., Kuzminykh, A., \u0026amp; Liut, M. (2023). \u003cem\u003eImpact of Guidance and Interaction Strategies for LLM Use on Learner Performance and Perception\u003c/em\u003e. https://doi.org/10.48550/ARXIV.2310.13712\u003c/li\u003e\n \u003cli\u003eLarrabee S\u0026oslash;nderlund, A., Hughes, E., \u0026amp; Smith, J. (2019). The efficacy of learning analytics interventions in higher education: A systematic review. \u003cem\u003eBritish Journal of Educational Technology\u003c/em\u003e, \u003cem\u003e50\u003c/em\u003e(5), 2594\u0026ndash;2618. https://doi.org/10.1111/bjet.12720\u003c/li\u003e\n \u003cli\u003eLavou\u0026eacute;, \u0026Eacute;., Ju, Q., Hallifax, S., \u0026amp; Serna, A. (2021). Analyzing the relationships between learners\u0026rsquo; motivation and observable engaged behaviors in a gamified learning environment. \u003cem\u003eInternational Journal of Human-Computer Studies\u003c/em\u003e, \u003cem\u003e154\u003c/em\u003e, 102670. https://doi.org/10.1016/j.ijhcs.2021.102670\u003c/li\u003e\n \u003cli\u003eLim, W. M., Gunasekara, A., Pallant, J. L., Pallant, J. I., \u0026amp; Pechenkina, E. (2023). Generative AI and the future of education: Ragnar\u0026ouml;k or reformation? A paradoxical perspective from management educators. \u003cem\u003eThe International Journal of Management Education\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e(2), 100790. https://doi.org/10.1016/j.ijme.2023.100790\u003c/li\u003e\n \u003cli\u003eLiu, S., Yu, Z., Huang, F., Bulbulia, Y., Bergen, A., \u0026amp; Liut, M. (2024). Can Small Language Models With Retrieval-Augmented Generation Replace Large Language Models When Learning Computer Science? \u003cem\u003eProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1\u003c/em\u003e, 388\u0026ndash;393. https://doi.org/10.1145/3649217.3653554\u003c/li\u003e\n \u003cli\u003eLiu, Y., Wang, T., Bo, H., \u0026amp; Zhang, N. (2022). The Influence of Personality on Epistemic Network in an Open-Ended Question Discussion Scenario. \u003cem\u003e2022 4th International Conference on Computer Science and Technologies in Education (CSTE)\u003c/em\u003e, 215\u0026ndash;220. https://doi.org/10.1109/CSTE55932.2022.00046\u003c/li\u003e\n \u003cli\u003eLopez, C. E., \u0026amp; Tucker, C. S. (2019). The effects of player type on performance: A gamification case study. \u003cem\u003eComputers in Human Behavior\u003c/em\u003e, \u003cem\u003e91\u003c/em\u003e, 333\u0026ndash;345. https://doi.org/10.1016/j.chb.2018.10.005\u003c/li\u003e\n \u003cli\u003eLounsbury, J. W., Sundstrom, E., Loveland, J. M., \u0026amp; Gibson, L. W. (2003). Intelligence, \u0026ldquo;Big Five\u0026rdquo; personality traits, and work drive as predictors of course grade. \u003cem\u003ePersonality and Individual Differences\u003c/em\u003e, \u003cem\u003e35\u003c/em\u003e(6), 1231\u0026ndash;1239. https://doi.org/10.1016/S0191-8869(02)00330-6\u003c/li\u003e\n \u003cli\u003eLudwig, S., Rausch, A., Deutscher, V., \u0026amp; Seifried, J. (2024). Predicting problem-solving success in an office simulation applying N-grams and a random forest to behavioral process data. \u003cem\u003eComputers \u0026amp; Education\u003c/em\u003e, \u003cem\u003e218\u003c/em\u003e, 105093. https://doi.org/10.1016/j.compedu.2024.105093\u003c/li\u003e\n \u003cli\u003eMacArthur, C. A., Jennings, A., \u0026amp; Philippakos, Z. A. (2019). Which linguistic features predict quality of argumentative writing for college basic writers, and how do those features change with instruction? \u003cem\u003eReading and Writing\u003c/em\u003e, \u003cem\u003e32\u003c/em\u003e(6), 1553\u0026ndash;1574. https://doi.org/10.1007/s11145-018-9853-6\u003c/li\u003e\n \u003cli\u003eMalvern, D., Richards, B., Chipere, N., \u0026amp; Dur\u0026aacute;n, P. (2004). \u003cem\u003eLexical Diversity and Language Development\u003c/em\u003e. Palgrave Macmillan UK. https://doi.org/10.1057/9780230511804\u003c/li\u003e\n \u003cli\u003eMammadov, S. (2022). Big Five personality traits and academic performance: A meta‐analysis. \u003cem\u003eJournal of Personality\u003c/em\u003e, \u003cem\u003e90\u003c/em\u003e(2), 222\u0026ndash;255. https://doi.org/10.1111/jopy.12663\u003c/li\u003e\n \u003cli\u003eMarczewski, A. (2015). \u003cem\u003eUser Types\u003c/em\u003e. In Even Ninja Monkeys Like to Play: Gamification, Game Thinking and Motivational Design (1st ed.); CreateSpace Independent Publishing Platform: Scotts Valley, CA, USA; Volume 65\u0026ndash;80, 240\u0026ndash;290. https://www.gamified.uk/user-types (October 2018 update).\u003c/li\u003e\n \u003cli\u003eMirkin, S., Nowson, S., Brun, C., \u0026amp; Perez, J. (2015). Motivating Personality-aware Machine Translation. \u003cem\u003eProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing\u003c/em\u003e, 1102\u0026ndash;1108. https://doi.org/10.18653/v1/D15-1130\u003c/li\u003e\n \u003cli\u003eMoore, D. S., McCabe, G. P., \u0026amp; Craig, B. A. (2009). \u003cem\u003eIntroduction to the Practice of Statistics\u003c/em\u003e. (Vol. 4). New York: WH Freeman.\u003c/li\u003e\n \u003cli\u003eMornar, M., Maru\u0026scaron;ić, I., \u0026amp; \u0026Scaron;abić, J. (2022). Academic self-efficacy and learning strategies as mediators of the relation between personality and elementary school students\u0026rsquo; achievement. \u003cem\u003eEuropean Journal of Psychology of Education\u003c/em\u003e, \u003cem\u003e37\u003c/em\u003e(4), 1237\u0026ndash;1254. https://doi.org/10.1007/s10212-021-00576-8\u003c/li\u003e\n \u003cli\u003eNababan, M., \u0026amp; Nuraeni, A. (2012). \u003cem\u003ePengembangan model penilaian kualitas terjemahan\u003c/em\u003e. Surakarta: Kajian Linguistik dan Sastra. 24(1), 39-57.\u003c/li\u003e\n \u003cli\u003eNaqshbandi, M. M., Ainin, S., Jaafar, N. I., \u0026amp; Mohd Shuib, N. L. (2017). To Facebook or to Face Book? An investigation of how academic performance of different personalities is affected through the intervention of Facebook usage. \u003cem\u003eComputers in Human Behavior\u003c/em\u003e, \u003cem\u003e75\u003c/em\u003e, 167\u0026ndash;176. https://doi.org/10.1016/j.chb.2017.05.012\u003c/li\u003e\n \u003cli\u003eNational Association of Colleges and Employers. (2019). \u003cem\u003eJob Outlook 2020. Arkansas: Bethlehem\u003c/em\u003e. https://in.nau.edu/wp-content/uploads/sites/204/2020-nace-job-outlook.pdf\u003c/li\u003e\n \u003cli\u003eNewmark, P. (1988). \u003cem\u003eA Textbook on Translation\u003c/em\u003e. New York: Prentice-Hall International.\u003c/li\u003e\n \u003cli\u003eNicol, D. J., \u0026amp; Macfarlane‐Dick, D. (2006). Formative assessment and self‐regulated learning: A model and seven principles of good feedback practice. \u003cem\u003eStudies in Higher Education\u003c/em\u003e, \u003cem\u003e31\u003c/em\u003e(2), 199\u0026ndash;218. https://doi.org/10.1080/03075070600572090\u003c/li\u003e\n \u003cli\u003eNippold, M. A. (2000). Language Development during the Adolescent Years: Aspects of Pragmatics, Syntax, and Semantics. \u003cem\u003eTopics in Language Disorders\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e(2), 15\u0026ndash;28. https://doi.org/10.1097/00011363-200020020-00004\u003c/li\u003e\n \u003cli\u003eNippold, M. A. (2002). Lexical learning in school-age children, adolescents, and adults: A process where language and literacy converge. \u003cem\u003eJournal of Child Language\u003c/em\u003e, \u003cem\u003e29\u003c/em\u003e(2), 449\u0026ndash;488. https://doi.org/10.1017/S0305000902275340\u003c/li\u003e\n \u003cli\u003eNippold, M. A. (2016). \u003cem\u003eLater language development: School-age children, adolescents, and young adults\u003c/em\u003e. Austin, TX: PRO-ED.\u003c/li\u003e\n \u003cli\u003eO\u0026rsquo;Connor, M. C., \u0026amp; Paunonen, S. V. (2007). \u003cem\u003eBig Five personality predictors of post-secondary academic performance. Personality and Individual Differences, 43(5), 971-990.\u003c/em\u003e\u003c/li\u003e\n \u003cli\u003eOECD. (2023). \u003cem\u003eDigital Education Outlook 2023\u003c/em\u003e. https://www.oecd.org/en/publications/oecd-digital-education-outlook-2023_c74f03de-en.html\u003c/li\u003e\n \u003cli\u003ePanadero, E. (2017). A Review of Self-regulated Learning: Six Models and Four Directions for Research. \u003cem\u003eFrontiers in Psychology\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e, 422. https://doi.org/10.3389/fpsyg.2017.00422\u003c/li\u003e\n \u003cli\u003ePankiewicz, M., \u0026amp; Baker, R. S. (2023). \u003cem\u003eLarge Language Models (GPT) for automating feedback on programming assignments\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.2307.00150\u003c/li\u003e\n \u003cli\u003ePeredo, R., Canales, A., Menchaca, A., \u0026amp; Peredo, I. (2011). Intelligent Web-based education system for adaptive learning. \u003cem\u003eExpert Systems with Applications\u003c/em\u003e, \u003cem\u003e38\u003c/em\u003e(12), 14690\u0026ndash;14702. https://doi.org/10.1016/j.eswa.2011.05.013\u003c/li\u003e\n \u003cli\u003ePhillips, P., Abraham, C., \u0026amp; Bond, R. (2003). Personality, cognition, and university students\u0026rsquo; examination performance. \u003cem\u003eEuropean Journal of Personality\u003c/em\u003e, \u003cem\u003e17\u003c/em\u003e(6), 435\u0026ndash;448. https://doi.org/10.1002/per.488\u003c/li\u003e\n \u003cli\u003ePintrich, P. R., \u0026amp; De Groot, E. V. (1990). Motivational and self-regulated learning components of classroom academic performance. \u003cem\u003eJournal of Educational Psychology\u003c/em\u003e, \u003cem\u003e82\u003c/em\u003e(1), 33.\u003c/li\u003e\n \u003cli\u003ePokrivcakova, S. (2019). Preparing teachers for the application of AI-powered technologies in foreign language education. \u003cem\u003eJournal of Language and Cultural Education\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(3), 135\u0026ndash;153. https://doi.org/10.2478/jolace-2019-0025\u003c/li\u003e\n \u003cli\u003ePoropat, A. E. (2014). A meta‐analysis of adult‐rated child personality and academic performance in primary education. \u003cem\u003eBritish Journal of Educational Psychology\u003c/em\u003e, \u003cem\u003e84\u003c/em\u003e(2), 239\u0026ndash;252. https://doi.org/10.1111/bjep.12019\u003c/li\u003e\n \u003cli\u003ePrabhumoye, S., Tsvetkov, Y., Black, A. W., \u0026amp; Salakhutdinov, R. (2018). \u003cem\u003eStyle Transfer Through Multilingual and Feedback-Based Back-Translation\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.1809.06284\u003c/li\u003e\n \u003cli\u003eProverbio, A. M., \u0026amp; Zani, A. (2005). Developmental changes in the linguistic brain after puberty. \u003cem\u003eTrends in Cognitive Sciences\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(4), 164\u0026ndash;167. https://doi.org/10.1016/j.tics.2005.02.001\u003c/li\u003e\n \u003cli\u003eRabinovich, E., Mirkin, S., Patel, R. N., Specia, L., \u0026amp; Wintner, S. (2016). \u003cem\u003ePersonalized Machine Translation: Preserving Original Author Traits\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.1610.05461\u003c/li\u003e\n \u003cli\u003eRaffaghelli, J. E., Rodr\u0026iacute;guez, M. E., Guerrero-Rold\u0026aacute;n, A.-E., \u0026amp; Ba\u0026ntilde;eres, D. (2022). Applying the UTAUT model to explain the students\u0026rsquo; acceptance of an early warning system in Higher Education. \u003cem\u003eComputers \u0026amp; Education\u003c/em\u003e, \u003cem\u003e182\u003c/em\u003e, 104468. https://doi.org/10.1016/j.compedu.2022.104468\u003c/li\u003e\n \u003cli\u003eRagnarsd\u0026oacute;ttir, H., Aparici, M., Cahana-Amitay, D., Van Hell, J. G., \u0026amp; Vigui\u0026eacute;-Simon, A. (2002). Verbal structure and content in written discourse: Expository and narrative texts. \u003cem\u003eWritten Language \u0026amp; Literacy\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e(1), 95\u0026ndash;126. https://doi.org/10.1075/wll.5.1.05rag\u003c/li\u003e\n \u003cli\u003eRahayuni̇Ngsi̇H, S., Si̇Rajuddi̇N, S., \u0026amp; Ikram, M. (2021). Using Open-ended Problem-solving Tests to Identify Students\u0026rsquo; Mathematical Creative Thinking Ability. \u003cem\u003eParticipatory Educational Research\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e(3), 285\u0026ndash;299. https://doi.org/10.17275/per.21.66.8.3\u003c/li\u003e\n \u003cli\u003eRavid, D. (2005). Emergence of Linguistic Complexity in Later Language Development: Evidence from Expository Text Construction. In D. D. Ravid \u0026amp; H. B.-Z. Shyldkrot (Eds.), \u003cem\u003ePerspectives on Language and Language Development\u003c/em\u003e (pp. 337\u0026ndash;355). Springer US. https://doi.org/10.1007/1-4020-7911-7_25\u003c/li\u003e\n \u003cli\u003eRavid, D. (2006). Semantic development in textual contexts during the school years: Noun Scale analyses. \u003cem\u003eJournal of Child Language\u003c/em\u003e, \u003cem\u003e33\u003c/em\u003e(4), 791\u0026ndash;821. https://doi.org/10.1017/S0305000906007586\u003c/li\u003e\n \u003cli\u003eRavid, D., \u0026amp; Tolchinsky, L. (2002). Developing linguistic literacy: A comprehensive model. \u003cem\u003eJournal of Child Language, 29, 417\u0026ndash;447\u003c/em\u003e.\u003c/li\u003e\n \u003cli\u003eRavid, D., Van Hell, J. G., Rosado, E., \u0026amp; Zamora, A. (2002). Subject NP patterning in the development of text production: Speech and writing. \u003cem\u003eWritten Language \u0026amp; Literacy\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e(1), 69\u0026ndash;93. https://doi.org/10.1075/wll.5.1.04rav\u003c/li\u003e\n \u003cli\u003eReilly, J., Zamora, A., \u0026amp; Mcgivern, R. (2005). Acquiring perspective in English: The development of stance. \u003cem\u003eJournal of Pragmatics\u003c/em\u003e, \u003cem\u003e37\u003c/em\u003e(2), 185\u0026ndash;208. https://doi.org/10.1016/S0378-2166(04)00191-2\u003c/li\u003e\n \u003cli\u003eRello, L., Baeza-Yates, R., Ali, A., Bigham, J. P., \u0026amp; Serra, M. (2020). Predicting risk of dyslexia with an online gamified test. \u003cem\u003ePLOS ONE\u003c/em\u003e, \u003cem\u003e15\u003c/em\u003e(12), e0241687. https://doi.org/10.1371/journal.pone.0241687\u003c/li\u003e\n \u003cli\u003eRichards, B. J. \u0026amp; Malvern, D. D. (1997). \u003cem\u003eThe new Bulmershe papers. Quantifying lexical diversity in the study of language development\u003c/em\u003e. Reading: The University of Reading.\u003c/li\u003e\n \u003cli\u003eRimmer, W. (2008). Putting grammatical complexity in context. \u003cem\u003eLiteracy\u003c/em\u003e, \u003cem\u003e42\u003c/em\u003e(1), 29\u0026ndash;35. https://doi.org/10.1111/j.1467-9345.2008.00478.x\u003c/li\u003e\n \u003cli\u003eR\u0026iacute;o Lugo, N. D. (1996). David R. Olson, The world on paper. The conceptual and cognitive implications of writing and reading. Cambridge University Press, Cambridge, 1994; 318 pp. \u003cem\u003eNueva Revista de Filolog\u0026iacute;a Hisp\u0026aacute;nica (NRFH)\u003c/em\u003e, \u003cem\u003e44\u003c/em\u003e(1), 197\u0026ndash;200. https://doi.org/10.24201/nrfh.v44i1.1919\u003c/li\u003e\n \u003cli\u003eRosselli, M., Ardila, A., Matute, E., \u0026amp; V\u0026eacute;lez-Uribe, I. (2014). Language Development across the Life Span: A Neuropsychological/Neuroimaging Perspective. \u003cem\u003eNeuroscience Journal\u003c/em\u003e, \u003cem\u003e2014\u003c/em\u003e, 1\u0026ndash;21. https://doi.org/10.1155/2014/585237\u003c/li\u003e\n \u003cli\u003eRovira, S., Puertas, E., \u0026amp; Igual, L. (2017). Data-driven system to predict academic grades and dropout. \u003cem\u003ePLOS ONE\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e(2), e0171207. https://doi.org/10.1371/journal.pone.0171207\u003c/li\u003e\n \u003cli\u003eRus, V., D\u0026rsquo;Mello, S., Hu, X., \u0026amp; Graesser, A. C. (2013). Recent Advances in Conversational Intelligent Tutoring Systems. \u003cem\u003eAI Magazine\u003c/em\u003e, \u003cem\u003e34\u003c/em\u003e(3), 42\u0026ndash;54. https://doi.org/10.1609/aimag.v34i3.2485\u003c/li\u003e\n \u003cli\u003eSajja, R., Sermet, Y., Cwiertny, D., \u0026amp; Demir, I. (2023). Platform-independent and curriculum-oriented intelligent assistant for higher education. \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e(1), 42. https://doi.org/10.1186/s41239-023-00412-7\u003c/li\u003e\n \u003cli\u003eSalas-Pilco, S. Z., \u0026amp; Yang, Y. (2022). Artificial intelligence applications in Latin American higher education: A systematic review. \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(1), 21. https://doi.org/10.1186/s41239-022-00326-w\u003c/li\u003e\n \u003cli\u003eSamuel Girard, Jill-J\u0026ecirc;nn Vie, Fran\u0026ccedil;oise Tort, \u0026amp; Amel Bouzeghoub. (2024). \u003cem\u003eOptimizing Human Learning using Reinforcement Learning\u003c/em\u003e. https://doi.org/10.5281/ZENODO.12730016\u003c/li\u003e\n \u003cli\u003eSantos, A. C. G., Oliveira, W., Hamari, J., Joaquim, S., \u0026amp; Isotani, S. (2023). The Consistency of Gamification User Types: A Study on the Change of Preferences over Time. \u003cem\u003eProceedings of the ACM on Human-Computer Interaction\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(CHI PLAY), 1253\u0026ndash;1281. https://doi.org/10.1145/3611068\u003c/li\u003e\n \u003cli\u003eSato, T. (2022). Assessing critical thinking through L2 argumentative essays: An investigation of relevant and salient criteria from raters\u0026rsquo; perspectives. \u003cem\u003eLanguage Testing in Asia\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e(1), 9. https://doi.org/10.1186/s40468-022-00159-4\u003c/li\u003e\n \u003cli\u003eSennrich, R., Haddow, B., \u0026amp; Birch, A. (2016). Controlling Politeness in Neural Machine Translation via Side Constraints. \u003cem\u003eProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies\u003c/em\u003e, 35\u0026ndash;40. https://doi.org/10.18653/v1/N16-1005\u003c/li\u003e\n \u003cli\u003eShakir, A., \u0026amp; Obeidat, H. (1991). Maturity in AFL Student-Written Texts: A case study. \u003cem\u003eAl-\u0026rsquo;Arabiyya, 24, 65\u0026ndash;81. Http://Www.Jstor.Org/Stable/43192653\u003c/em\u003e.\u003c/li\u003e\n \u003cli\u003eShen, T., Lei, T., Barzilay, R., \u0026amp; Jaakkola, T. (2017). \u003cem\u003eStyle Transfer from Non-Parallel Text by Cross-Alignment\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.1705.09655\u003c/li\u003e\n \u003cli\u003eShuliang, M. (2025). \u003cem\u003eDevelopmental constructivism\u003c/em\u003e. Springer Nature Singapore.\u003c/li\u003e\n \u003cli\u003eSlade, S., \u0026amp; Prinsloo, P. (2013). Learning Analytics: Ethical Issues and Dilemmas. \u003cem\u003eAmerican Behavioral Scientist\u003c/em\u003e, \u003cem\u003e57\u003c/em\u003e(10), 1510\u0026ndash;1529. https://doi.org/10.1177/0002764213479366\u003c/li\u003e\n \u003cli\u003eSubirats, L., Nousiainen, T., Hooda, A., Rubio-Andrada, L., Fort, S., Vesisenaho, M., \u0026amp; Sacha, G. M. (2023). Gamification Based on User Types: When and Where It Is Worth Applying. \u003cem\u003eApplied Sciences\u003c/em\u003e, \u003cem\u003e13\u003c/em\u003e(4), 2269. https://doi.org/10.3390/app13042269\u003c/li\u003e\n \u003cli\u003eSubirats, L., Palacios Corral, A., P\u0026eacute;rez-Ruiz, S., Fort, S., \u0026amp; Sacha, G.-M. (2023). Temporal analysis of academic performance in higher education before, during and after COVID-19 confinement using artificial intelligence. \u003cem\u003ePLOS ONE\u003c/em\u003e, \u003cem\u003e18\u003c/em\u003e(2), e0282306. https://doi.org/10.1371/journal.pone.0282306\u003c/li\u003e\n \u003cli\u003eSun, K., \u0026amp; Xiong, W. (2019). A computational model for measuring discourse complexity. \u003cem\u003eDiscourse Studies\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e(6), 690\u0026ndash;712. https://doi.org/10.1177/1461445619866985\u003c/li\u003e\n \u003cli\u003eTempelaar, D. T., Rienties, B., \u0026amp; Nguyen, Q. (2017). Towards Actionable Learning Analytics Using Dispositions. \u003cem\u003eIEEE Transactions on Learning Technologies\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(1), 6\u0026ndash;16. https://doi.org/10.1109/TLT.2017.2662679\u003c/li\u003e\n \u003cli\u003eTimms, M. J. (2016). Letting Artificial Intelligence in Education Out of the Box: Educational Cobots and Smart Classrooms. \u003cem\u003eInternational Journal of Artificial Intelligence in Education\u003c/em\u003e, \u003cem\u003e26\u003c/em\u003e(2), 701\u0026ndash;712. https://doi.org/10.1007/s40593-016-0095-y\u003c/li\u003e\n \u003cli\u003eTolchinsky, L., \u0026amp; Rosado, E. (2005). The effect of literacy, text type, and modality on the use of grammatical means for agency alternation in Spanish. \u003cem\u003eJournal of Pragmatics\u003c/em\u003e, \u003cem\u003e37\u003c/em\u003e(2), 209\u0026ndash;237. https://doi.org/10.1016/S0378-2166(04)00195-X\u003c/li\u003e\n \u003cli\u003eTomlinson, C. A. (2017). \u003cem\u003eHow to differentiate instruction in academically diverse classrooms (3rd ed.). ASCD.\u003c/em\u003e\u003c/li\u003e\n \u003cli\u003eTondello, G. F., Mora, A., Marczewski, A., \u0026amp; Nacke, L. E. (2019). Empirical validation of the Gamification User Types Hexad scale in English and Spanish. \u003cem\u003eInternational Journal of Human-Computer Studies\u003c/em\u003e, \u003cem\u003e127\u003c/em\u003e, 95\u0026ndash;111. https://doi.org/10.1016/j.ijhcs.2018.10.002\u003c/li\u003e\n \u003cli\u003eTondello, G. F., Wehbe, R. R., Diamond, L., Busch, M., Marczewski, A., \u0026amp; Nacke, L. E. (2016). The Gamification User Types Hexad Scale. \u003cem\u003eProceedings of the 2016 Annual Symposium on Computer-Human Interaction in Play\u003c/em\u003e, 229\u0026ndash;243. https://doi.org/10.1145/2967934.2968082\u003c/li\u003e\n \u003cli\u003eUNESCO. (2019). \u003cem\u003eBeijing Consensus on Artificial Intelligence and Education\u003c/em\u003e. https://unesdoc.unesco.org/ark:/48223/pf0000368303\u003c/li\u003e\n \u003cli\u003eVanmassenhove, E. (2024). \u003cem\u003eGender Bias in Machine Translation and The Era of Large Language Models\u003c/em\u003e. arXiv. https://doi.org/10.48550/ARXIV.2401.10016\u003c/li\u003e\n \u003cli\u003eVygotsky, L. S. (1978). \u003cem\u003eMind in society: The development of higher psychological processes\u003c/em\u003e. Harvard University Press.\u003c/li\u003e\n \u003cli\u003eWagner, R. K., Puranik, C. S., Foorman, B., Foster, E., Wilson, L. G., Tschinkel, E., \u0026amp; Kantor, P. T. (2011). Modeling the development of written language. \u003cem\u003eReading and Writing\u003c/em\u003e, \u003cem\u003e24\u003c/em\u003e(2), 203\u0026ndash;220. https://doi.org/10.1007/s11145-010-9266-7\u003c/li\u003e\n \u003cli\u003eWijaksono, R. N., Hilman, E. H., \u0026amp; Mustolih, A. (2022). TRANSLATION METHODS AND QUALITY OF IDIOMATIC EXPRESSION IN MY SISTER\u0026rsquo;S KEEPER MOVIE. \u003cem\u003eJURNAL BASIS\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(1), 73\u0026ndash;84. https://doi.org/10.33884/basisupb.v9i1.5428\u003c/li\u003e\n \u003cli\u003eZhang, Z., Zhang, E., Liu, H., \u0026amp; Han, S. (2024). Examining the association between discussion strategies and learners\u0026rsquo; critical thinking in asynchronous online discussion. \u003cem\u003eThinking Skills and Creativity\u003c/em\u003e, \u003cem\u003e53\u003c/em\u003e, 101588. https://doi.org/10.1016/j.tsc.2024.101588\u003c/li\u003e\n \u003cli\u003eZimmerman, B. J. (2002). Becoming a Self-Regulated Learner: An Overview. \u003cem\u003eTheory Into Practice\u003c/em\u003e, \u003cem\u003e41\u003c/em\u003e(2), 64\u0026ndash;70. https://doi.org/10.1207/s15430421tip4102_2\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Natural Language Processing, gamification, self-regulated learning, machine learning, academic performance, higher education","lastPublishedDoi":"10.21203/rs.3.rs-7801827/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7801827/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eArtificial Intelligence and Natural Language Processing are widely used to predict academic performance across education levels. This study examines how students\u0026rsquo; personality and maturity relate to academic progress by comparing their actual degree program year with predictions generated through advanced Artificial Intelligence and Natural Language Processing techniques. Our predictive model achieved a mean absolute error of 0.794. Maturity and personality traits were derived from two sources: (a) Type-Token Ratio and Flesch-Kincaid Grade used to assess developmental writing features; and (b) Hexad Model player-types to classify motivational traits. Both data types were collected through post-task surveys following a gamified activity. The model was applied to students across all four years of a Tourism degree program. Results show that NLP-based linguistic features and gamification profiles were the strongest predictors. The data confirmed expected gains in lexical diversity and readability across academic years. Additionally, shifts in player types \u0026ndash; from first to fourth year \u0026ndash; suggest evolving motivational orientations and personality development. These findings offer valuable insights for identifying students whose developmental trajectory may not align with their academic standing. When predicted and actual program year differ, educators can use this signal to provide targeted, timely support. AI-based writing analysis fosters maturity monitoring, metacognitive awareness and self-regulated learning. Meanwhile, personality profiling enables differentiated instruction based on motivational drivers. This methodology offers a scalable tool for inclusive, personalized education \u0026ndash; particularly in multilingual settings where accurate translation preserves students\u0026rsquo; linguistic voice.\u003c/p\u003e","manuscriptTitle":"Predictive AI for Academic Performance: Integrating Natural Language Processing, Gamification and Self-regulated Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-21 06:09:20","doi":"10.21203/rs.3.rs-7801827/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e1ddf4c5-b0d3-4f02-ad3f-9c511169ba08","owner":[],"postedDate":"November 21st, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-03T09:12:06+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-21 06:09:20","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7801827","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7801827","identity":"rs-7801827","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.