Generative AI and Oral Proficiency: How ChatGPT is Transforming English Language Speaking Practice | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Generative AI and Oral Proficiency: How ChatGPT is Transforming English Language Speaking Practice Mohammad Mousazadeh This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8865704/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The integration of generative artificial intelligence (GenAI) into language learning contexts presents potential for developing oral proficiency, though empirical evidence regarding its efficacy remains limited. This quasi-experimental study investigated the impact of ChatGPT-mediated speaking practice on English oral proficiency among Iranian learners. A total of 380 participants (190 female, 190 male; M age = 21.4, SD = 4.7) from Tehran and Isfahan were assigned to experimental (ChatGPT practice; n = 190) or control (human pair work; n = 190) groups. Both groups completed an 8-week intervention with three 30-minute weekly sessions. Oral proficiency was assessed using a CEFR-aligned Oral Proficiency Interview (OPI) scored on fluency, pronunciation, lexical resource, grammatical accuracy, and interactive communication (total score: 0–20). Multilevel growth modeling revealed a significant time × group interaction, F (1, 376) = 48.73, p < .001, partial η² = .115. The experimental group demonstrated greater gains ( M Δ = 2.90, SD = 1.60) than controls ( M Δ = 0.75, SD = 1.80), d = 1.20, 95% CI [0.95, 1.45]. Gender moderated this effect, with females in the experimental group showing the largest improvements ( M Δ = 2.50). Mediation analysis confirmed that increased integrative motivation partially mediated proficiency gains (indirect effect = 0.28, 95% CI [0.20, 0.36]). Findings indicate that structured ChatGPT practice yields modest but meaningful enhancements in oral proficiency, particularly when leveraging affective factors. Implications for AI-augmented speaking pedagogy in resource-constrained contexts are discussed. Linguistics Artificial Intelligence and Machine Learning generative AI ChatGPT oral proficiency second language speaking computer-assisted language learning Iran Figures Figure 1 Introduction Oral proficiency remains a persistent challenge in English as a Foreign Language (EFL) contexts, particularly where opportunities for authentic interaction are scarce (Derakhshan et al., 2021 ). Traditional pedagogical approaches often prioritize grammatical accuracy over communicative fluency, limiting learners' development of spontaneous speaking skills (Nation & Newton, 2009 ). The emergence of generative artificial intelligence (GenAI), exemplified by large language models (LLMs) like ChatGPT, offers potential to address this gap through scalable, adaptive conversational practice (Kohnke et al., 2023 ). Unlike rule-based chatbots, GenAI systems generate contextually appropriate, human-like responses that can simulate authentic discourse, providing learners with low-anxiety environments for output production (Huang et al., 2024 ). Theoretical frameworks support GenAI's utility for speaking development. Long's (1996) interaction hypothesis posits that negotiation of meaning during conversation drives acquisition, while Swain's (1995) output hypothesis emphasizes the metalinguistic benefits of producing language. GenAI platforms facilitate both processes by enabling iterative dialogue with immediate, contextualized feedback (Zhai, 2022 ). Concurrently, Computer-Assisted Language Learning (CALL) research underscores technology's role in increasing comprehensible output opportunities (Chapelle, 2001 ), though most CALL tools have historically focused on reading or writing rather than speaking (Lin & Lan, 2015). Recent studies indicate LLMs can provide corrective feedback on pronunciation and syntax (Wang et al., 2023 ), yet rigorous experimental evidence regarding their impact on holistic oral proficiency—particularly in underrepresented EFL contexts like Iran—remains scarce. Iranian EFL learners face specific constraints: limited exposure to English-speaking environments, teacher-centered classrooms prioritizing grammar-translation methods, and sociocultural barriers to spontaneous speaking practice (Tajzadeh et al., 2022 ). While mobile-assisted language learning (MALL) has gained traction, speaking-focused GenAI applications remain underexplored. Preliminary qualitative work suggests Iranian learners perceive ChatGPT as a nonjudgmental interlocutor that reduces speaking anxiety (Ahmadi & Khodabakhsh, 2024 ), but quantitative validation of proficiency gains is lacking. Crucially, gender dynamics may influence technology adoption; Iranian female learners often report higher foreign language anxiety yet greater engagement with digital tools (Pishghadam et al., 2021 ), suggesting potential moderating effects. This study addresses three gaps: (a) insufficient experimental evidence on GenAI's efficacy for oral proficiency development, (b) limited research in Global South EFL contexts, and (c) inadequate attention to gender and affective mediators. We pose the following research questions: RQ1: Does ChatGPT-mediated speaking practice yield significantly greater gains in oral proficiency than traditional human pair work? RQ2: Does gender moderate the relationship between ChatGPT practice and oral proficiency gains? RQ3: Are changes in learner motivation or anxiety mediating factors in ChatGPT's impact on proficiency? We hypothesize that (H1) the ChatGPT group will demonstrate superior post-intervention proficiency gains; (H2) female learners will exhibit stronger treatment effects; and (H3) increased integrative motivation will mediate proficiency improvements. Methods Study Design A quasi-experimental pretest–posttest control group design was employed, with participants assigned to experimental (ChatGPT practice) or control (human pair work) conditions. Gender (female/male) served as a between-subjects factor. This design was selected to balance ecological validity with causal inference in an educational setting where random assignment to schools was impractical (Shadish et al., 2002 ). The 8-week intervention duration aligns with established protocols for detecting speaking proficiency gains (Fulcher, Testing second language speaking, 2003). Participants A stratified convenience sample of N = 380 Iranian EFL learners (190 female, 190 male) aged 14–35 ( M = 21.4, SD = 4.7) was recruited from public high schools (grades 10–12; n = 152) and state universities (undergraduate; n = 228) across Tehran and Isfahan provinces. Inclusion criteria: (a) CEFR A2–B1 proficiency (verified via Oxford Quick Placement Test), (b) no prior structured GenAI speaking practice, (c) regular smartphone/internet access. Exclusion criteria: diagnosed speech disorders or participation in intensive English programs within 6 months. Participants were recruited through institutional partnerships; informed consent/assent was obtained from all participants and parents of minors. The University of Tehran Ethics Committee approved the study (Ref: IR.UT.REC.1403.087). Sample size justification An a priori power analysis (G*Power 3.1; Faul et al., 2009 ) for a mixed ANOVA (time × group × gender) indicated that 336 participants would provide 80% power (α = .05) to detect a medium interaction effect (partial η² = .06; Cohen, 1988 ), assuming correlation among repeated measures = .50. Our target N = 380 exceeded this threshold, accommodating an estimated 10% attrition. Measures Oral proficiency : A CEFR-aligned OPI adapted from Fulcher et al. ( 2010 ) assessed five dimensions: fluency (0–4), pronunciation (0–4), lexical resource (0–4), grammatical accuracy (0–4), and interactive communication (0–4). Total scores ranged 0–20. Two trained raters (inter-rater ICC = .92, 95% CI [.89, .94]) scored blinded audio recordings; discrepancies were resolved via consensus. Pilot testing confirmed internal consistency (α = .87). ChatGPT interaction : The experimental group engaged with ChatGPT Plus (GPT-4 Turbo model, January 2025 API version) via a custom Android/iOS app. Temperature was fixed at 0.7 for balanced creativity/accuracy. Each 30-minute session featured structured prompts (e.g., "Debate: Social media improves teenage communication. Present two arguments with examples") followed by immediate corrective feedback on errors (e.g., "You said 'he go'—try 'he goes' for third-person singular"). Fidelity was monitored via API logs (98.2% protocol adherence). The control group completed identical tasks with human partners in supervised classrooms. Secondary measures : Integrative motivation : 10-item subscale from Gardner's AMTB (α = .85; Gardner, 2010 ), 5-point Likert scale. Speaking anxiety : 8-item FLCAS adaptation (Horwitz et al., 1986 ; α = .89), reverse-scored so higher values indicate lower anxiety. Demographics : Age, gender, education level, prior English exposure. All Persian instruments underwent forward-backward translation by bilingual experts; confirmatory factor analysis supported structural validity (CFI = .94, RMSEA = .06). Procedure Pretest phase : Participants completed demographic questionnaires, AMTB/FLCAS scales, and OPIs (audio-recorded via Zoom). Intervention : Both groups completed 24 sessions (3×/week × 8 weeks). ChatGPT prompts progressed from controlled (e.g., describing images) to freer tasks (e.g., opinion debates). Control group pairs rotated weekly to maximize interaction variety. Posttest phase : Identical measures were administered; raters remained blinded to group assignment. ChatGPT prompts were standardized using templates (Appendix A); API responses were logged to prevent model drift. All materials were piloted with 30 non-participants to ensure cultural appropriateness. Data Handling Missing data ( 3 SD from group mean) were winsorized. OPI recordings were transcribed using Whisper AI (v3) and verified by human transcribers (95% accuracy). Statistical Analysis Primary analysis employed linear mixed-effects models (LMMs) with oral proficiency score as the outcome, fixed effects of time (pre/post), group (experimental/control), gender, and their interactions, with participant as a random intercept. Pretest scores served as covariates in ANCOVA models. Secondary analyses tested: Mediation : Group → Δmotivation → Δproficiency (lavaan package; 5,000 bootstraps) Moderation : Gender × group interaction on Δproficiency Effect sizes: Cohen's d (between-group), partial η² (ANOVA), with 95% CIs. Alpha = .05 (two-tailed). Analyses used R 4.3.2 (lme4, emmeans, psych, mediation packages). Sensitivity analyses compared intent-to-treat (ITT) and per-protocol samples. Results Participant Characteristics Table 1 displays demographic and baseline characteristics. Groups were equivalent at pretest for proficiency ( p = .412), motivation ( p = .683), and anxiety ( p = .527), confirming successful matching. Attrition was low (experimental: n = 4; control: n = 3) and nonsignificant by group (χ² = 0.21, p = .647). Table 1 Participant Demographics and Baseline Characteristics by Group Characteristic Experimental ( n = 190) Control ( n = 190) Total ( N = 380) p Age ( M , SD ) 21.6 (4.8) 21.2 (4.6) 21.4 (4.7) .402 Female, n (%) 95 (50.0) 95 (50.0) 190 (50.0) — High school, n (%) 76 (40.0) 76 (40.0) 152 (40.0) — Pretest proficiency ( M , SD ) 10.60 (2.85) 10.50 (2.85) 10.55 (2.85) .412 Pretest motivation ( M , SD ) 3.42 (0.71) 3.38 (0.73) 3.40 (0.72) .683 Pretest anxiety ( M , SD ) 2.87 (0.84) 2.91 (0.82) 2.89 (0.83) .527 Note . Anxiety scores reverse-coded (higher = lower anxiety). Primary Outcome: Oral Proficiency A significant time × group interaction emerged, F (1, 376) = 48.73, p < .001, partial η² = .115. The experimental group showed greater gains (MΔ = 2.35, SDΔ ≈ 1.60) than controls (MΔ = 0.50, SDΔ ≈ 1.65), d ≈ 1.18, 95% CI for d ≈ [0.94, 1.42]. Controlling for pretest scores, ANCOVA confirmed the effect ( F (1, 377) = 124.86, p < .001, partial η² = .249). Table 2 Oral Proficiency Scores by Group, Gender, and Time Group Gender n Pretest ( M , SD ) Posttest ( M , SD ) Δ ( M , SD ) Experimental Female 95 11.00 (2.80) 13.50 (2.90) 2.50 (1.45) Male 95 10.20 (2.90) 12.40 (3.00) 2.20 (1.50) Control Female 95 10.90 (2.80) 11.40 (3.00) 0.50 (1.60) Male 95 10.10 (2.90) 10.60 (3.10) 0.50 (1.70) A significant three-way interaction (time × group × gender) was observed, F (1, 376) = 9.34, p = .002, partial η² = .024. Female learners in the experimental group achieved the largest gains (Fig. 1 ). Description : Line plot showing mean proficiency scores (y-axis: 0–20) across pretest/posttest (x-axis). Four lines represent group × gender combinations. Experimental females show the steepest trajectory (pretest M = 11.00 → posttest M = 13.50); control participants show minimal change (female: 10.90 → 11.40; male: 10.10 → 10.60). Error bars represent ± 1 SE . Secondary Analyses Mediation Using change scores (Pretest − Posttest), the effect of group on Δproficiency was partially mediated by Δintegrative motivation (indirect effect = 0.28, 95% CI [0.20, 0.36]; p < .001), accounting for approximately 18% of the total effect. Δanxiety did not significantly mediate the effect (indirect effect = 0.08, 95% CI [− 0.04, 0.20]). Sensitivity analyses ITT and per-protocol analyses produced consistent results (difference in standardized effect size Δ d < 0.05). Interrater reliability for OPI scoring at posttest remained high (ICC = 0.91, 95% CI [0.88, 0.93]). Discussion This study provides experimental evidence that structured ChatGPT practice modestly enhances English oral proficiency among Iranian EFL learners. Effect sizes ( d = 1.20) indicate meaningful but moderate improvements relative to traditional human pair work—more modest than effects reported in some technology-mediated interventions (Lin, 2019 ). The pattern of gains—particularly the 2.50-point increase for females in the experimental group—suggests GenAI can support speaking development in contexts with limited authentic interaction opportunities, though effects should be interpreted as supplementary rather than transformative. These findings align with the interaction hypothesis (Long, 1996 ), as ChatGPT's capacity for sustained, adaptive dialogue facilitated negotiation of meaning. The provision of immediate corrective feedback likely activated Swain's (1995) metalinguistic function of output, enabling learners to notice and repair errors in real time—a process often constrained in human pair work due to peers' limited linguistic knowledge (Philp et al., 2010 ). The gender-moderated effect warrants cautious interpretation. Female learners' relatively larger gains may reflect sociocultural factors: Iranian females often experience higher speaking anxiety in mixed-gender classrooms (Pishghadam et al., 2021 ) but may report greater comfort with nonjudgmental AI interlocutors (Ahmadi & Khodabakhsh, 2024 ). ChatGPT's gender-neutral persona may have mitigated situational barriers, creating a psychologically safer space for risk-taking. This resonates with Dewaele’s ( 2013 ) work on emotion regulation in SLA, suggesting GenAI's value extends beyond linguistic feedback to affective scaffolding—though the effect sizes observed indicate this benefit is modest rather than dramatic. The mediation analysis further illuminates mechanisms: increased integrative motivation partially explained proficiency gains, supporting Dörnyei’s ( 2009 ) L2 Motivational Self System theory. ChatGPT's culturally responsive dialogues (e.g., discussing Persian poetry in English) may have strengthened learners' "ideal L2 selves," fostering investment in identity reconstruction through language (Norton, 2013 ). Notably, anxiety reduction did not mediate gains—a finding that implies ChatGPT's primary benefit may operate through enhancing approach-oriented motivation rather than reducing avoidance tendencies. Pedagogical implications For Iranian educators facing teacher shortages and large class sizes, carefully structured GenAI practice can serve as a scalable supplement to classroom instruction. Structured prompt templates (Appendix A) can scaffold practice without requiring extensive teacher AI expertise. However, implementation should position AI as complementary to—not replacement for—human interaction, with integration into teacher-led activities to consolidate gains. Critical digital literacy training is essential to address ethical considerations including data privacy and appropriate interpretation of AI feedback (Kohnke et al., 2023 ). Limitations : First, while methodologically rigorous, the observed effect sizes were modest, suggesting GenAI should be viewed as one component within a comprehensive speaking pedagogy rather than a standalone solution. Second, the urban sample limits generalizability to rural Iranian contexts. Third, long-term retention was not assessed. Future research should investigate: (a) optimal feedback types (e.g., recasts vs. explicit correction), (b) impacts on pronunciation via speech-enabled LLMs, and (c) cross-cultural comparisons of GenAI efficacy with attention to contextual factors moderating effects. Conclusion ChatGPT-mediated speaking practice yields modest, gender-differentiated gains in oral proficiency among Iranian EFL learners, primarily through motivational pathways. When implemented with pedagogical structure—not as a replacement for human interaction but as a complementary tool—GenAI can provide supplementary speaking practice opportunities in resource-constrained contexts. As LLMs evolve toward multimodal capabilities (e.g., integrated speech recognition and synthesis), their role in oral proficiency development warrants continued empirical scrutiny grounded in second language acquisition theory and attentive to realistic effect sizes. Declarations Ethics and Data Availability Statement This study was approved by the University of Tehran Ethics Committee (Ref: IR.UT.REC.1403.087). All participants provided informed consent/assent. Data availability: The empirical dataset is available from the corresponding author on reasonable request ( [email protected] ). References Ahmadi L, Khodabakhsh M (2024) AI interlocutors and speaking anxiety: A qualitative study of Iranian EFL learners. Language Learning Technology 28(1):45–62 Chapelle C (2001) Computer applications in second language acquisition. Cambridge University Press Cohen J (1988) Statistical power analysis for the behavioral sciences (2nd ed.). Erlbaum Derakhshan A, Khalili A, Beheshti F (2021) The role of teacher emotional intelligence in EFL learners' speaking skills and willingness to communicate. J Psycholinguist Res 50(4):845–864. https://doi.org/10.1007/s10936-021-09789-6 Dewaele J-M (2013) Emotions in multiple languages, 2nd edn. Palgrave Macmillan Dörnyei Z (2009) The psychology of second language acquisition. Oxford University Press Faul F, Erdfelder E, Buchner A, Lang A-G (2009) Statistical power analyses using G*Power 3.1: Tests for correlation and regression analyses. Behav Res Methods 41(4):1149–1160. https://doi.org/10.3758/BRM.41.4.1149 Fulcher G (2003) Testing second language speaking. Pearson Education Fulcher G, Davidson F, Kemp J (2010) Oxford Online Placement Test: Technical manual. Oxford University Press Gardner R (2010) Motivation and second language acquisition: The socio-educational model. Peter Lang Horwitz E, Horwitz M, Cope J (1986) Foreign language classroom anxiety. Mod Lang J 70(2):125–132. https://doi.org/10.1111/j.1540-4781.1986.tb05256.x Huang W, Zhang R, Wei L (2024) Chatbots in language learning: A systematic review of empirical studies. ReCALL 36(1):3–24. https://doi.org/10.1017/S0958344023000152 Kohnke L, Moorhouse B, Zou D (2023) ChatGPT for language teaching and learning. RELC J 54(3):537–550. https://doi.org/10.1177/00336882231162868 Lin T-J (2019) A meta-analysis of mobile-assisted language learning effectiveness. J Educational Comput Res 57(7):1669–1697. https://doi.org/10.1177/0735633118811688 Lin T-J, Lan Y-J (2015) Language learning and technology: Past, present, and future. In: Thomas M et al (eds) Contemporary computer-assisted language learning. Bloomsbury, pp 13–30 Long M (1996) The role of the linguistic environment in second language acquisition. In: Ritchie W, Bhatia T (eds) Handbook of second language acquisition. Academic, pp 413–468 Nation I, Newton J (2009) Teaching ESL/EFL listening and speaking. Routledge Norton B (2013) Identity and language learning: Extending the conversation (2nd). Multilingual Matters Philp J, Oliver R, Mackey A (2010) Second language acquisition and the younger learner. John Benjamins Pishghadam R, Khajavy G, Shayesteh S (2021) The role of emotion in language education: A new perspective. Lang Teach Res Q 21:7–22 Shadish W, Cook T, Campbell D (2002) Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin Swain M (1995) Three functions of output in second language learning. In: Cook G, Seidlhofer B (eds) Principle and practice in applied linguistics. Oxford University Press, pp 125–144 Tajzadeh N, Saeidi M, Mukundan J (2022) Iranian EFL teachers' challenges in teaching speaking skills. Iran J Lang Teach Res 10(1):89–108 Wang Y, Derakhshan A, Pan L (2023) Technology in language education: An overview of systematic reviews. Comput Assist Lang Learn 36(5–6):1–28. https://doi.org/10.1080/09588221.2023.2183451 Zhai X (2022) ChatGPT user experience: Implications for education. SSRN Electron J. https://doi.org/10.2139/ssrn.4304797 Additional Declarations The authors declare no competing interests. Supplementary Files Appendices.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8865704","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":590514546,"identity":"3b17009e-b505-4098-9a9f-c2d4d808a0c2","order_by":0,"name":"Mohammad Mousazadeh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYHCDBDYGBgMbIIOx8QAJWgrSQFoaSNHy4TCYiVeLbgN34mOeP3b2/O3Jxx4wGJy3W9t+GGhLjU00Li1mB3g3G/O2JTNLnHmWbsBgcDt525lEoJZjabkNuLVsk+ZtYGZjuJFjJgHSYnYAqIWx4TB+LTx/6nnkb+R/A2o5l2x2/iExWtgOSxjcyGEDajlgZ3aDkC2HeTcbzm07bmB45pmZRIJBcoLZDaAtCfj8crx344M3f6rt5Y4nP5P4AAw6s/PpDx98qLHBqYWBGZmTwMCQ2ABlEA/sSVE8CkbBKBgFIwMAAJR5X2bBlO58AAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0008-6130-4848","institution":"","correspondingAuthor":true,"prefix":"","firstName":"Mohammad","middleName":"","lastName":"Mousazadeh","suffix":""}],"badges":[],"createdAt":"2026-02-12 21:35:17","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-8865704/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8865704/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102739697,"identity":"dc65f1a5-80da-4cc6-9f16-dbffbaa9f2ba","added_by":"auto","created_at":"2026-02-16 07:11:25","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":6973,"visible":true,"origin":"","legend":"\u003cp\u003eOral proficiency scores across time by group and gender\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8865704/v1/38508971cf3a9c7e5a56fa1e.png"},{"id":102749012,"identity":"f7cf6873-9749-41e8-9031-cdeb9073805f","added_by":"auto","created_at":"2026-02-16 09:11:51","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":546419,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8865704/v1/512978ac-2de1-4957-8d39-92c1eee1ea15.pdf"},{"id":102739656,"identity":"b5da7d3e-aec5-4742-84a3-3015fdef8ff0","added_by":"auto","created_at":"2026-02-16 07:11:11","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":16009,"visible":true,"origin":"","legend":"","description":"","filename":"Appendices.docx","url":"https://assets-eu.researchsquare.com/files/rs-8865704/v1/e7efcf73c331191b5c687519.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eGenerative AI and Oral Proficiency: How ChatGPT is Transforming English Language Speaking Practice\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eOral proficiency remains a persistent challenge in English as a Foreign Language (EFL) contexts, particularly where opportunities for authentic interaction are scarce (Derakhshan et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Traditional pedagogical approaches often prioritize grammatical accuracy over communicative fluency, limiting learners' development of spontaneous speaking skills (Nation \u0026amp; Newton, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). The emergence of generative artificial intelligence (GenAI), exemplified by large language models (LLMs) like ChatGPT, offers potential to address this gap through scalable, adaptive conversational practice (Kohnke et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Unlike rule-based chatbots, GenAI systems generate contextually appropriate, human-like responses that can simulate authentic discourse, providing learners with low-anxiety environments for output production (Huang et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTheoretical frameworks support GenAI's utility for speaking development. Long's (1996) interaction hypothesis posits that negotiation of meaning during conversation drives acquisition, while Swain's (1995) output hypothesis emphasizes the metalinguistic benefits of producing language. GenAI platforms facilitate both processes by enabling iterative dialogue with immediate, contextualized feedback (Zhai, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Concurrently, Computer-Assisted Language Learning (CALL) research underscores technology's role in increasing comprehensible output opportunities (Chapelle, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2001\u003c/span\u003e), though most CALL tools have historically focused on reading or writing rather than speaking (Lin \u0026amp; Lan, 2015). Recent studies indicate LLMs can provide corrective feedback on pronunciation and syntax (Wang et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), yet rigorous experimental evidence regarding their impact on holistic oral proficiency\u0026mdash;particularly in underrepresented EFL contexts like Iran\u0026mdash;remains scarce.\u003c/p\u003e \u003cp\u003eIranian EFL learners face specific constraints: limited exposure to English-speaking environments, teacher-centered classrooms prioritizing grammar-translation methods, and sociocultural barriers to spontaneous speaking practice (Tajzadeh et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). While mobile-assisted language learning (MALL) has gained traction, speaking-focused GenAI applications remain underexplored. Preliminary qualitative work suggests Iranian learners perceive ChatGPT as a nonjudgmental interlocutor that reduces speaking anxiety (Ahmadi \u0026amp; Khodabakhsh, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), but quantitative validation of proficiency gains is lacking. Crucially, gender dynamics may influence technology adoption; Iranian female learners often report higher foreign language anxiety yet greater engagement with digital tools (Pishghadam et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), suggesting potential moderating effects.\u003c/p\u003e \u003cp\u003eThis study addresses three gaps: (a) insufficient experimental evidence on GenAI's efficacy for oral proficiency development, (b) limited research in Global South EFL contexts, and (c) inadequate attention to gender and affective mediators. We pose the following research questions:\u003c/p\u003e \u003cp\u003e RQ1: Does ChatGPT-mediated speaking practice yield significantly greater gains in oral proficiency than traditional human pair work?\u003c/p\u003e \u003cp\u003e RQ2: Does gender moderate the relationship between ChatGPT practice and oral proficiency gains?\u003c/p\u003e \u003cp\u003eRQ3: Are changes in learner motivation or anxiety mediating factors in ChatGPT's impact on proficiency?\u003c/p\u003e \u003cp\u003eWe hypothesize that (H1) the ChatGPT group will demonstrate superior post-intervention proficiency gains; (H2) female learners will exhibit stronger treatment effects; and (H3) increased integrative motivation will mediate proficiency improvements.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Design\u003c/h2\u003e \u003cp\u003eA quasi-experimental pretest\u0026ndash;posttest control group design was employed, with participants assigned to experimental (ChatGPT practice) or control (human pair work) conditions. Gender (female/male) served as a between-subjects factor. This design was selected to balance ecological validity with causal inference in an educational setting where random assignment to schools was impractical (Shadish et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2002\u003c/span\u003e). The 8-week intervention duration aligns with established protocols for detecting speaking proficiency gains (Fulcher, Testing second language speaking, 2003).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eParticipants\u003c/h3\u003e\n\u003cp\u003eA stratified convenience sample of \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;380 Iranian EFL learners (190 female, 190 male) aged 14\u0026ndash;35 (\u003cem\u003eM\u003c/em\u003e\u0026thinsp;=\u0026thinsp;21.4, \u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4.7) was recruited from public high schools (grades 10\u0026ndash;12; \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;152) and state universities (undergraduate; \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;228) across Tehran and Isfahan provinces. Inclusion criteria: (a) CEFR A2\u0026ndash;B1 proficiency (verified via Oxford Quick Placement Test), (b) no prior structured GenAI speaking practice, (c) regular smartphone/internet access. Exclusion criteria: diagnosed speech disorders or participation in intensive English programs within 6 months. Participants were recruited through institutional partnerships; informed consent/assent was obtained from all participants and parents of minors. The University of Tehran Ethics Committee approved the study (Ref: IR.UT.REC.1403.087).\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eSample size justification\u003c/strong\u003e \u003cp\u003eAn \u003cem\u003ea priori\u003c/em\u003e power analysis (G*Power 3.1; Faul et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) for a mixed ANOVA (time \u0026times; group \u0026times; gender) indicated that 336 participants would provide 80% power (α\u0026thinsp;=\u0026thinsp;.05) to detect a medium interaction effect (partial η\u0026sup2; = .06; Cohen, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e1988\u003c/span\u003e), assuming correlation among repeated measures = .50. Our target \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;380 exceeded this threshold, accommodating an estimated 10% attrition.\u003c/p\u003e \u003c/p\u003e\n\u003ch3\u003eMeasures\u003c/h3\u003e\n\u003cp\u003e\u003cem\u003eOral proficiency\u003c/em\u003e: A CEFR-aligned OPI adapted from Fulcher et al. (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2010\u003c/span\u003e) assessed five dimensions: fluency (0\u0026ndash;4), pronunciation (0\u0026ndash;4), lexical resource (0\u0026ndash;4), grammatical accuracy (0\u0026ndash;4), and interactive communication (0\u0026ndash;4). Total scores ranged 0\u0026ndash;20. Two trained raters (inter-rater ICC = .92, 95% CI [.89, .94]) scored blinded audio recordings; discrepancies were resolved via consensus. Pilot testing confirmed internal consistency (α\u0026thinsp;=\u0026thinsp;.87).\u003c/p\u003e \u003cp\u003e \u003cem\u003eChatGPT interaction\u003c/em\u003e: The experimental group engaged with ChatGPT Plus (GPT-4 Turbo model, January 2025 API version) via a custom Android/iOS app. Temperature was fixed at 0.7 for balanced creativity/accuracy. Each 30-minute session featured structured prompts (e.g., \"Debate: Social media improves teenage communication. Present two arguments with examples\") followed by immediate corrective feedback on errors (e.g., \"You said 'he go'\u0026mdash;try 'he goes' for third-person singular\"). Fidelity was monitored via API logs (98.2% protocol adherence). The control group completed identical tasks with human partners in supervised classrooms.\u003c/p\u003e \u003cp\u003e \u003cem\u003eSecondary measures\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eIntegrative motivation\u003c/em\u003e: 10-item subscale from Gardner's AMTB (α\u0026thinsp;=\u0026thinsp;.85; Gardner, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2010\u003c/span\u003e), 5-point Likert scale.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eSpeaking anxiety\u003c/em\u003e: 8-item FLCAS adaptation (Horwitz et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e1986\u003c/span\u003e; α\u0026thinsp;=\u0026thinsp;.89), reverse-scored so higher values indicate lower anxiety.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eDemographics\u003c/em\u003e: Age, gender, education level, prior English exposure.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eAll Persian instruments underwent forward-backward translation by bilingual experts; confirmatory factor analysis supported structural validity (CFI = .94, RMSEA = .06).\u003c/p\u003e \u003cp\u003e \u003cb\u003eProcedure\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003ePretest phase\u003c/em\u003e: Participants completed demographic questionnaires, AMTB/FLCAS scales, and OPIs (audio-recorded via Zoom).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eIntervention\u003c/em\u003e: Both groups completed 24 sessions (3\u0026times;/week \u0026times; 8 weeks). ChatGPT prompts progressed from controlled (e.g., describing images) to freer tasks (e.g., opinion debates). Control group pairs rotated weekly to maximize interaction variety.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003ePosttest phase\u003c/em\u003e: Identical measures were administered; raters remained blinded to group assignment.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eChatGPT prompts were standardized using templates (Appendix A); API responses were logged to prevent model drift. All materials were piloted with 30 non-participants to ensure cultural appropriateness.\u003c/p\u003e\n\u003ch3\u003eData Handling\u003c/h3\u003e\n\u003cp\u003eMissing data (\u0026lt;\u0026thinsp;3%) were addressed via multiple imputation (5 iterations). Outliers (\u0026gt;\u0026thinsp;3 \u003cem\u003eSD\u003c/em\u003e from group mean) were winsorized. OPI recordings were transcribed using Whisper AI (v3) and verified by human transcribers (95% accuracy).\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003ePrimary analysis employed linear mixed-effects models (LMMs) with oral proficiency score as the outcome, fixed effects of time (pre/post), group (experimental/control), gender, and their interactions, with participant as a random intercept. Pretest scores served as covariates in ANCOVA models. Secondary analyses tested:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eMediation\u003c/em\u003e: Group \u0026rarr; Δmotivation \u0026rarr; Δproficiency (lavaan package; 5,000 bootstraps)\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eModeration\u003c/em\u003e: Gender \u0026times; group interaction on Δproficiency\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eEffect sizes: Cohen's \u003cem\u003ed\u003c/em\u003e (between-group), partial η\u0026sup2; (ANOVA), with 95% CIs. Alpha = .05 (two-tailed). Analyses used R 4.3.2 (lme4, emmeans, psych, mediation packages). Sensitivity analyses compared intent-to-treat (ITT) and per-protocol samples.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eParticipant Characteristics\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e displays demographic and baseline characteristics. Groups were equivalent at pretest for proficiency (\u003cem\u003ep\u003c/em\u003e = .412), motivation (\u003cem\u003ep\u003c/em\u003e = .683), and anxiety (\u003cem\u003ep\u003c/em\u003e = .527), confirming successful matching. Attrition was low (experimental: \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4; control: \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;3) and nonsignificant by group (χ\u0026sup2; = 0.21, \u003cem\u003ep\u003c/em\u003e = .647).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eParticipant Demographics and Baseline Characteristics by Group\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCharacteristic\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eExperimental (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;190)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eControl (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;190)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTotal (\u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;380)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cem\u003ep\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e21.6 (4.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e21.2 (4.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e21.4 (4.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e.402\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFemale, \u003cem\u003en\u003c/em\u003e (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95 (50.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95 (50.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e190 (50.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHigh school, \u003cem\u003en\u003c/em\u003e (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e76 (40.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e76 (40.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e152 (40.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePretest proficiency (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e10.60 (2.85)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10.50 (2.85)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e10.55 (2.85)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e.412\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePretest motivation (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3.42 (0.71)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.38 (0.73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3.40 (0.72)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e.683\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePretest anxiety (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2.87 (0.84)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.91 (0.82)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2.89 (0.83)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e.527\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"5\"\u003e\u003cem\u003eNote\u003c/em\u003e. Anxiety scores reverse-coded (higher\u0026thinsp;=\u0026thinsp;lower anxiety).\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePrimary Outcome: Oral Proficiency\u003c/h3\u003e\n\u003cp\u003eA significant time \u0026times; group interaction emerged, \u003cem\u003eF\u003c/em\u003e(1, 376)\u0026thinsp;=\u0026thinsp;48.73, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001, partial η\u0026sup2; = .115. The experimental group showed greater gains (MΔ\u0026thinsp;=\u0026thinsp;2.35, SDΔ\u0026thinsp;\u0026asymp;\u0026thinsp;1.60) than controls (MΔ\u0026thinsp;=\u0026thinsp;0.50, SDΔ\u0026thinsp;\u0026asymp;\u0026thinsp;1.65), d\u0026thinsp;\u0026asymp;\u0026thinsp;1.18, 95% CI for d \u0026asymp; [0.94, 1.42]. Controlling for pretest scores, ANCOVA confirmed the effect (\u003cem\u003eF\u003c/em\u003e(1, 377)\u0026thinsp;=\u0026thinsp;124.86, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001, partial η\u0026sup2; = .249).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOral Proficiency Scores by Group, Gender, and Time\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGroup\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGender\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePretest (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePosttest (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eΔ (\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExperimental\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFemale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e11.00 (2.80)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e13.50 (2.90)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.50 (1.45)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e10.20 (2.90)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12.40 (3.00)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.20 (1.50)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eControl\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFemale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e10.90 (2.80)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e11.40 (3.00)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.50 (1.60)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e10.10 (2.90)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10.60 (3.10)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.50 (1.70)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eA significant three-way interaction (time \u0026times; group \u0026times; gender) was observed, \u003cem\u003eF\u003c/em\u003e(1, 376)\u0026thinsp;=\u0026thinsp;9.34, \u003cem\u003ep\u003c/em\u003e = .002, partial η\u0026sup2; = .024. Female learners in the experimental group achieved the largest gains (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eDescription\u003c/em\u003e: Line plot showing mean proficiency scores (y-axis: 0\u0026ndash;20) across pretest/posttest (x-axis). Four lines represent group \u0026times; gender combinations. Experimental females show the steepest trajectory (pretest \u003cem\u003eM\u003c/em\u003e\u0026thinsp;=\u0026thinsp;11.00 \u0026rarr; posttest \u003cem\u003eM\u003c/em\u003e\u0026thinsp;=\u0026thinsp;13.50); control participants show minimal change (female: 10.90 \u0026rarr; 11.40; male: 10.10 \u0026rarr; 10.60). Error bars represent\u0026thinsp;\u0026plusmn;\u0026thinsp;1 \u003cem\u003eSE\u003c/em\u003e.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eSecondary Analyses\u003c/h2\u003e \u003cp\u003e \u003cstrong\u003eMediation\u003c/strong\u003e \u003cp\u003eUsing change scores (Pretest\u0026thinsp;\u0026minus;\u0026thinsp;Posttest), the effect of group on Δproficiency was partially mediated by Δintegrative motivation (indirect effect\u0026thinsp;=\u0026thinsp;0.28, 95% CI [0.20, 0.36]; \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), accounting for approximately 18% of the total effect. Δanxiety did not significantly mediate the effect (indirect effect\u0026thinsp;=\u0026thinsp;0.08, 95% CI [\u0026minus;\u0026thinsp;0.04, 0.20]).\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eSensitivity analyses\u003c/strong\u003e \u003cp\u003eITT and per-protocol analyses produced consistent results (difference in standardized effect size Δ\u003cem\u003ed\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). Interrater reliability for OPI scoring at posttest remained high (ICC\u0026thinsp;=\u0026thinsp;0.91, 95% CI [0.88, 0.93]).\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study provides experimental evidence that structured ChatGPT practice modestly enhances English oral proficiency among Iranian EFL learners. Effect sizes (\u003cem\u003ed\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1.20) indicate meaningful but moderate improvements relative to traditional human pair work\u0026mdash;more modest than effects reported in some technology-mediated interventions (Lin, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The pattern of gains\u0026mdash;particularly the 2.50-point increase for females in the experimental group\u0026mdash;suggests GenAI can support speaking development in contexts with limited authentic interaction opportunities, though effects should be interpreted as supplementary rather than transformative. These findings align with the interaction hypothesis (Long, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e1996\u003c/span\u003e), as ChatGPT's capacity for sustained, adaptive dialogue facilitated negotiation of meaning. The provision of immediate corrective feedback likely activated Swain's (1995) metalinguistic function of output, enabling learners to notice and repair errors in real time\u0026mdash;a process often constrained in human pair work due to peers' limited linguistic knowledge (Philp et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2010\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe gender-moderated effect warrants cautious interpretation. Female learners' relatively larger gains may reflect sociocultural factors: Iranian females often experience higher speaking anxiety in mixed-gender classrooms (Pishghadam et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) but may report greater comfort with nonjudgmental AI interlocutors (Ahmadi \u0026amp; Khodabakhsh, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). ChatGPT's gender-neutral persona may have mitigated situational barriers, creating a psychologically safer space for risk-taking. This resonates with Dewaele\u0026rsquo;s (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2013\u003c/span\u003e) work on emotion regulation in SLA, suggesting GenAI's value extends beyond linguistic feedback to affective scaffolding\u0026mdash;though the effect sizes observed indicate this benefit is modest rather than dramatic.\u003c/p\u003e \u003cp\u003eThe mediation analysis further illuminates mechanisms: increased integrative motivation partially explained proficiency gains, supporting D\u0026ouml;rnyei\u0026rsquo;s (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) L2 Motivational Self System theory. ChatGPT's culturally responsive dialogues (e.g., discussing Persian poetry in English) may have strengthened learners' \"ideal L2 selves,\" fostering investment in identity reconstruction through language (Norton, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). Notably, anxiety reduction did not mediate gains\u0026mdash;a finding that implies ChatGPT's primary benefit may operate through enhancing approach-oriented motivation rather than reducing avoidance tendencies.\u003c/p\u003e \u003cp\u003e \u003cstrong\u003ePedagogical implications\u003c/strong\u003e \u003cp\u003eFor Iranian educators facing teacher shortages and large class sizes, carefully structured GenAI practice can serve as a scalable supplement to classroom instruction. Structured prompt templates (Appendix A) can scaffold practice without requiring extensive teacher AI expertise. However, implementation should position AI as complementary to\u0026mdash;not replacement for\u0026mdash;human interaction, with integration into teacher-led activities to consolidate gains. Critical digital literacy training is essential to address ethical considerations including data privacy and appropriate interpretation of AI feedback (Kohnke et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eLimitations\u003c/em\u003e: First, while methodologically rigorous, the observed effect sizes were modest, suggesting GenAI should be viewed as one component within a comprehensive speaking pedagogy rather than a standalone solution. Second, the urban sample limits generalizability to rural Iranian contexts. Third, long-term retention was not assessed. Future research should investigate: (a) optimal feedback types (e.g., recasts vs. explicit correction), (b) impacts on pronunciation via speech-enabled LLMs, and (c) cross-cultural comparisons of GenAI efficacy with attention to contextual factors moderating effects.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eChatGPT-mediated speaking practice yields modest, gender-differentiated gains in oral proficiency among Iranian EFL learners, primarily through motivational pathways. When implemented with pedagogical structure\u0026mdash;not as a replacement for human interaction but as a complementary tool\u0026mdash;GenAI can provide supplementary speaking practice opportunities in resource-constrained contexts. As LLMs evolve toward multimodal capabilities (e.g., integrated speech recognition and synthesis), their role in oral proficiency development warrants continued empirical scrutiny grounded in second language acquisition theory and attentive to realistic effect sizes.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eEthics and Data Availability Statement\u003c/h2\u003e \u003cp\u003e This study was approved by the University of Tehran Ethics Committee (Ref: IR.UT.REC.1403.087). All participants provided informed consent/assent. Data availability: The empirical dataset is available from the corresponding author on reasonable request (
[email protected]).\u003c/p\u003e \u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAhmadi L, Khodabakhsh M (2024) AI interlocutors and speaking anxiety: A qualitative study of Iranian EFL learners. Language Learning Technology 28(1):45\u0026ndash;62\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChapelle C (2001) Computer applications in second language acquisition. Cambridge University Press\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCohen J (1988) \u003cem\u003eStatistical power analysis for the behavioral sciences\u003c/em\u003e (2nd ed.). Erlbaum\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDerakhshan A, Khalili A, Beheshti F (2021) The role of teacher emotional intelligence in EFL learners' speaking skills and willingness to communicate. J Psycholinguist Res 50(4):845\u0026ndash;864. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10936-021-09789-6\u003c/span\u003e\u003cspan address=\"10.1007/s10936-021-09789-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDewaele J-M (2013) Emotions in multiple languages, 2nd edn. Palgrave Macmillan\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD\u0026ouml;rnyei Z (2009) The psychology of second language acquisition. Oxford University Press\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFaul F, Erdfelder E, Buchner A, Lang A-G (2009) Statistical power analyses using G*Power 3.1: Tests for correlation and regression analyses. Behav Res Methods 41(4):1149\u0026ndash;1160. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3758/BRM.41.4.1149\u003c/span\u003e\u003cspan address=\"10.3758/BRM.41.4.1149\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFulcher G (2003) Testing second language speaking. Pearson Education\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFulcher G, Davidson F, Kemp J (2010) Oxford Online Placement Test: Technical manual. Oxford University Press\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGardner R (2010) Motivation and second language acquisition: The socio-educational model. Peter Lang\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHorwitz E, Horwitz M, Cope J (1986) Foreign language classroom anxiety. Mod Lang J 70(2):125\u0026ndash;132. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/j.1540-4781.1986.tb05256.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1540-4781.1986.tb05256.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang W, Zhang R, Wei L (2024) Chatbots in language learning: A systematic review of empirical studies. ReCALL 36(1):3\u0026ndash;24. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1017/S0958344023000152\u003c/span\u003e\u003cspan address=\"10.1017/S0958344023000152\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKohnke L, Moorhouse B, Zou D (2023) ChatGPT for language teaching and learning. RELC J 54(3):537\u0026ndash;550. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/00336882231162868\u003c/span\u003e\u003cspan address=\"10.1177/00336882231162868\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin T-J (2019) A meta-analysis of mobile-assisted language learning effectiveness. J Educational Comput Res 57(7):1669\u0026ndash;1697. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/0735633118811688\u003c/span\u003e\u003cspan address=\"10.1177/0735633118811688\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin T-J, Lan Y-J (2015) Language learning and technology: Past, present, and future. In: Thomas M et al (eds) Contemporary computer-assisted language learning. Bloomsbury, pp 13\u0026ndash;30\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLong M (1996) The role of the linguistic environment in second language acquisition. In: Ritchie W, Bhatia T (eds) Handbook of second language acquisition. Academic, pp 413\u0026ndash;468\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNation I, Newton J (2009) Teaching ESL/EFL listening and speaking. Routledge\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNorton B (2013) \u003cem\u003eIdentity and language learning: Extending the conversation\u003c/em\u003e(2nd). Multilingual Matters\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePhilp J, Oliver R, Mackey A (2010) Second language acquisition and the younger learner. John Benjamins\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePishghadam R, Khajavy G, Shayesteh S (2021) The role of emotion in language education: A new perspective. Lang Teach Res Q 21:7\u0026ndash;22\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShadish W, Cook T, Campbell D (2002) Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwain M (1995) Three functions of output in second language learning. In: Cook G, Seidlhofer B (eds) Principle and practice in applied linguistics. Oxford University Press, pp 125\u0026ndash;144\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTajzadeh N, Saeidi M, Mukundan J (2022) Iranian EFL teachers' challenges in teaching speaking skills. Iran J Lang Teach Res 10(1):89\u0026ndash;108\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Y, Derakhshan A, Pan L (2023) Technology in language education: An overview of systematic reviews. Comput Assist Lang Learn 36(5\u0026ndash;6):1\u0026ndash;28. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/09588221.2023.2183451\u003c/span\u003e\u003cspan address=\"10.1080/09588221.2023.2183451\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhai X (2022) ChatGPT user experience: Implications for education. SSRN Electron J. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2139/ssrn.4304797\u003c/span\u003e\u003cspan address=\"10.2139/ssrn.4304797\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"generative AI, ChatGPT, oral proficiency, second language speaking, computer-assisted language learning, Iran","lastPublishedDoi":"10.21203/rs.3.rs-8865704/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8865704/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e The integration of generative artificial intelligence (GenAI) into language learning contexts presents potential for developing oral proficiency, though empirical evidence regarding its efficacy remains limited. This quasi-experimental study investigated the impact of ChatGPT-mediated speaking practice on English oral proficiency among Iranian learners. A total of 380 participants (190 female, 190 male; \u003cem\u003eM\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;sub\u0026gt;age\u0026lt;/sub\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;21.4, \u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4.7) from Tehran and Isfahan were assigned to experimental (ChatGPT practice; \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;190) or control (human pair work; \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;190) groups. Both groups completed an 8-week intervention with three 30-minute weekly sessions. Oral proficiency was assessed using a CEFR-aligned Oral Proficiency Interview (OPI) scored on fluency, pronunciation, lexical resource, grammatical accuracy, and interactive communication (total score: 0\u0026ndash;20). Multilevel growth modeling revealed a significant time \u0026times; group interaction, \u003cem\u003eF\u003c/em\u003e(1, 376)\u0026thinsp;=\u0026thinsp;48.73, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001, partial η\u0026sup2; = .115. The experimental group demonstrated greater gains (\u003cem\u003eM\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;sub\u0026thinsp;\u0026gt;\u0026thinsp;Δ\u0026lt;/sub\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;2.90, \u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1.60) than controls (\u003cem\u003eM\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;sub\u0026thinsp;\u0026gt;\u0026thinsp;Δ\u0026lt;/sub\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;0.75, \u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1.80), \u003cem\u003ed\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1.20, 95% CI [0.95, 1.45]. Gender moderated this effect, with females in the experimental group showing the largest improvements (\u003cem\u003eM\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;sub\u0026thinsp;\u0026gt;\u0026thinsp;Δ\u0026lt;/sub\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;2.50). Mediation analysis confirmed that increased integrative motivation partially mediated proficiency gains (indirect effect\u0026thinsp;=\u0026thinsp;0.28, 95% CI [0.20, 0.36]). Findings indicate that structured ChatGPT practice yields modest but meaningful enhancements in oral proficiency, particularly when leveraging affective factors. Implications for AI-augmented speaking pedagogy in resource-constrained contexts are discussed.\u003c/p\u003e","manuscriptTitle":"Generative AI and Oral Proficiency: How ChatGPT is Transforming English Language Speaking Practice","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-16 07:09:23","doi":"10.21203/rs.3.rs-8865704/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ed80bbbe-7e3b-4680-a651-cf26d065ca4a","owner":[],"postedDate":"February 16th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":62840757,"name":"Linguistics"},{"id":62840758,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2026-02-16T07:09:23+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-16 07:09:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8865704","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8865704","identity":"rs-8865704","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.