Human-Centered GenAI Feedback Design in Higher Education: A Multisite Experiment on Direct, Reflective, and Hybrid Approaches to Scientific Argumentation

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Generative artificial intelligence (GenAI) is increasingly used for feedback in higher education, yet evidence remains limited on how alternative human–AI feedback designs shape learning processes and durable outcomes. This study addresses that gap through a multisite, cluster-randomized, longitudinal field experiment comparing four feedback designs in introductory university science courses: peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and a hybrid design combining self-evaluation, peer feedback, and GenAI critique. The analytic sample comprised 1,176 first-year undergraduate students from 48 course sections across four universities and three science domains. Primary and secondary outcomes were four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer; feedback uptake and self-regulated learning during revision were modeled as process mediators. Direct GenAI-supported feedback improved immediate argument-quality gain relative to peer feedback, but reflective and hybrid designs produced stronger outcomes on feedback uptake, self-regulated learning, conceptual learning, and delayed AI-free transfer. The hybrid condition yielded the highest adjusted mean for immediate argument-quality gain, whereas both reflective and hybrid conditions outperformed direct GenAI-supported feedback on delayed AI-free transfer. Multilevel mediation analyses indicated that feedback uptake and self-regulated learning partially explained these advantages. By combining four feedback designs, process mediation, and delayed AI-free transfer in a multisite field experiment, the study shows that the educational value of GenAI in higher education depends less on AI access per se than on whether feedback environments preserve student agency, evaluative judgment, and ownership during revision.
Full text 272,810 characters · extracted from preprint-html · click to expand
Human-Centered GenAI Feedback Design in Higher Education: A Multisite Experiment on Direct, Reflective, and Hybrid Approaches to Scientific Argumentation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Human-Centered GenAI Feedback Design in Higher Education: A Multisite Experiment on Direct, Reflective, and Hybrid Approaches to Scientific Argumentation Huseyin ATES This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9396658/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Generative artificial intelligence (GenAI) is increasingly used for feedback in higher education, yet evidence remains limited on how alternative human–AI feedback designs shape learning processes and durable outcomes. This study addresses that gap through a multisite, cluster-randomized, longitudinal field experiment comparing four feedback designs in introductory university science courses: peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and a hybrid design combining self-evaluation, peer feedback, and GenAI critique. The analytic sample comprised 1,176 first-year undergraduate students from 48 course sections across four universities and three science domains. Primary and secondary outcomes were four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer; feedback uptake and self-regulated learning during revision were modeled as process mediators. Direct GenAI-supported feedback improved immediate argument-quality gain relative to peer feedback, but reflective and hybrid designs produced stronger outcomes on feedback uptake, self-regulated learning, conceptual learning, and delayed AI-free transfer. The hybrid condition yielded the highest adjusted mean for immediate argument-quality gain, whereas both reflective and hybrid conditions outperformed direct GenAI-supported feedback on delayed AI-free transfer. Multilevel mediation analyses indicated that feedback uptake and self-regulated learning partially explained these advantages. By combining four feedback designs, process mediation, and delayed AI-free transfer in a multisite field experiment, the study shows that the educational value of GenAI in higher education depends less on AI access per se than on whether feedback environments preserve student agency, evaluative judgment, and ownership during revision. Generative artificial intelligence feedback design higher education digital learning scientific argumentation self-regulated learning epistemic agency Figures Figure 1 Figure 2 Figure 3 Figure 4 1. Introduction Generative artificial intelligence (GenAI) is increasingly embedded in higher education, including feedback, assessment, and learning support, yet its educational significance remains unsettled (Wang & Fan, 2025; Xia et al., 2024). Existing evidence suggests that GenAI can support learning performance, learning perceptions, and higher-order thinking, but these benefits vary substantially depending on the role assigned to the tool, the instructional design, and the characteristics of the task (Wang & Fan, 2025). In feedback contexts, this variation is especially important. Feedback is not educationally valuable simply because comments are delivered; rather, it becomes productive when learners interpret feedback, compare it against criteria, judge its relevance, and use it to improve subsequent work (Carless & Boud, 2018; Hattie & Timperley, 2007; Nicol & Macfarlane-Dick, 2006). Thus, the central issue in GenAI-supported feedback is not automation alone, but whether feedback design sustains active learner engagement, feedback uptake, and evaluative agency (Buckingham Shum et al., 2023; Darvishi et al., 2024). This question is particularly consequential in higher education science learning. Scientific argumentation requires students not only to produce acceptable answers, but also to formulate claims, justify them with evidence and reasoning, consider alternative explanations, and revise disciplinary explanations in response to critique (Driver et al., 2000; Osborne et al., 2004). Compared with more generic writing tasks, scientific argumentation places stronger demands on disciplinary judgment because revision depends not only on coherence or fluency, but also on the adequacy of evidence, the quality of reasoning, and the treatment of uncertainty or alternative explanations (Osborne et al., 2004; Zhu et al., 2020). In this context, GenAI-supported feedback is pedagogically promising but also potentially risky: it may help students improve written products and, when designed dialogically, support deeper reasoning, yet it may also reduce the need for students to make their own evaluative judgments if the feedback process encourages passive uptake (Darvishi et al., 2024; Tang & Putra, 2025; Zhu et al., 2020). Recent reviews suggest that GenAI is expanding rapidly in higher education, yet three substantive gaps remain in the feedback literature (Wang & Fan, 2025; Xia et al., 2024). First, much of the literature on AI-supported feedback has focused on general writing tasks or broad comparisons among feedback sources, whereas disciplinary reasoning in university science courses has received less sustained attention (Banihashem et al., 2024; Tang & Putra, 2025; Wang & Fan, 2025). Second, many studies compare a single AI condition with a single non-AI condition, leaving underexplored how alternative feedback designs may distribute evaluative work differently across learners, peers, and GenAI, especially when self-evaluation, peer input, and GenAI critique are differently combined or sequenced (Banihashem et al., 2024; Chang et al., 2026; Darvishi et al., 2024). Third, immediate product improvement is often examined without simultaneously investigating feedback uptake, self-regulated learning, and later performance without AI support, making it difficult to determine whether observed gains reflect durable learning, stronger revision processes, or temporary performance support (Wang & Fan, 2025; Xia et al., 2024; Zhu et al., 2020). Beyond the immediate question of whether GenAI can generate feedback efficiently, the broader issue for higher education digital learning is how feedback environments should be designed so that speed, scalability, and accessibility do not come at the expense of student agency, evaluative judgment, and ownership of learning (Buckingham Shum et al., 2023; Darvishi et al., 2024; Xia et al., 2024). Under these conditions, the key pedagogical question is not simply whether students have access to feedback, but whether feedback design positions them as active interpreters, comparators, and decision-makers in relation to critique—learners who can exercise evaluative judgment and generate internal feedback rather than merely implement external suggestions (Carless & Boud, 2018; Tai et al., 2018; Molloy et al., 2020; Nicol, 2021). Scientific argumentation provides a critical case for examining this problem because it offers a demanding test of whether GenAI functions as a support for reasoning or as a shortcut around reasoning. Against this background, the present study addresses this problem through a multisite, cluster-randomized, longitudinal field experiment in introductory university science courses. The study compares four feedback designs—peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and a hybrid design combining self-evaluation, peer feedback, and GenAI critique—and examines not only immediate four-dimension argument-quality gain, but also conceptual learning, delayed AI-free transfer, and the process pathways of feedback uptake and self-regulated learning. In doing so, the study contributes to higher education digital learning research by shifting attention from access to AI-generated feedback toward the design conditions under which GenAI can support meaningful revision, disciplinary reasoning, and learner agency. More specifically, by combining four feedback designs, process mediation, and delayed AI-free transfer in a multisite field experiment, the study asks how AI can be integrated into feedback processes in ways that amplify rather than displace students’ evaluative work.. 2. Literature review and theoretical framework 2.1. Feedback as uptake, judgment, and revision The present study is grounded in a process-oriented understanding of feedback. In this perspective, feedback is not reducible to information transmitted from a source to a learner. Rather, it becomes educationally meaningful when learners compare their current performance against criteria or standards, interpret the significance of comments, and use those judgments to guide subsequent action (Carless & Boud, 2018; Nicol & Macfarlane-Dick, 2006; Nicol, 2021). This view shifts attention from feedback provision to feedback use and from the delivery of comments to the learner processes through which comments are evaluated and transformed into revision. This learner-centered view is closely aligned with evaluative judgment, defined as the capacity to make decisions about the quality of work, whether one’s own or that of others (Tai et al., 2018). From this perspective, the educational value of feedback depends not simply on the accuracy or quantity of comments, but on whether students are positioned to judge quality, compare alternatives, and decide how revision should proceed. Carless and Boud (2018) conceptualized this broader learner capability through feedback literacy, which includes appreciating feedback, making judgments, managing affective responses, and taking action. Similarly, Molloy et al. (2020) argued that feedback-literate learning environments should support students’ active role in interpreting and using feedback rather than treating them as passive recipients of commentary. This distinction is especially important in technology-mediated feedback. In higher education, feedback increasingly occurs through digitally mediated environments, including learning-management systems, platform-based review processes, and AI-enabled tools that make rapid and scalable commentary more readily available (Buckingham Shum et al., 2023). Yet access to larger volumes of feedback does not automatically generate learning. Nicol (2021) argued that feedback becomes most productive when it stimulates internal feedback, that is, when learners engage in comparison processes between current performance, criteria, alternatives, and desired standards. For the purposes of the present study, feedback literacy refers to a broader learner capability or disposition, whereas feedback uptake refers to a task-specific process: how learners attend to, interpret, select, and implement critique during a particular episode of revision. Accordingly, the present study treats feedback uptake as the more proximal process through which external comments are translated into revision decisions and learning-relevant action. 2.2. Self-regulated learning and epistemic agency in AI-mediated feedback Feedback uptake is closely related to self-regulated learning (SRL). SRL theories describe learning as an active process in which students plan, monitor, evaluate, and adapt their cognition, motivation, and behavior in response to task demands (Nicol & Macfarlane-Dick, 2006; Panadero, 2017). In feedback contexts, SRL is central because students must decide how to interpret comments, whether to trust them, which revisions to prioritize, and how to integrate them into future performance. Feedback, in this sense, does not merely inform learning; it becomes one of the mechanisms through which learners regulate learning. This relationship can be understood more precisely through evaluative judgment and internal feedback. When students compare external comments with task criteria, prior understanding, and their own initial diagnosis of the work, they generate internal feedback that can inform planning, monitoring, and strategic adjustment during revision (Nicol, 2021; Tai et al., 2018). In the present study, feedback uptake is therefore treated as a proximal revision process that helps shape subsequent self-regulation. When students actively attend to, interpret, and select among feedback messages, they create the conditions under which revision becomes strategic rather than merely reactive. AI-mediated feedback intensifies this issue. GenAI may support SRL when it helps learners identify weaknesses, clarify evaluative criteria, compare alternatives, and plan revision moves in ways that remain cognitively engaging for the learner (Buckingham Shum et al., 2023). However, it may also undermine SRL if it encourages students to accept suggestions uncritically or to outsource difficult evaluative judgments. Darvishi et al. (2024) showed that AI assistance in peer-feedback environments may increase dependence on the tool, raising questions about student agency in AI-supported learning contexts. This concern can be framed more precisely as a question of epistemic agency: whether students remain responsible for judging the adequacy of claims, evidence, and reasoning, or whether these judgments are displaced by AI-generated suggestions. In this sense, epistemic agency is closely related to ownership, evaluative judgment, and responsibility for disciplinary adequacy (Tai et al., 2018; Nicol, 2021). Research on self-assessment provides an important implication for this problem. Meta-analytic findings indicate that self-assessment can positively influence self-regulated learning and self-efficacy because it requires learners to compare their work against explicit criteria and to generate internal judgments before receiving or acting on external feedback (Panadero et al., 2017). This point is especially relevant for AI-supported revision: if students are first required to articulate their own diagnosis of strengths, weaknesses, or uncertainties, they may be better positioned to engage with GenAI critique as a resource for comparison rather than as a substitute for judgment. From this perspective, feedback designs that require prior self-evaluation before exposure to AI-generated commentary may offer more supportive conditions for SRL and epistemic agency than designs in which AI feedback is presented immediately and directly. 2.3. Scientific argumentation as disciplinary and epistemic work Scientific argumentation provides a particularly demanding context for studying AI-supported feedback because it is epistemically different from generic academic writing. In introductory university science courses, students must do more than produce coherent text. They are expected to formulate claims, coordinate claims with evidence and reasoning, address limitations or counterpositions, and revise explanations in light of critique (Driver et al., 2000; Osborne et al., 2004). These demands make scientific argumentation a domain in which revision quality depends on disciplinary judgment rather than only on stylistic improvement. This distinction has direct implications for feedback. Feedback on scientific argumentation must address the epistemic adequacy of an argument: whether evidence is relevant and sufficient, whether reasoning justifies the claim, whether uncertainty or alternative explanations are acknowledged, and whether revisions improve explanatory power. Zhu et al. (2020) demonstrated that, in the formative assessment of scientific argument writing, revisions were positively associated with learning gains and that contextualized feedback was more effective than generic feedback. Emerging work also suggests that GenAI may support disciplinary engagement when it is designed in dialogic rather than answer-providing ways. Tang and Putra (2025), for example, found that a customized chatbot could support perspective-taking, reasoning, and argumentation in science learning when it functioned as a dialogic partner rather than an authoritative answer source. For this reason, scientific argumentation serves as a critical case for testing whether GenAI-supported feedback amplifies or displaces students’ evaluative work in higher education. If student judgment and ownership can be preserved in this context, the implications are likely to extend beyond science to other higher education settings that require evidence-based explanation, justification, and revision. 2.4. Why direct, reflective, and hybrid feedback designs may differ The core theoretical claim of the present study is that alternative feedback designs should not be expected to generate identical learning processes because they distribute evaluative work differently across learners, peers, and AI-supported tools. From a human-centered digital learning perspective, the central issue is not AI access alone, but whether the feedback design preserves learners’ responsibility for interpreting critique, making judgments, and deciding how revision should proceed. In this sense, feedback design matters because it shapes whether students engage in evaluative judgment and internal feedback processes or whether those processes are prematurely displaced by external suggestions (Tai et al., 2018; Nicol, 2021). In direct GenAI-supported feedback, students receive AI-generated critique without first externalizing their own diagnosis of the task. Such a design may improve efficiency and support immediate four-dimension argument-quality gain by providing rapid, criterion-referenced suggestions. At the same time, because evaluative work is front-loaded onto the system rather than the learner, direct GenAI-supported feedback may offer weaker conditions for internal comparison, self-generated judgment, and strategic revision. For this reason, it may also provide weaker conditions for feedback uptake, self-regulated revision, and delayed AI-free transfer than designs that require more active learner judgment. In reflective GenAI-supported feedback, students first evaluate their own work against explicit criteria or identify areas of uncertainty before engaging with AI-generated critique. Its theoretical advantage is that it requires learners to articulate an initial judgment and then compare that judgment with external feedback. This sequence is important because evaluative judgment develops through opportunities to make and test quality-related decisions rather than merely receive answers about what to change (Tai et al., 2018). It is also consistent with Nicol’s (2021) argument that learning is strengthened when feedback triggers internal comparison processes. Because students are not positioned as passive recipients of commentary, reflective GenAI-supported feedback may be expected to strengthen feedback uptake, planning, monitoring, and strategy adjustment during revision, thereby supporting not only four-dimension argument-quality gain but also more durable conceptual learning and AI-independent transfer. In hybrid feedback, self-evaluation, peer feedback, and GenAI critique are combined within the same instructional sequence. This design is theoretically important because these feedback sources appear to offer partially complementary rather than identical affordances. Self-evaluation foregrounds internal criteria and initial judgment, peer feedback highlights reader-oriented comprehensibility and ambiguity, and GenAI critique may identify rubric-aligned omissions, weaknesses, or alternative revision moves (Banihashem et al., 2024; Chang et al., 2026). Taken together, these sources may create the strongest comparison ecology for active judgment and revision: students are required not only to receive multiple forms of critique, but also to compare them, select among them, and justify how revision decisions should proceed. In this sense, hybrid feedback may provide especially strong support for internal feedback processes, evaluative judgment, and ownership of revision, although it also entails greater coordination demands. Based on this framework, the present study proposes a process model in which feedback design is expected to affect learning outcomes both directly and indirectly. Direct effects remain plausible because alternative designs may immediately alter the quality, timing, and diversity of critique available to learners. Indirect effects are expected to operate through feedback uptake, through self-regulated learning during revision, and through the serial pathway from uptake to self-regulated revision. On this account, the more agentic designs, especially reflective and hybrid feedback, should be better positioned to support stronger four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer because they require learners to compare, select, and justify revision decisions rather than merely implement external suggestions. The proposed process model is presented in Figure 1. The model proposes that feedback design affects four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer both directly and indirectly through feedback uptake and self-regulated learning during revision. Feedback uptake is modeled as a proximal process through which learners interpret, select, and implement critique, while self-regulated learning captures the planning, monitoring, and strategy adjustment that follow during revision. The more agentic feedback designs are expected to strengthen the serial pathway from feedback uptake to self-regulated revision. The model is situated in introductory university science courses and scientific argumentation tasks. 3. Research questions and hypotheses For clarity, the primary immediate outcome is defined as four-dimension argument-quality gain, that is, improvement from first draft to revised submission on the four content dimensions shared across draft stages. The strategic quality of revision is treated as a separate revised-draft indicator and is used in a supplementary interpretive role rather than as the primary gain outcome. RQ1. How do peer feedback, direct GenAI-supported feedback, reflective GenAI-supported feedback, and hybrid feedback differ in their effects on students’ immediate four-dimension argument-quality gain in scientific argumentation tasks? RQ2. How do these four feedback conditions differ in their effects on students’ conceptual learning and delayed AI-free transfer in introductory university science courses? RQ3. How do these feedback conditions differ in terms of students’ feedback uptake and self-regulated learning during the revision process? RQ4. To what extent do feedback uptake and self-regulated learning mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer? H1a. Reflective GenAI-supported feedback and hybrid feedback will yield higher levels of feedback uptake than direct GenAI-supported feedback because they require learners to make evaluative judgments before and during engagement with AI-generated critique. H1b. No directional hypothesis is specified between reflective GenAI-supported feedback and hybrid feedback, although the hybrid design may provide additional advantages through multi-source comparison. H2a. Reflective GenAI-supported feedback and hybrid feedback will yield higher levels of self-regulated learning during revision than direct GenAI-supported feedback. H2b. No directional hypothesis is specified between reflective GenAI-supported feedback and hybrid feedback for self-regulated learning during revision. H3. Reflective GenAI-supported feedback and hybrid feedback are expected to support stronger conceptual learning and delayed AI-free transfer than direct GenAI-supported feedback because they preserve learners’ evaluative judgment and promote more agentic revision. H4. Direct GenAI-supported feedback may be expected to improve immediate four-dimension argument-quality gain relative to peer feedback by providing rapid, criterion-referenced critique; however, reflective and hybrid feedback designs are expected to yield comparable or stronger gains when deeper feedback uptake and self-regulated revision are supported. H5a. Feedback uptake will mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer. H5b. Self-regulated learning during revision will mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer. H5c. Feedback uptake and self-regulated learning will jointly mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer through a serial pathway from uptake to self-regulated revision. 4. Method 4.1. Design This study employed a multisite, cluster-randomized, longitudinal field experiment to examine how different feedback designs shaped students’ scientific argumentation and learning in introductory university science courses. Cluster randomization was used because the intervention was implemented at the level of intact course sections, thereby reducing contamination across conditions during peer interaction and revision activities. The unit of randomization was the course section, whereas the primary unit of analysis was the student, with clustering handled through multilevel modeling. Four conditions were compared: peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and hybrid feedback combining self-evaluation, peer feedback, and GenAI critique. The study followed a pre-test, post-test, and delayed post-test structure. The primary immediate outcome was four-dimension argument-quality gain in scientific argumentation tasks, defined as improvement from the initial draft to the revised submission on the four content dimensions shared across draft stages. A separate revised-draft indicator captured the strategic quality of revision and was used as a concurrent supplementary indicator rather than as the primary gain outcome. Secondary outcomes were conceptual learning and delayed AI-free transfer. Two process variables—feedback uptake and self-regulated learning during revision—were specified a priori as mediators linking feedback design to learning outcomes. The design, primary outcomes, focal contrasts, and analysis plan were preregistered on the Open Science Framework (OSF) prior to post-test analysis. The preregistered focal contrasts emphasized comparisons of the more agentic feedback designs with direct GenAI-supported feedback while also retaining planned comparisons with the peer-feedback condition. An a priori power analysis for a four-arm cluster-randomized design indicated that, assuming an intraclass correlation of .05, an average cluster size of 26 students, a two-sided alpha of .05, and a minimum detectable standardized effect of d = .25 for the planned between-condition contrasts, 48 sections would provide statistical power of .82. Power calculations were conducted in Optimal Design 3.01 under a balanced allocation scenario. On this basis, the target sample was set at approximately 1,200–1,300 students across four institutions. 4.2. Setting and participants The study was conducted across four universities offering compulsory first-year undergraduate courses in biology, chemistry, and physics. These courses were selected because they required students to complete written tasks involving explanation, evidence use, and revision of scientific arguments. Across the participating universities, the intervention was embedded in regular coursework rather than delivered as an external experimental activity. Background data collected at baseline included demographic and academic variables relevant to the analyses, together with prior GenAI experience and feedback-related measures; descriptive characteristics of the analytic sample are reported in the Results section. The initial sample comprised 1,248 first-year undergraduate students from 48 course sections. The section-level distribution was balanced across disciplines: 16 biology sections (n = 416), 16 chemistry sections (n = 420), and 16 physics sections (n = 412). Mean section size at enrollment was 26.0 students (SD = 2.7, range = 21–31). Sections were randomly assigned to one of the four feedback conditions, resulting in 12 sections per condition. Randomization was conducted by a researcher not involved in instruction using a computer-generated random-number sequence. To preserve balance, randomization was blocked by institution and discipline, and section meeting time (morning/afternoon) was used as a secondary balancing variable. Allocation was completed after course enrollment closed and before the Week 1 pre-test. Participation in the instructional activities formed part of normal course requirements; participation in the research component, including questionnaire, draft, interview, and log-data analysis, was voluntary and based on informed consent. Inclusion in the analytic dataset required consent for research use of student work, survey responses, and platform data. Students who did not consent completed the same class activities but were excluded from the analytic dataset. Additional exclusions related to course withdrawal and missing outcome data are reported in the Results section. Because research consent was voluntary while instructional participation was compulsory, the analytic sample may reflect a degree of consent-based selection. To reduce institution-level heterogeneity, the study used a common intervention schedule, shared task-design procedures, and a common analytic rubric across sites. 4.3. Instructional tasks and rubric The intervention comprised three scientific argumentation cycles distributed across a 10-week period. In each cycle, students produced an initial draft, received condition-specific feedback, revised their text, and submitted a final version. Tasks were designed to require explicit coordination of claim, evidence, and reasoning and were aligned with the disciplinary content of each course. Two task formats were used: laboratory-based interpretive writing tasks and short socio-scientific argumentation tasks. In the former, students explained data patterns, justified conclusions, and addressed sources of uncertainty; in the latter, they defended a position using disciplinary concepts and evidence. To support comparability across sites and disciplines, tasks were co-developed by participating instructors through a shared design protocol and reviewed in cross-institution calibration meetings focusing on epistemic demands, expected difficulty, and rubric fit across biology, chemistry, and physics. Each cycle required an initial draft of approximately 350–500 words and a revised submission of 400–600 words. Students had 72 hours to complete the first draft, 48 hours to review assigned feedback, and 72 hours to submit the revision. In the hybrid condition, the revision memo was limited to 100–150 words. Students’ work was evaluated using an analytic scientific argumentation framework adapted from prior research on scientific argumentation and automated feedback in scientific writing (e.g., Osborne et al., 2004; Zhu et al., 2020). The framework comprised five dimensions: claim quality, evidence relevance and sufficiency, coherence of reasoning, treatment of limitations or alternative explanations, and strategic quality of revision. The first four dimensions were scored for both initial and revised drafts and were used to estimate four-dimension argument-quality gain. The fifth dimension, strategic quality of revision, was scored only for revised submissions on the basis of paired comparison with the initial draft and was retained as a supplementary indicator of revised performance. The same four common content dimensions were applied across intervention and transfer tasks, with topic-specific descriptors added where necessary to preserve disciplinary appropriateness. Prior to data collection, raters completed two calibration sessions using anchor responses from all three disciplinary domains. Two trained raters independently scored all first and final drafts used in outcome assessment, and discrepancies greater than one point on any rubric dimension were resolved through discussion. Full rubric details and anchor descriptors are provided in the supplementary materials. 4.4. Experimental conditions An overview of the four intervention conditions is provided in Table 1. 4.4.1. Peer feedback only Students in the peer-feedback condition exchanged drafts with two peers and provided rubric-based comments through a structured template covering claim, evidence, reasoning, and one revision suggestion. No AI support was available in this condition. 4.4.2. Direct GenAI-supported feedback Students in the direct GenAI condition uploaded their draft to the course-integrated GenAI interface immediately after drafting. The system generated criterion-referenced feedback on claim quality, evidence use, reasoning, limitations, and possible revision moves. Students could consult this feedback during revision, but they did not complete a structured self-evaluation before viewing the AI output. 4.4.3. Reflective GenAI-supported feedback Students in the reflective GenAI condition completed a brief structured self-evaluation aligned with the rubric before receiving AI-generated feedback. They identified strengths, weaknesses, and uncertainties in their draft, so that engagement with AI critique followed an initial evaluative judgment. 4.4.4. Hybrid feedback Students in the hybrid condition completed a fixed sequence of self-evaluation, peer feedback, GenAI critique, and revision memo. They first evaluated their own draft, then received comments from two peers, consulted the GenAI system for additional critique and revision suggestions, and finally completed a short memo explaining which suggestions they accepted, rejected, or modified. This condition was designed to maximize comparison across self-, peer-, and AI-based feedback. Because the reflective and hybrid conditions incorporated additional guided evaluative activity, observed differences may reflect not only feedback-source configuration and sequencing but also differences in reflective workload. Access to condition-specific feedback resources was monitored through platform records, and potential contamination through non-assigned resources is addressed under implementation fidelity. Table 1. Overview of the four feedback conditions Condition Feedback sources Sequence Student role AI role Intended pedagogical function Peer feedback only Peer feedback Draft → peer feedback from two peers → revision → final submission Reviewer and reviser None To support evaluative judgment through peer review and revision without AI assistance Direct GenAI-supported feedback GenAI feedback Draft → GenAI-generated feedback → revision → final submission Reviser Provides criterion-referenced critique and revision suggestions To provide immediate, scalable, rubric-aligned feedback with minimal preparatory judgment by the learner Reflective GenAI-supported feedback Self-evaluation + GenAI feedback Draft → structured self-evaluation → GenAI-generated feedback → revision → final submission Self-evaluator and reviser Provides critique after the learner articulates an initial judgment To strengthen feedback uptake by requiring prior evaluative judgment before AI-supported revision Hybrid feedback Self-evaluation + peer feedback + GenAI feedback Draft → self-evaluation → peer feedback from two peers → GenAI-generated feedback → revision memo → final submission Self-evaluator, reviewer, and reviser Provides additional critique and revision suggestions after self- and peer-based judgment To maximize active comparison across internal judgment, peer judgment, and AI-supported critique Note. All conditions used the same scientific argumentation tasks and common analytic rubric. The intervention differed in the source and sequencing of feedback and, to some extent, in the amount of guided evaluative activity required of students. 4.5. GenAI system and prompting protocol The AI-supported conditions used a secure, institutionally hosted large language model interface based on OpenAI’s GPT-5 model, accessed through the universities’ approved application environment and integrated into the learning-management platform so that interaction logs could be captured while maintaining institutional data-governance requirements. Students were instructed not to enter personally identifying information into the system. The GenAI tool was configured to provide criterion-referenced feedback rather than replacement text. Prompts directed the system to identify strengths and weaknesses in the argument, flag missing or weak evidence, question unsupported claims, suggest revision moves, and pose follow-up questions intended to deepen reasoning, while explicitly prohibiting full-answer rewriting. To ensure consistency across institutions, all AI-supported conditions used the same prompt architecture, rubric anchors, interface settings, and fixed model parameters across all three intervention cycles. Instructors were not permitted to alter prompts during the intervention period. To reduce variability and prompt gaming, students were limited to a single feedback generation per draft and could not iteratively regenerate responses within the same cycle, although they could review the generated feedback during revision. Chat history was reset between cycles to avoid cross-task carryover. Full prompt materials, model settings, interface screenshots, and implementation rules are provided in the supplementary materials. 4.6. Procedure The study ran for 10 instructional weeks (Figure 2). In Week 1, students completed a pre-test assessing domain-specific prior knowledge and baseline scientific argumentation, together with a background survey covering prior achievement, prior GenAI experience, and baseline feedback literacy. In Week 2, all students received a common orientation to scientific argumentation, the analytic rubric, and constructive feedback. Students in AI-supported conditions also received additional training on ethical and effective use of the GenAI system. Although this additional training was necessary for procedural consistency, it may also have increased preparation time relative to the non-AI condition. During Weeks 3–8, students completed three argumentation cycles, each involving an initial draft, a condition-specific feedback sequence, and a revised submission. The two task formats were distributed across the intervention in alignment with course content, while sections within the same course followed the same task sequence. Interaction logs, timestamps, revision traces, and revision memos were collected throughout this period. In Week 9, students completed a post-test assessing conceptual learning and a novel transfer task without AI access. In Week 10, they completed a delayed AI-free transfer task and a post-intervention questionnaire assessing self-regulated learning during revision and perceived engagement with feedback. The delayed transfer task was conceptually related to, but topically distinct from, the Week 9 transfer task. After the intervention, a purposive subsample of students and instructors participated in semi-structured interviews. 4.7. Measures An overview of all measures, data sources, and timing is presented in Table 2. 4.7.1. Four-dimension argument-quality gain The primary immediate outcome was four-dimension argument-quality gain, operationalized as the change in students’ common-content argument score from initial draft to revised submission. Gain was calculated on the four dimensions shared across draft stages—claim quality, evidence relevance and sufficiency, coherence of reasoning, and treatment of limitations or alternative explanations—each scored on a four-point scale, yielding directly comparable common-content subtotals ranging from 4 to 16 at both stages. Higher values indicated greater improvement in scientific argumentation quality from draft to revision. To distinguish change in common argument quality from revised-product quality, the study also retained a separate revised-draft indicator: strategic quality of revision. This fifth dimension was scored only for revised submissions and was interpreted as a concurrent supplementary indicator rather than as part of pre/post gain estimation. Revision depth was coded separately on a 3-point scale (1 = predominantly surface-level, 2 = mixed, 3 = predominantly substantive) and averaged across the three intervention cycles to support interpretation of whether observed gains reflected substantive rather than predominantly surface-level changes. 4.7.2. Conceptual learning Conceptual learning was measured using discipline-specific pre/post assessments co-developed by instructors across the four institutions. Each assessment included 12–15 items targeting the core concepts underlying the argumentation tasks and combined selected-response and short constructed-response formats. Raw scores were converted to percentages within discipline. Because the discipline-specific forms were not identical, post-test scores were standardized within discipline before pooled modeling, allowing conceptual-learning outcomes to be compared on a common relative metric across biology, chemistry, and physics. Internal consistency at post-test was acceptable across disciplines (McDonald’s ω = .79 in biology, .82 in chemistry, and .80 in physics). 4.7.3. Delayed AI-free transfer Delayed AI-free transfer was assessed through a novel scientific argumentation task completed individually without access to the course-integrated GenAI interface. Although the task differed in topic from both the intervention activities and the Week 9 transfer assessment, it required the same underlying epistemic practices captured by the common-content scoring framework: formulating a claim, selecting relevant evidence, linking evidence to reasoning, and addressing uncertainty or competing explanations. Novelty was ensured by using prompts that had not appeared in course activities, practice materials, or post-test instruments. Scoring was based on the same four common content dimensions used across the intervention tasks, with topic-specific descriptors added where necessary to preserve disciplinary appropriateness. Because the delayed transfer task was completed as a single independent performance rather than as a paired draft sequence, strategic quality of revision was not applicable. In the present study, AI-free task completion meant that students had no access to the study’s GenAI tool and completed the task independently in a supervised classroom setting without internet-enabled devices. 4.7.4. Feedback uptake Feedback uptake was treated as a multi-indicator construct capturing the extent to which students attended to, interpreted, and incorporated critique during revision. It drew on four sources of evidence: students’ post-revision rationales, coder-rated alignment between received feedback and implemented revisions, time spent reviewing assigned feedback resources in platform logs, and the proportion of substantive revisions traceable to feedback content. In the hybrid condition, the revision memo served as the post-revision rationale; in the other conditions, the same indicator was captured through a shorter structured rationale submitted with the final draft. Each indicator was averaged across the three intervention cycles and z-standardized across the full analytic sample. Indicators were then examined using a one-factor measurement model, with standardized loadings ranging from .58 to .81. The final uptake index used in the multilevel models was computed as the equal-weight mean of the four z-standardized indicators, with higher values indicating stronger uptake. 4.7.5. Feedback literacy Baseline feedback literacy was assessed using the full 22-item Student Feedback Literacy Instrument (SFLI), a higher-education measure comprising the dimensions of feedback attitudes and feedback practices (Weidlich et al., 2025). The instrument was selected because it was specifically developed for use with higher-education students and builds on earlier initial scale-validation work conducted in the same context (Woitt et al., 2025). Students responded on a 5-point Likert scale ranging from 1 (strongly disagree) to 5 (strongly agree), and scores were calculated as the mean across items, with higher values indicating greater feedback literacy. Internal consistency in the present sample was high (McDonald’s ω = .90). 4.7.6. Self-regulated learning during revision Self-regulated learning (SRL) during revision was measured using a 12-item task-specific scale adapted from the higher-education self-regulated learning tradition represented by the Motivated Strategies for Learning Questionnaire (MSLQ; Pintrich et al., 1991) and from subsequent work emphasizing planning, monitoring, and strategic adaptation as core features of SRL (Panadero, 2017). The scale sampled four revision-relevant domains—planning, monitoring, strategy adjustment, and evaluation—with three items per domain. Responses were recorded on a 7-point Likert scale ranging from 1 (not at all true of me) to 7 (very true of me). After reverse-coding where necessary, the self-report component was scored as the mean across items, with higher scores indicating stronger self-regulated learning during revision. Internal consistency was satisfactory (McDonald’s ω = .88). To complement self-report data, the study also incorporated digital trace indicators derived from the learning-management system, including revisiting feedback, spacing of revision activity, repeated consultation of the rubric, and sequencing of draft-comparison actions. These traces were treated as behavioral indicators of revision regulation rather than as stand-alone proxies for self-regulation. A trace-based SRL subindex was computed by z-standardizing the four indicators and averaging them with equal weight. For the primary multilevel analyses, the overall SRL construct was represented by the mean of the z-standardized self-report score and the z-standardized trace-based subindex, with higher values indicating stronger self-regulated learning during revision. 4.7.7. Covariates The models adjusted for prior knowledge, baseline scientific argumentation, prior achievement, prior GenAI experience, gender, and science domain. Prior achievement was operationalized as each student’s standardized university-entry score obtained from institutional records; because entry metrics differed across institutions, scores were standardized within institution before inclusion in the pooled analyses. Institutional variation was handled through institution fixed effects in the mixed-effects models. Exploratory moderation analyses additionally examined whether baseline feedback literacy conditioned the effects of the feedback designs. Table 2. Overview of measures Variable Construct type Operational definition Instrument/source Time of measurement Scoring/format Prior knowledge Covariate Domain-specific understanding of core course concepts before the intervention Instructor-developed pre-test Week 1 Standardized test score Baseline scientific argumentation Covariate Initial quality of written scientific argumentation before the intervention Scientific argumentation task scored on the four common content dimensions of the analytic framework Week 1 Four-dimension common-content subtotal (range = 4–16) Prior achievement Covariate Academic achievement prior to university study Institutional entry records Week 1 Institution-standardized university-entry score Prior GenAI experience Covariate Self-reported familiarity and prior use of GenAI tools for study-related tasks Background questionnaire Week 1 5-point self-report scale Feedback literacy Covariate / exploratory moderator Students’ readiness to interpret, value, and act on feedback 22-item SFLI Week 1 Mean scale score (1–5) Four-dimension argument-quality gain Primary outcome Improvement from first draft to final draft on the four content dimensions shared across draft stages Analytic scientific argumentation framework; strategic quality of revision and revision-depth coding used as supplementary indicators Weeks 3–8 Change in common-content subtotal (final minus initial; four shared dimensions only) Conceptual learning Secondary outcome Understanding of disciplinary concepts addressed in the intervention tasks Discipline-specific post-test Week 9 Discipline-standardized post-test score Delayed AI-free transfer Secondary outcome Ability to construct a scientific argument on a novel task without AI support Novel scientific argumentation task scored on the four common content dimensions of the analytic framework Week 10 Four-dimension common-content rubric score Feedback uptake Process variable / mediator Extent to which feedback was attended to, interpreted, and incorporated into revision Post-revision rationales, coded alignment between feedback and revisions, platform log data, traceable substantive revisions Weeks 3–8 Equal-weight mean of four z-standardized indicators averaged across cycles Self-regulated learning during revision Process variable / mediator Planning, monitoring, strategic adjustment, and evaluation during revision 12-item task-specific SRL scale plus digital trace indicators Weeks 3–10 Mean of z-standardized self-report score and z-standardized trace-based subindex Implementation fidelity Quality-control indicator Degree to which the intervention was delivered as intended across sites and sections Instructor logs, platform records, manual audits Throughout intervention Fidelity checklist / audit record Note . For pooled analyses across biology, chemistry, and physics, discipline-specific conceptual-learning scores were standardized within discipline before modeling. Institutional variation in prior achievement was addressed by standardizing university-entry scores within institution before inclusion in the pooled analyses. Baseline scientific argumentation, four-dimension argument-quality gain, and delayed AI-free transfer were all based on the four common content dimensions shared across draft stages; the strategic quality of revision was scored only for revised intervention submissions and was interpreted as a concurrent supplementary indicator rather than as part of pre/post gain estimation. 4.8. Fidelity of implementation Implementation fidelity was monitored through four procedures: use of a common intervention schedule, rubric, and instructional materials across sites; standardized prompt architecture and fixed system settings in the AI-supported conditions; platform logs verifying that students accessed only the feedback resources assigned to their condition; and manual audits of a random 15% sample of student cases across sites. Fidelity was assessed using a predefined checklist covering sequence integrity, timing compliance, access restrictions, and completion of condition-specific activities. Audits were conducted by trained research assistants who were not involved in instruction, and 20% of audited cases were double-coded, with high agreement on checklist items (Cohen’s κ = .89). Detailed audit criteria and procedures are reported in the supplementary materials. 4.9. Data analysis Because students were nested within course sections and only four institutions participated, the primary quantitative analyses were estimated using mixed-effects models with students at Level 1, sections at Level 2, and institution entered as a fixed effect. This approach was preferred to specifying institutions as a random third level because the number of higher-level units was too small to support stable estimation of institution-level random effects. Analyses were conducted in R 4.3.2. For RQ1 and RQ2, separate mixed-effects models were estimated for four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer. Experimental condition was treated as the principal fixed effect, section was specified as a random intercept, and baseline scores and background variables were included as covariates. Planned pairwise contrasts were estimated from the fitted models using estimated marginal means and Holm-adjusted p values. In line with the preregistered comparison strategy, outcomes collected across the three intervention cycles were averaged at the student level before modeling. The primary immediate outcome was modeled as a gain score, consistent with both the instructional design and the preregistered outcome definition. For RQ3, between-condition differences in feedback uptake and self-regulated learning during revision were examined using the same fixed- and random-effects structure. For RQ4, 2-1-1 multilevel mediation models were estimated with feedback condition at the section level and both mediators and outcomes at the student level. Indirect effects were estimated separately through feedback uptake, through self-regulated learning during revision, and through the serial pathway from uptake to self-regulated revision. All primary analyses followed an intention-to-treat approach. Missing data were handled using multiple imputation by chained equations with 20 imputations under a missing-at-random assumption. For mediation analyses, indirect effects were estimated using Monte Carlo confidence intervals based on 20,000 simulations. To assess robustness, the primary models were also re-estimated using per-protocol analyses, models excluding low-fidelity sections, and an alternative missing-data specification based on full-information maximum likelihood. Additional analytic details and robustness results are reported in the supplementary materials. 4.10. Qualitative analysis To complement the quantitative analyses, semi-structured interviews were conducted with a purposive subsample of 32 students and 12 instructors. The student subsample was balanced across conditions and approximately balanced across disciplines, while instructor interviews captured site-level variation in feasibility and implementation. Interviews focused on how participants interpreted feedback, made revision decisions, and experienced ownership in AI-supported revision. All interviews were audio-recorded, transcribed verbatim, and analyzed using reflexive thematic analysis following Braun and Clarke (2006). The qualitative component was interpretive rather than reliability-seeking. Themes were developed iteratively through memoing and team discussion, with reflexive attention to how researchers’ assumptions informed coding. No inter-coder reliability coefficient was calculated, in keeping with the reflexive thematic analysis approach. The qualitative findings were used in an explanatory role to clarify how participants experienced the feedback designs and to interpret the mechanisms underlying the quantitative patterns. 4.11. Ethical considerations Ethical approval was obtained from the institutional review boards of the participating universities before data collection. Participation in the research component was voluntary, and students were informed that non-participation would not affect course grades. All data were de-identified prior to analysis, and identifiable platform data were stored on institutionally approved encrypted servers with access restricted to the research team in accordance with local ethics procedures. Participants were informed of their right to withdraw their research data up to the point at which the dataset had been de-identified and merged for analysis. Because the intervention involved GenAI tools, students received explicit guidance on ethical AI use, data-entry limits, and the requirement that submitted work remain their own. Monitoring of AI use was limited to interactions within the study platform. To reduce the risk of unequal access across conditions, all groups received the same core instruction in scientific argumentation, rubric use, and feedback principles, and students in the non-AI condition were provided with access to the GenAI orientation materials and a demonstration version of the interface after the intervention. 5. Results 5.1. Preliminary analyses The final analytic sample comprised 1,176 students nested within 48 course sections across four institutions. Of the 1,248 students initially enrolled, 72 were excluded because of non-consent for research use of their data, course withdrawal before the first intervention cycle, or absence from both the post-test and delayed transfer sessions. Final sample sizes were balanced across conditions, and attrition did not differ significantly by condition, χ²(3) = 0.33, p = .954. Missing data on the retained sample ranged from 0.6% to 3.4% across variables and were handled using multiple imputation with 20 imputations. Baseline equivalence analyses indicated that the four conditions did not differ significantly on age, gender distribution, prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, or prior GenAI experience (all p ≥ .472), suggesting that the randomization procedure produced broadly comparable groups prior to the intervention. Descriptive statistics and baseline-equivalence tests are reported in Table 3. The internal consistency of the self-report measures was satisfactory, and inter-rater reliability for scientific argumentation scoring and revision-depth coding was high. Implementation fidelity was also high across sites: based on instructor logs, platform records, and checklist-based audits of 176 randomly selected student cases, 95.8% of applicable instructional steps were delivered as intended, with high agreement on fidelity checklist items (Cohen’s κ = .89). Cross-condition contamination was minimal. Taken together, these findings support the adequacy of the dataset and the integrity of the intervention. Section-level variance components and intraclass correlation coefficients are reported in Supplementary Table S1. Table 3. Sample characteristics and baseline equivalence across conditions Variable Peer feedback only (n = 293) Direct GenAI-supported feedback (n = 295) Reflective GenAI-supported feedback (n = 294) Hybrid feedback (n = 294) Total (N = 1,176) Test statistic p Age, years, M (SD) 18.92 (1.31) 18.88 (1.27) 18.95 (1.29) 18.90 (1.34) 18.91 (1.30) F(3, 1172) = 0.14 .936 Women, n (%) 167 (57.0) 171 (58.0) 166 (56.5) 169 (57.5) 673 (57.2) χ²(3) = 0.18 .981 Prior knowledge (0–100), M (SD) 58.63 (10.84) 59.11 (10.52) 60.02 (10.41) 59.38 (10.76) 59.29 (10.63) F(3, 1172) = 0.84 .472 Baseline scientific argumentation (4–16), M (SD) 10.32 (2.08) 10.51 (2.03) 10.64 (2.10) 10.57 (2.07) 10.51 (2.07) F(3, 1172) = 1.10 .349 Prior achievement, M (SD) -0.05 (0.98) 0.03 (1.01) 0.04 (0.96) -0.01 (1.00) 0.00 (0.99) F(3, 1172) = 0.71 .548 Feedback literacy (1–5), M (SD) 3.41 (0.52) 3.44 (0.50) 3.47 (0.49) 3.45 (0.51) 3.44 (0.51) F(3, 1172) = 0.67 .571 Prior GenAI experience (1–5), M (SD) 2.78 (0.89) 2.93 (0.92) 2.88 (0.90) 2.99 (0.91) 2.89 (0.91) F(3, 1172) = 1.42 .236 Note. Baseline-equivalence tests are based on the final analytic sample. Prior achievement scores were standardized within institution before pooled comparison. No significant between-condition differences were observed at baseline. Baseline scientific argumentation scores reflect the four common content dimensions of the analytic framework and therefore ranged from 4 to 16 rather than including the revised-draft strategic-quality-of-revision indicator. 5.2. Effects of feedback condition on four-dimension argument-quality gain To address RQ1, a linear mixed-effects model was estimated using students’ mean four-dimension argument-quality gain across the three intervention cycles as the dependent variable, with section included as a random intercept and institution entered as a fixed effect. Feedback condition significantly predicted four-dimension argument-quality gain, F (3, 43.8) = 18.76, p < .001. As shown in Figure 3 and Table 4, argument-quality gain was highest in the hybrid condition, followed by the reflective GenAI-supported feedback condition, the direct GenAI-supported feedback condition, and the peer-feedback condition. Relative to peer feedback, direct GenAI-supported feedback improved four-dimension argument-quality gain, and both reflective and hybrid feedback produced still larger gains. Both reflective and hybrid conditions also significantly outperformed direct GenAI-supported feedback, whereas the difference between the hybrid and reflective conditions was not statistically significant. Supplementary analyses of coder-rated revision depth showed a parallel pattern. Relative to peer feedback, revision depth was higher in the direct GenAI-supported feedback condition and higher still in the reflective and hybrid conditions; both reflective and hybrid feedback also exceeded direct GenAI-supported feedback, with no significant difference between the reflective and hybrid conditions. This supplementary pattern indicates that the advantages of the more agentic conditions were accompanied by more substantive rather than merely surface-level revision. Detailed estimates are reported in Table 4 and Supplementary Table S2. Table 4. Multilevel model predicting four-dimension argument-quality gain Panel A. Estimated marginal means by condition Condition Adjusted M SE 95% CI Peer feedback only 2.84 0.15 [2.55, 3.13] Direct GenAI-supported feedback 3.41 0.14 [3.14, 3.68] Reflective GenAI-supported feedback 4.07 0.14 [3.80, 4.34] Hybrid feedback 4.31 0.13 [4.05, 4.57] Panel B. Pairwise contrasts Contrast b SE p Cohen’s d 95% CI for b Direct GenAI – Peer 0.57 0.18 .002 0.25 [0.22, 0.92] Reflective GenAI – Peer 1.23 0.18 < .001 0.54 [0.88, 1.58] Hybrid – Peer 1.47 0.18 < .001 0.64 [1.12, 1.82] Reflective GenAI – Direct GenAI 0.66 0.17 < .001 0.29 [0.33, 0.99] Hybrid – Direct GenAI 0.90 0.17 < .001 0.40 [0.57, 1.23] Hybrid – Reflective GenAI 0.24 0.16 .138 0.11 [-0.08, 0.56] Panel C. Overall model summary Effect Test statistic p Feedback condition F(3, 43.8) = 18.76 < .001 Note. The model was adjusted for prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, prior GenAI experience, gender, science domain, and institution fixed effects, with section included as a random intercept. Four-dimension argument-quality gain scores were computed as change scores from initial draft to revised submission on the four common content dimensions shared across draft stages and were averaged across the three intervention cycles. Higher scores indicate greater improvement in common-content scientific argument quality from draft to revision. The separate strategic-quality-of-revision indicator was not included in pre/post gain estimation. Holm-adjusted p values are reported for pairwise contrasts. 5.3. Effects on conceptual learning and delayed AI-free transfer To address RQ2, separate linear mixed-effects models were estimated for conceptual learning and delayed AI-free transfer, with section included as a random intercept and institution entered as a fixed effect. Feedback condition significantly predicted both conceptual learning, F (3, 43.5) = 7.62, p < .001, and delayed AI-free transfer, F (3, 43.9) = 12.03, p < .001. For conceptual learning, the pattern favored the more agentic conditions. Direct GenAI-supported feedback did not differ significantly from peer feedback, whereas reflective and hybrid feedback outperformed peer feedback. Hybrid feedback also significantly outperformed direct GenAI-supported feedback, while the reflective–direct contrast was not statistically significant after adjustment, and the hybrid–reflective contrast was also not significant. For delayed AI-free transfer, the pattern was clearer. Direct GenAI-supported feedback did not outperform peer feedback, whereas both reflective and hybrid feedback significantly outperformed both peer feedback and direct GenAI-supported feedback. The difference between reflective and hybrid feedback was not statistically significant. Taken together, these findings indicate that the more agentic feedback designs were associated with stronger distal learning outcomes, particularly AI-independent transfer. Detailed estimates and pairwise contrasts are reported in Table 5. Table 5. Mixed-effects models predicting conceptual learning and delayed AI-free transfer . Panel A. Conceptual learning Condition Adjusted M SE 95% CI Peer feedback only -0.08 0.06 [-0.20, 0.04] Direct GenAI-supported feedback 0.01 0.06 [-0.11, 0.13] Reflective GenAI-supported feedback 0.19 0.06 [0.07, 0.31] Hybrid feedback 0.26 0.06 [0.14, 0.38] Contrast b SE p Cohen’s d 95% CI for b Direct GenAI – Peer 0.09 0.09 .314 0.08 [-0.09, 0.27] Reflective GenAI – Peer 0.27 0.10 .006 0.25 [0.08, 0.46] Hybrid – Peer 0.34 0.09 < .001 0.32 [0.16, 0.52] Reflective GenAI – Direct GenAI 0.18 0.10 .072 0.17 [-0.02, 0.38] Hybrid – Direct GenAI 0.25 0.10 .014 0.23 [0.05, 0.45] Hybrid – Reflective GenAI 0.07 0.09 .438 0.07 [-0.11, 0.25] Effect Test statistic p Feedback condition F(3, 43.5) = 7.62 < .001 Panel B. Delayed AI-free transfer Condition Adjusted M SE 95% CI Peer feedback only 11.82 0.21 [11.41, 12.23] Direct GenAI-supported feedback 11.46 0.20 [11.07, 11.85] Reflective GenAI-supported feedback 12.61 0.19 [12.24, 12.98] Hybrid feedback 12.88 0.20 [12.49, 13.27] Contrast b SE p Cohen’s d 95% CI for b Direct GenAI – Peer -0.36 0.25 .149 -0.12 [-0.85, 0.13] Reflective GenAI – Peer 0.79 0.26 .003 0.27 [0.28, 1.30] Hybrid – Peer 1.06 0.25 < .001 0.36 [0.57, 1.55] Reflective GenAI – Direct GenAI 1.15 0.26 < .001 0.39 [0.64, 1.66] Hybrid – Direct GenAI 1.42 0.27 < .001 0.48 [0.89, 1.95] Hybrid – Reflective GenAI 0.27 0.28 .327 0.09 [-0.28, 0.82] Effect Test statistic p Feedback condition F(3, 43.9) = 12.03 < .001 Note. Models were adjusted for prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, prior GenAI experience, gender, science domain, and institution fixed effects, with section included as a random intercept. Conceptual learning scores were standardized within discipline before pooled analysis. Higher delayed-transfer scores indicate stronger AI-independent scientific argumentation on the novel task. Holm-adjusted p values are reported for pairwise contrasts. 5.4. Effects on feedback uptake and self-regulated learning To address RQ3, separate linear mixed-effects models were estimated for feedback uptake and self-regulated learning during revision, with section included as a random intercept and institution entered as a fixed effect. Feedback condition significantly predicted both feedback uptake, F (3, 43.6) = 16.28, p < .001, and self-regulated learning, F (3, 43.2) = 11.91, p < .001. For both process variables, the same overall pattern emerged. Direct GenAI-supported feedback did not differ significantly from peer feedback, whereas reflective and hybrid feedback yielded significantly higher feedback uptake and self-regulated learning than both peer feedback and direct GenAI-supported feedback. No significant difference emerged between reflective and hybrid feedback on either process variable. These results are consistent with H1a and H2a, and with the nondirectional expectations stated in H1b and H2b. Full estimates are presented in Table 6. Table 6. Mixed-effects models predicting feedback uptake and self-regulated learning during revision Panel A. Feedback uptake Condition Adjusted M SE 95% CI Peer feedback only -0.06 0.05 [-0.16, 0.04] Direct GenAI-supported feedback -0.02 0.05 [-0.12, 0.08] Reflective GenAI-supported feedback 0.24 0.05 [0.14, 0.34] Hybrid feedback 0.31 0.05 [0.21, 0.41] Contrast b SE p Direct GenAI – Peer 0.04 0.07 .563 Reflective GenAI – Peer 0.30 0.07 < .001 Hybrid – Peer 0.37 0.07 < .001 Reflective GenAI – Direct GenAI 0.26 0.07 < .001 Hybrid – Direct GenAI 0.33 0.07 < .001 Hybrid – Reflective GenAI 0.07 0.07 .314 Effect Test statistic p Feedback condition F(3, 43.6) = 16.28 < .001 Panel B. Self-regulated learning during revision Condition Adjusted M SE 95% CI Peer feedback only -0.03 0.05 [-0.13, 0.07] Direct GenAI-supported feedback -0.09 0.05 [-0.19, 0.01] Reflective GenAI-supported feedback 0.17 0.05 [0.07, 0.27] Hybrid feedback 0.26 0.05 [0.16, 0.36] Contrast b SE p Direct GenAI – Peer -0.06 0.07 .401 Reflective GenAI – Peer 0.20 0.08 .015 Hybrid – Peer 0.29 0.08 < .001 Reflective GenAI – Direct GenAI 0.26 0.08 .002 Hybrid – Direct GenAI 0.35 0.08 < .001 Hybrid – Reflective GenAI 0.09 0.08 .226 Effect Test statistic p Feedback condition F(3, 43.2) = 11.91 < .001 Note. Both process variables were standardized composite indices. Models were adjusted for prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, prior GenAI experience, gender, science domain, and institution fixed effects, with section included as a random intercept. Higher scores indicate stronger engagement with feedback and greater self-regulation during revision. Holm-adjusted p values are reported for pairwise contrasts. 5.5. Mediation analyses To address RQ4, 2-1-1 multilevel mediation models were estimated to examine whether feedback uptake and self-regulated learning during revision mediated the relationship between feedback condition and four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer. Because the theoretically focal contrasts concerned the more agentic feedback designs, the mediation analyses compared the reflective and hybrid conditions primarily against direct GenAI-supported feedback. Standardized coefficients are reported in Table 7. Across outcomes, the mediation results showed that the advantages of reflective and hybrid feedback relative to direct GenAI-supported feedback were explained in part by stronger feedback uptake and self-regulated learning. For four-dimension argument-quality gain, both reflective and hybrid feedback showed significant indirect effects through feedback uptake, self-regulated learning, and the serial pathway from uptake to self-regulated revision. For conceptual learning, the strongest indirect effects operated through self-regulated learning and the serial pathway, whereas the pathway through feedback uptake alone did not reach significance. For delayed AI-free transfer, all three indirect pathways were statistically significant for both reflective and hybrid feedback. Taken together, these findings indicate that the more agentic feedback designs improved learning outcomes not only by changing the feedback resources available to students, but also by changing how students engaged with and acted on those resources. Detailed direct, indirect, and total effects are reported in Table 7. Table 7. Direct, indirect, and total effects of feedback condition on outcomes Contrast Outcome Direct effect, β [95% CI] Total indirect effect, β [95% CI] Via feedback uptake, β [95% CI] Via SRL, β [95% CI] Via uptake → SRL, β [95% CI] Total effect, β [95% CI] Reflective GenAI vs. Direct GenAI Four-dimension argument-quality gain 0.36 [0.11, 0.61] 0.30 [0.18, 0.44] 0.18 [0.08, 0.31] 0.07 [0.02, 0.14] 0.05 [0.01, 0.10] 0.66 [0.41, 0.91] Hybrid vs. Direct GenAI Four-dimension argument-quality gain 0.52 [0.27, 0.77] 0.38 [0.25, 0.52] 0.22 [0.11, 0.35] 0.09 [0.03, 0.16] 0.07 [0.02, 0.13] 0.90 [0.65, 1.15] Reflective GenAI vs. Direct GenAI Conceptual learning 0.05 [-0.04, 0.14] 0.13 [0.06, 0.21] 0.03 [-0.01, 0.08] 0.06 [0.02, 0.12] 0.04 [0.01, 0.08] 0.18 [0.08, 0.28] Hybrid vs. Direct GenAI Conceptual learning 0.08 [-0.02, 0.18] 0.17 [0.09, 0.25] 0.04 [-0.00, 0.09] 0.08 [0.03, 0.14] 0.05 [0.02, 0.10] 0.25 [0.15, 0.35] Reflective GenAI vs. Direct GenAI Delayed AI-free transfer 0.65 [0.30, 1.00] 0.50 [0.31, 0.72] 0.21 [0.09, 0.36] 0.18 [0.07, 0.32] 0.11 [0.04, 0.19] 1.15 [0.80, 1.50] Hybrid vs. Direct GenAI Delayed AI-free transfer 0.80 [0.44, 1.16] 0.62 [0.42, 0.84] 0.26 [0.12, 0.41] 0.22 [0.10, 0.36] 0.14 [0.06, 0.23] 1.42 [1.06, 1.78] Note. Standardized coefficients are reported for all focal contrasts. Indirect-effect confidence intervals were derived from Monte Carlo simulation; direct- and total-effect confidence intervals were derived from model-based standard errors. Confidence intervals that do not cross zero indicate statistically reliable effects. For the primary immediate outcome, mediation models used four-dimension argument-quality gain rather than the separate revised-draft strategic-quality-of-revision indicator, because gain estimation required directly comparable scoring across draft stages. 5.6. Qualitative findings on feedback use and epistemic agency The qualitative findings were used to clarify how students and instructors experienced the feedback designs and to interpret the mechanisms underlying the quantitative patterns. Although both student and instructor interviews informed the thematic analysis, the themes reported below focus primarily on student accounts of engaging with different feedback designs. Instructor interviews largely corroborated these patterns, particularly with respect to the trade-off between efficiency and ownership and the greater deliberation observed in the reflective and hybrid conditions. Representative quotations are provided in Table 8. The first theme, AI as a fast but sometimes over-directive feedback source, captured students’ descriptions of GenAI feedback as immediate, accessible, and efficient for identifying obvious weaknesses in their drafts. Students in the direct GenAI-supported feedback condition frequently valued the speed and specificity of the feedback but also described a risk of following AI-suggested structures too readily, particularly when they had not first articulated their own judgment about the task. The second theme, reflective prompting as support for evaluative judgment, captured students’ reports that the self-evaluation step changed how they read AI-generated feedback. Rather than treating the feedback as an answer key, students described using it as a comparison point against their own diagnosis of strengths, weaknesses, and uncertainties. This pattern helps explain why the reflective condition was associated with stronger feedback uptake and self-regulated learning than direct GenAI-supported feedback alone. The third theme, hybrid feedback as comparison across perspectives, reflected students’ descriptions of the combined value of self-, peer-, and AI-based feedback. Participants often noted that these sources highlighted different weaknesses in the same draft, thereby supporting more deliberate judgment about which revisions to prioritize. Students in the hybrid condition also reported that the revision memo encouraged them to justify rather than simply enact revision decisions. The fourth theme, tensions between efficiency and ownership of revision, described students’ sense that AI-enabled feedback reduced the effort required to generate revision ideas while sometimes weakening their sense of authorship over the final text. This tension was most visible in AI-supported conditions and was described by some instructors as a pedagogical challenge: the same features that made AI feedback attractive could also reduce the extent to which students saw revision as their own evaluative work. Taken together, the qualitative findings help explain why the reflective and hybrid conditions were associated with stronger feedback uptake and self-regulated learning than the direct GenAI-supported feedback condition. They also reinforce the broader interpretation that more agentic feedback designs preserve student ownership over revision by requiring comparison, selection, and justification rather than simple implementation of external critique. Table 8. Qualitative themes and representative quotations Theme Description Representative quotation Typical conditions AI as a fast but sometimes over-directive feedback source Students valued the speed and clarity of AI-generated critique but sometimes felt that it encouraged formulaic revision. “It helped me see what was missing almost immediately, but sometimes it felt as though I was borrowing its structure instead of testing my own reasoning.” (S07, direct GenAI-supported feedback, biology) Mostly direct GenAI-supported feedback Reflective prompting as support for evaluative judgment Completing self-evaluation before reading AI feedback prompted comparison, verification, and more active judgment. “Because I had already written down what I thought was weak, I was not just accepting the AI’s comments. I was checking where it agreed with me and where it made me rethink the argument.” (S18, reflective GenAI-supported feedback, chemistry) Mostly reflective GenAI-supported feedback Hybrid feedback as comparison across perspectives Students described hybrid feedback as helpful because self, peer, and AI input highlighted different weaknesses in the same draft. “The peer comments showed me what was unclear to a reader, and the AI pointed out where my evidence still was not doing enough work. Seeing both made it easier to decide what actually needed revision.” (S24, hybrid feedback, physics) Mostly hybrid feedback Tensions between efficiency and ownership of revision Students reported a trade-off between faster revision and maintaining a sense of authorship and control over argument quality. “It definitely saved time, but I had to stop and ask whether this was still my explanation or whether I was just polishing the version the tool seemed to prefer.” (S31, hybrid feedback, chemistry) Across AI-supported conditions 5.7. Sensitivity analyses Several sensitivity analyses were conducted to examine the robustness of the main findings. First, the primary models were re-estimated using a per-protocol sample excluding students who completed fewer than two of the three intervention cycles (n = 1,102). The overall pattern of results remained substantively unchanged. Second, models excluding the three lowest-fidelity sections yielded parameter estimates that differed by less than 0.04 standard deviations from the main intention-to-treat analyses. Third, alternative missing-data specifications using full-information maximum likelihood produced conclusions identical to those obtained from multiple imputation. Exploratory moderation analyses examined whether the effects of feedback condition varied as a function of baseline feedback literacy and science domain. The interaction between condition and baseline feedback literacy was significant for delayed AI-free transfer, F(3, 1158) = 2.83, p = .037, indicating that the disadvantage of the direct GenAI-supported feedback condition relative to the reflective and hybrid conditions was more pronounced among students with lower initial feedback literacy. Because these moderation tests were exploratory and not the focus of the preregistered confirmatory analyses, they should be interpreted cautiously. By contrast, the interaction between condition and science domain was not significant for any primary or secondary outcome (all p > .10), indicating that the overall pattern of effects was largely consistent across biology, chemistry, and physics. Detailed sensitivity results are reported in the Supplementary Materials. These additional analyses support the robustness of the main findings while also suggesting that feedback literacy may condition the extent to which students benefit from more direct forms of AI-supported feedback. 6. Discussion 6.1. Interpretation of the principal findings The present study examined how alternative GenAI-supported feedback designs shaped revision, learning, and transfer in introductory university science courses. The findings indicate that the educational value of GenAI feedback depended less on AI access per se than on how students were positioned within the feedback process. Direct GenAI-supported feedback improved immediate four-dimension argument-quality gain relative to peer feedback, but reflective and hybrid feedback designs produced stronger outcomes on feedback uptake, self-regulated learning, conceptual learning, and delayed AI-free transfer. The hybrid condition yielded the highest adjusted mean for four-dimension argument-quality gain, whereas both reflective and hybrid designs outperformed direct GenAI-supported feedback on delayed AI-free transfer. Taken together, these findings suggest that GenAI can support rapid improvement in common-content argument quality, but more agentic feedback designs are more effective when the goal is durable learning and AI-independent transfer. The overall pattern is important because it differentiates between short-term improvement in argument quality and broader educational value. Direct GenAI-supported feedback appears to be effective when students need rapid, criterion-referenced critique to strengthen a draft in the moment. However, the stronger outcomes associated with reflective and hybrid designs indicate that longer-term conceptual gains and transfer are more likely when feedback processes require students to engage in their own evaluative work. In this respect, the study supports the view that immediate improvement and durable learning should not be treated as interchangeable outcomes in AI-mediated feedback environments. The supplementary indicators help sharpen this interpretation. Although the primary immediate outcome was defined as four-dimension argument-quality gain, the more agentic conditions were also associated with more substantive revision and stronger revised-draft strategic quality of revision. This pattern suggests that the advantages of reflective and hybrid feedback were not limited to larger gain scores alone; they were also linked to revision that appeared more purposeful, better integrated, and more epistemically substantive. The distinction matters because it indicates that stronger performance in the more agentic conditions was accompanied by qualitatively stronger revision activity rather than by superficial editing alone. The absence of statistically significant differences between reflective and hybrid conditions on several outcomes should also be interpreted cautiously. Rather than demonstrating equivalence in any strong sense, the pattern is consistent with the possibility that requiring learners to externalize an initial evaluative judgment before receiving AI critique is itself a particularly important ingredient of effective GenAI-supported feedback. The additional components included in the hybrid condition may still offer practical and pedagogical benefits, but these advantages were not always statistically separable from those associated with reflective GenAI-supported feedback alone in the present sample. 6.2. Theoretical and empirical implications These findings support the claim that feedback design is a more meaningful analytic unit than feedback source alone. In the present study, direct GenAI-supported feedback provided rapid, criterion-referenced critique, but it did so without requiring learners to externalize an initial judgment about the quality of their own work. By contrast, reflective and hybrid designs required learners to compare, interpret, and justify revision decisions before or alongside engagement with AI critique. The stronger outcomes associated with these more agentic designs are therefore consistent with process-oriented accounts of feedback, which emphasize that learning depends not merely on receiving comments, but on learners’ active interpretation and use of them (Carless & Boud, 2018 ; Nicol & Macfarlane-Dick, 2006 ). The findings also extend work on self-regulated learning in AI-mediated environments. The mediation analyses suggest that reflective and hybrid designs outperformed direct GenAI-supported feedback partly because they strengthened feedback uptake and self-regulated learning during revision. This pattern is consistent with SRL-oriented perspectives that view feedback as educationally powerful when it becomes part of learners’ own regulatory activity rather than remaining external advice (Nicol & Macfarlane-Dick, 2006 ; Panadero, 2017 ). The distinction between four-dimension argument-quality gain and the strategic quality of revision sharpens this interpretation further: the more agentic designs were associated not only with larger gains in common-content argument quality, but also with stronger supplementary indicators of revision substance and revised-draft strategic quality. Their advantages therefore appear to reflect a more productive organization of evaluative activity rather than short-term score improvement alone (Carless & Boud, 2018 ; Buckingham Shum et al., 2023 ). The study also contributes to emerging discussions of epistemic agency in AI-supported learning. Recent work has suggested that the educational use of automated feedback depends on how human and machine roles are distributed within feedback processes, and that AI assistance may weaken student agency when critique is accepted too readily or when difficult judgments are outsourced to the system (Buckingham Shum et al., 2023 ; Darvishi et al., 2024 ). The present results are consistent with those concerns: students in the direct GenAI-supported feedback condition described AI as fast and useful, but sometimes overly directive, whereas students in the reflective and hybrid conditions described self-evaluation, peer input, and revision justification as preserving a stronger sense of ownership. In this respect, the findings add empirical weight to the argument that the pedagogical value of GenAI depends not only on what the tool can generate, but also on whether surrounding feedback designs amplify or displace students’ evaluative responsibility (Buckingham Shum et al., 2023 ; Darvishi et al., 2024 ). They also align with prior work showing that AI- and peer-generated feedback offer partially complementary affordances (Banihashem et al., 2024 ), that AI-supported peer feedback is most effective when collaborative argumentation is scaffolded (Chang et al., 2026 ), and that GenAI in science learning is most educationally useful when it functions as a dialogic partner rather than an authoritative answer source (Tang & Putra, 2025 ). More broadly, the study adds nuance to recent syntheses of GenAI in higher education by suggesting that apparent benefits are highly design-sensitive and that concerns about authorship, evidence of learning, and learner responsibility are not only matters of policy or integrity, but also matters of feedback design (Wang & Fan, 2025 ; Xia et al., 2024 ). 6.3. Practical implications for higher education digital learning The findings have several practical implications for higher education digital learning. First, GenAI feedback should not be treated as pedagogically neutral or self-sufficient. When the goal is quick draft improvement, direct GenAI-supported feedback may be useful because it provides rapid, criterion-referenced critique with relatively low orchestration demands. However, when the goal is deeper learning, stronger reasoning, or AI-independent transfer, instructors should not present AI critique as the first or only evaluative input. The results instead suggest that AI critique should be preceded by a brief, structured self-evaluation that asks students to identify strengths, weaknesses, uncertainties, and revision priorities before viewing external feedback. Second, hybrid feedback designs appear particularly well suited to complex reasoning tasks in which students must compare perspectives, weigh evidence, and justify revision decisions. In such contexts, peer feedback can be scaffolded through a structured template, while AI critique can be positioned as a third perspective rather than a final answer. A short revision rationale or memo may further help students explain why particular comments were accepted, rejected, or modified. More broadly, the results imply that GenAI feedback systems should be designed around comparison, selection, and justification rather than passive implementation of suggestions. In practical terms, this means constraining regeneration, making rubric criteria visible, and prompting students to record at least one revision decision in their own words. At the institutional level, the study suggests that decisions about GenAI adoption should not be made only in terms of scalability or efficiency, but also in terms of whether digital-learning designs preserve student agency, evaluative judgment, and ownership of learning when AI tools are embedded in routine coursework. 6.4. Limitations Several limitations should be considered when interpreting the findings. First, the reflective and hybrid conditions involved greater amounts of guided evaluative activity, meaning that feedback-source configuration and reflective workload were not fully separable. Second, students in AI-supported conditions received additional system training, which may have increased condition-specific preparation time relative to the peer-feedback condition. Third, the study was conducted in first-year university science courses, and the extent to which the findings generalize to other disciplines, levels of study, or institutional contexts remains uncertain. Fourth, although AI-free transfer tasks were completed without access to the study’s GenAI system in supervised settings, the design cannot eliminate broader concerns about students’ external familiarity with AI beyond the intervention context. Fifth, the analytic sample may reflect a degree of consent-based selection, because instructional participation was compulsory whereas research participation was voluntary. Sixth, although peer feedback was structured through a common template, peer-comment quality was not independently scored as a separate explanatory variable, making it difficult to determine the extent to which condition differences may partly reflect variation in the quality of peer input. Seventh, the participating institutions were modeled as fixed effects because only four institutions were included. Finally, the moderation analyses were exploratory and should therefore be interpreted with caution. 6.5. Future research directions Future research should extend this work in several directions. First, studies should examine the minimum effective amount of self-evaluation required before AI critique, since the present findings suggest that prior self-evaluation matters but do not establish whether shorter or lighter reflective prompts would produce comparable benefits. Second, future work should compare teacher-plus-AI, peer-plus-AI, and hybrid self–peer–AI designs more directly in order to clarify which combinations are most effective for different types of higher education tasks. Third, longer-duration studies are needed to determine whether the advantages of reflective and hybrid designs persist across a semester or academic year and whether repeated cycles of AI-supported revision produce cumulative gains in feedback literacy, self-regulated learning, or AI-independent transfer. Finally, further qualitative and mixed-method research should examine how students experience responsibility, authorship, and evaluative judgment in AI-supported feedback environments across disciplines and institutional contexts. 7. Conclusion This study shows that the pedagogical value of GenAI feedback in higher education depends less on the availability of AI-generated critique than on how feedback processes are designed. In the present study, direct GenAI-supported feedback supported immediate improvement in four-dimension argument-quality gain, but reflective and hybrid designs were more consistently associated with stronger feedback uptake, self-regulated learning during revision, conceptual learning, and AI-independent transfer. These findings reinforce the view that feedback is educationally valuable not simply when it is delivered efficiently, but when learners are positioned to interpret, compare, and act on critique as active evaluative agents (Carless & Boud, 2018 ; Nicol & Macfarlane-Dick, 2006 ). More broadly, the findings suggest that the benefits of GenAI in higher education are highly design-sensitive. When AI feedback is organized primarily around speed and efficiency, it may support short-term improvement without necessarily strengthening the broader learning processes that underlie durable understanding and transfer. By contrast, when AI-supported feedback is embedded in reflective and comparative designs that preserve student judgment, self-regulation, and ownership, it is more likely to support meaningful learning (Buckingham Shum et al., 2023 ; Darvishi et al., 2024 ; Wang & Fan, 2025 ). For higher education digital learning, the central implication is therefore not whether students can access GenAI feedback, but how GenAI feedback is pedagogically organized so that human evaluative work remains at the center (Buckingham Shum et al., 2023 ; Xia et al., 2024 ). Declarations Availability of data and materials The datasets generated and/or analyzed during the current study are not publicly available because they include de-identified student performance, survey, interview, and platform-trace data collected under institutional ethics approvals and data-governance restrictions. De-identified data are available from the corresponding author on reasonable request, subject to institutional and ethical requirements. Competing interests The author declares that there are no competing interests. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Authors’ contributions The author was solely responsible for the conception and design of the study, the development of the intervention materials, data collection, data analysis, interpretation of the findings, and preparation of the manuscript. The author read and approved the final manuscript. Acknowledgements Not applicable. References Banihashem, S. K., Kerman, N. T., Noroozi, O., Moon, J., & Drachsler, H. (2024). Feedback sources in essay writing: Peer-generated or AI-generated feedback? International Journal of Educational Technology in Higher Education , 21 . https://doi.org/10.1186/s41239-024-00455-4 . Article 23. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology , 3 (2), 77–101. https://doi.org/10.1191/1478088706qp063oa Buckingham Shum, S., Lim, L. A., Boud, D., Bearman, M., & Dawson, P. (2023). A comparative analysis of the skilled use of automated feedback tools through the lens of teacher feedback literacy. International Journal of Educational Technology in Higher Education , 20 ., Article 40. https://doi.org/10.1186/s41239-023-00410-9 Carless, D., & Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. Assessment & Evaluation in Higher Education , 43 (8), 1315–1325. https://doi.org/10.1080/02602938.2018.1463354 Chang, Y., Liu, Q., Lu, Y., & Miao, E. (2026). Leveraging generative AI to facilitate peer feedback in collaborative argumentation learning. International Journal of Educational Technology in Higher Education , 23 ., Article 10. https://doi.org/10.1186/s41239-026-00586-w Darvishi, A., Khosravi, H., Sadiq, S. W., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education , 210 . https://doi.org/10.1016/j.compedu.2023.104967 . Article 104967. Driver, R., Newton, P., & Osborne, J. (2000). Establishing the norms of scientific argumentation in classrooms. Science Education , 84 (3), 287–312. https://doi.org/10.1002/(SICI)1098-237X(200005)84:3%3C287::AID-SCE1%3E3.0.CO;2-A Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research , 77 (1), 81–112. https://doi.org/10.3102/003465430298487 Molloy, E., Boud, D., & Henderson, M. (2020). Developing a learning-centred framework for feedback literacy. Assessment & Evaluation in Higher Education , 45 (4), 527–540. https://doi.org/10.1080/02602938.2019.1667955 Nicol, D. (2021). The power of internal feedback: Exploiting natural comparison processes. Assessment & Evaluation in Higher Education , 46 (5), 756–778. https://doi.org/10.1080/02602938.2020.1823314 Nicol, D. J., & Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. Studies in Higher Education , 31 (2), 199–218. https://doi.org/10.1080/03075070600572090 Osborne, J., Erduran, S., & Simon, S. (2004). Enhancing the quality of argumentation in school science. Journal of Research in Science Teaching , 41 (10), 994–1020. https://doi.org/10.1002/tea.20035 Panadero, E. (2017). A review of self-regulated learning: Six models and four directions for research. Frontiers in Psychology , 8 ., Article 422. https://doi.org/10.3389/fpsyg.2017.00422 Panadero, E., Jonsson, A., & Botella, J. (2017). Effects of self-assessment on self-regulated learning and self-efficacy: Four meta-analyses. Educational Research Review , 22 , 74–98. https://doi.org/10.1016/j.edurev.2017.08.004 Pintrich, P. R., Smith, D. A. F., Garcia, T., & McKeachie, W. J. (1991). A manual for the use of the Motivated Strategies for Learning Questionnaire (MSLQ) (NCRIPTAL Technical Report No. 91-B-004). National Center for Research to Improve Postsecondary Teaching and Learning, University of Michigan. https://files.eric.ed.gov/fulltext/ED338122.pdf Tai, J., Ajjawi, R., Boud, D., Dawson, P., & Panadero, E. (2018). Developing evaluative judgement: Enabling students to make decisions about the quality of work. Higher Education , 76 (3), 467–481. https://doi.org/10.1007/s10734-017-0220-3 Tang, K. S., & Putra, G. B. S. (2025). Generative AI as a dialogic partner: Enhancing multiple perspectives, reasoning, and argumentation in science education with customized chatbots. Journal of Science Education and Technology , 35 , 128–140. https://doi.org/10.1007/s10956-025-10240-1 Wang, J., & Fan, W. (2025). The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. Humanities and Social Sciences Communications , 12 , 621. https://doi.org/10.1057/s41599-025-04787-y Weidlich, J., Jivet, I., Woitt, S., Orhan Göksün, D., Kraus, J., & Drachsler, H. (2025). The student feedback literacy instrument (SFLI): Multilingual validation and introduction of a short-form version. Assessment & Evaluation in Higher Education , 50 (5), 677–693. https://doi.org/10.1080/02602938.2025.2451729 Woitt, S., Weidlich, J., Jivet, I., Orhan Göksün, D., Drachsler, H., & Kalz, M. (2025). Students’ feedback literacy in higher education: An initial scale validation study. Teaching in Higher Education , 30 (1), 257–276. https://doi.org/10.1080/13562517.2023.2263838 Xia, Q., Weng, X., Ouyang, F., Lin, T. J., & Chiu, T. K. F. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education , 21 ., Article 40. https://doi.org/10.1186/s41239-024-00468-z Zhu, M., Liu, O. L., & Lee, H. S. (2020). The effect of automated feedback on revision behavior and learning gains in formative assessment of scientific argument writing. Computers & Education , 143 , 103668. https://doi.org/10.1016/j.compedu.2019.103668 Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterialsImprovedv4withfigures.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9396658","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":631045071,"identity":"21bb8f5f-153d-4f18-9e7d-09c12572c768","order_by":0,"name":"Huseyin ATES","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJElEQVRIie2RMUvDQBTHXzw4lwdZE2LrVzg5iIrFfhVFSJYIHR0clEI2O0cQ/Ar9CIEHlyUfoGCXEKhLh0BAChU1URSVS10d7jcc3L378d7dH8Bg+IeIbytABSl+VqyrPxXe3Ep+KmyDAh8KQ0i/9p3KvhMWJY7mYN+PqR5cznds57yoRzDoTVNbVRrlMImkRLEAR/HAi9QC3SSUXgKBnKaMJbrBZhH3UBCAQp9FnFDkCjwEOm0V7Vtm4eO6VXaVXdcHL4TDXLE1wusG5cRnrSIUgmfFTZftuOkLabeSL6V7Jwj3VOC7NxNCJ4v5EYozeUtMapUsLKrlM/X7RGW1eqKhPebsAS+Oe5PsutT+csNWEx/+Omvz0sfyjrXqrhkMBoMB4A2g2l8xlIOpUgAAAABJRU5ErkJggg==","orcid":"","institution":"Ahi Evran University","correspondingAuthor":true,"prefix":"","firstName":"Huseyin","middleName":"","lastName":"ATES","suffix":""}],"badges":[],"createdAt":"2026-04-12 20:23:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9396658/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9396658/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108135094,"identity":"9446721c-49a5-4c53-aac3-eda05ebc718b","added_by":"auto","created_at":"2026-04-29 17:22:16","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":19938,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eProcess model of GenAI-supported feedback in scientific argumentation.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9396658/v1/9edea37ecf70a9e6f1c69a61.jpg"},{"id":108182616,"identity":"b1244082-a4bc-43d9-b74a-15155e21a04d","added_by":"auto","created_at":"2026-04-30 08:59:27","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":39594,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy procedure across the 10-week intervention.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-9396658/v1/dc3dc29b8ac1feae4f0b5d72.png"},{"id":108135096,"identity":"3a6ed9e7-073c-42a3-88ff-f505bf654ea0","added_by":"auto","created_at":"2026-04-29 17:22:16","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":24470,"visible":true,"origin":"","legend":"\u003cp\u003eEstimated marginal means for four-dimension argument-quality gain across conditions.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-9396658/v1/ee122719658fc2451ada2f20.png"},{"id":108135097,"identity":"75bc85f7-b130-42ac-bc40-4122beffb614","added_by":"auto","created_at":"2026-04-29 17:22:16","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":34188,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEstimated marginal means for conceptual learning and delayed AI-free transfer\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-9396658/v1/774368e8c72d5e49e37d2df3.png"},{"id":108183506,"identity":"b896a107-d0b9-4a71-8ef6-61b0db493ad4","added_by":"auto","created_at":"2026-04-30 09:01:51","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":849433,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9396658/v1/0af54e72-1a57-4aa5-8424-989c6136968a.pdf"},{"id":108135093,"identity":"5bc4a9bd-b0a9-4880-ab7e-e5b68de23356","added_by":"auto","created_at":"2026-04-29 17:22:16","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":244915,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialsImprovedv4withfigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-9396658/v1/5e8b3b7e3647b39dda539804.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Human-Centered GenAI Feedback Design in Higher Education: A Multisite Experiment on Direct, Reflective, and Hybrid Approaches to Scientific Argumentation","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eGenerative artificial intelligence (GenAI) is increasingly embedded in higher education, including feedback, assessment, and learning support, yet its educational significance remains unsettled (Wang \u0026amp; Fan, 2025; Xia et al., 2024). Existing evidence suggests that GenAI can support learning performance, learning perceptions, and higher-order thinking, but these benefits vary substantially depending on the role assigned to the tool, the instructional design, and the characteristics of the task (Wang \u0026amp; Fan, 2025). In feedback contexts, this variation is especially important. Feedback is not educationally valuable simply because comments are delivered; rather, it becomes productive when learners interpret feedback, compare it against criteria, judge its relevance, and use it to improve subsequent work (Carless \u0026amp; Boud, 2018; Hattie \u0026amp; Timperley, 2007; Nicol \u0026amp; Macfarlane-Dick, 2006). Thus, the central issue in GenAI-supported feedback is not automation alone, but whether feedback design sustains active learner engagement, feedback uptake, and evaluative agency (Buckingham Shum et al., 2023; Darvishi et al., 2024).\u003c/p\u003e\n\u003cp\u003eThis question is particularly consequential in higher education science learning. Scientific argumentation requires students not only to produce acceptable answers, but also to formulate claims, justify them with evidence and reasoning, consider alternative explanations, and revise disciplinary explanations in response to critique (Driver et al., 2000; Osborne et al., 2004). Compared with more generic writing tasks, scientific argumentation places stronger demands on disciplinary judgment because revision depends not only on coherence or fluency, but also on the adequacy of evidence, the quality of reasoning, and the treatment of uncertainty or alternative explanations (Osborne et al., 2004; Zhu et al., 2020). In this context, GenAI-supported feedback is pedagogically promising but also potentially risky: it may help students improve written products and, when designed dialogically, support deeper reasoning, yet it may also reduce the need for students to make their own evaluative judgments if the feedback process encourages passive uptake (Darvishi et al., 2024; Tang \u0026amp; Putra, 2025; Zhu et al., 2020).\u003c/p\u003e\n\u003cp\u003eRecent reviews suggest that GenAI is expanding rapidly in higher education, yet three substantive gaps remain in the feedback literature (Wang \u0026amp; Fan, 2025; Xia et al., 2024). First, much of the literature on AI-supported feedback has focused on general writing tasks or broad comparisons among feedback sources, whereas disciplinary reasoning in university science courses has received less sustained attention (Banihashem et al., 2024; Tang \u0026amp; Putra, 2025; Wang \u0026amp; Fan, 2025). Second, many studies compare a single AI condition with a single non-AI condition, leaving underexplored how alternative feedback designs may distribute evaluative work differently across learners, peers, and GenAI, especially when self-evaluation, peer input, and GenAI critique are differently combined or sequenced (Banihashem et al., 2024; Chang et al., 2026; Darvishi et al., 2024). Third, immediate product improvement is often examined without simultaneously investigating feedback uptake, self-regulated learning, and later performance without AI support, making it difficult to determine whether observed gains reflect durable learning, stronger revision processes, or temporary performance support (Wang \u0026amp; Fan, 2025; Xia et al., 2024; Zhu et al., 2020).\u003c/p\u003e\n\u003cp\u003eBeyond the immediate question of whether GenAI can generate feedback efficiently, the broader issue for higher education digital learning is how feedback environments should be designed so that speed, scalability, and accessibility do not come at the expense of student agency, evaluative judgment, and ownership of learning (Buckingham Shum et al., 2023; Darvishi et al., 2024; Xia et al., 2024). Under these conditions, the key pedagogical question is not simply whether students have access to feedback, but whether feedback design positions them as active interpreters, comparators, and decision-makers in relation to critique—learners who can exercise evaluative judgment and generate internal feedback rather than merely implement external suggestions (Carless \u0026amp; Boud, 2018; Tai et al., 2018; Molloy et al., 2020; Nicol, 2021). Scientific argumentation provides a critical case for examining this problem because it offers a demanding test of whether GenAI functions as a support for reasoning or as a shortcut around reasoning.\u003c/p\u003e\n\u003cp\u003eAgainst this background, the present study addresses this problem through a multisite, cluster-randomized, longitudinal field experiment in introductory university science courses. The study compares four feedback designs—peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and a hybrid design combining self-evaluation, peer feedback, and GenAI critique—and examines not only immediate four-dimension argument-quality gain, but also conceptual learning, delayed AI-free transfer, and the process pathways of feedback uptake and self-regulated learning. In doing so, the study contributes to higher education digital learning research by shifting attention from access to AI-generated feedback toward the design conditions under which GenAI can support meaningful revision, disciplinary reasoning, and learner agency. More specifically, by combining four feedback designs, process mediation, and delayed AI-free transfer in a multisite field experiment, the study asks how AI can be integrated into feedback processes in ways that amplify rather than displace students’ evaluative work..\u003c/p\u003e"},{"header":"2. Literature review and theoretical framework","content":"\u003cp\u003e\u003cstrong\u003e2.1. Feedback as uptake, judgment, and revision\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe present study is grounded in a process-oriented understanding of feedback. In this perspective, feedback is not reducible to information transmitted from a source to a learner. Rather, it becomes educationally meaningful when learners compare their current performance against criteria or standards, interpret the significance of comments, and use those judgments to guide subsequent action (Carless \u0026amp; Boud, 2018; Nicol \u0026amp; Macfarlane-Dick, 2006; Nicol, 2021). This view shifts attention from feedback provision to feedback use and from the delivery of comments to the learner processes through which comments are evaluated and transformed into revision.\u003c/p\u003e\n\u003cp\u003eThis learner-centered view is closely aligned with evaluative judgment, defined as the capacity to make decisions about the quality of work, whether one\u0026rsquo;s own or that of others (Tai et al., 2018). From this perspective, the educational value of feedback depends not simply on the accuracy or quantity of comments, but on whether students are positioned to judge quality, compare alternatives, and decide how revision should proceed. Carless and Boud (2018) conceptualized this broader learner capability through feedback literacy, which includes appreciating feedback, making judgments, managing affective responses, and taking action. Similarly, Molloy et al. (2020) argued that feedback-literate learning environments should support students\u0026rsquo; active role in interpreting and using feedback rather than treating them as passive recipients of commentary.\u003c/p\u003e\n\u003cp\u003eThis distinction is especially important in technology-mediated feedback. In higher education, feedback increasingly occurs through digitally mediated environments, including learning-management systems, platform-based review processes, and AI-enabled tools that make rapid and scalable commentary more readily available (Buckingham Shum et al., 2023). Yet access to larger volumes of feedback does not automatically generate learning. Nicol (2021) argued that feedback becomes most productive when it stimulates internal feedback, that is, when learners engage in comparison processes between current performance, criteria, alternatives, and desired standards. For the purposes of the present study, feedback literacy refers to a broader learner capability or disposition, whereas feedback uptake refers to a task-specific process: how learners attend to, interpret, select, and implement critique during a particular episode of revision. Accordingly, the present study treats feedback uptake as the more proximal process through which external comments are translated into revision decisions and learning-relevant action.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.2. Self-regulated learning and epistemic agency in AI-mediated feedback\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFeedback uptake is closely related to self-regulated learning (SRL). SRL theories describe learning as an active process in which students plan, monitor, evaluate, and adapt their cognition, motivation, and behavior in response to task demands (Nicol \u0026amp; Macfarlane-Dick, 2006; Panadero, 2017). In feedback contexts, SRL is central because students must decide how to interpret comments, whether to trust them, which revisions to prioritize, and how to integrate them into future performance. Feedback, in this sense, does not merely inform learning; it becomes one of the mechanisms through which learners regulate learning.\u003c/p\u003e\n\u003cp\u003eThis relationship can be understood more precisely through evaluative judgment and internal feedback. When students compare external comments with task criteria, prior understanding, and their own initial diagnosis of the work, they generate internal feedback that can inform planning, monitoring, and strategic adjustment during revision (Nicol, 2021; Tai et al., 2018). In the present study, feedback uptake is therefore treated as a proximal revision process that helps shape subsequent self-regulation. When students actively attend to, interpret, and select among feedback messages, they create the conditions under which revision becomes strategic rather than merely reactive.\u003c/p\u003e\n\u003cp\u003eAI-mediated feedback intensifies this issue. GenAI may support SRL when it helps learners identify weaknesses, clarify evaluative criteria, compare alternatives, and plan revision moves in ways that remain cognitively engaging for the learner (Buckingham Shum et al., 2023). However, it may also undermine SRL if it encourages students to accept suggestions uncritically or to outsource difficult evaluative judgments. Darvishi et al. (2024) showed that AI assistance in peer-feedback environments may increase dependence on the tool, raising questions about student agency in AI-supported learning contexts. This concern can be framed more precisely as a question of epistemic agency: whether students remain responsible for judging the adequacy of claims, evidence, and reasoning, or whether these judgments are displaced by AI-generated suggestions. In this sense, epistemic agency is closely related to ownership, evaluative judgment, and responsibility for disciplinary adequacy (Tai et al., 2018; Nicol, 2021).\u003c/p\u003e\n\u003cp\u003eResearch on self-assessment provides an important implication for this problem. Meta-analytic findings indicate that self-assessment can positively influence self-regulated learning and self-efficacy because it requires learners to compare their work against explicit criteria and to generate internal judgments before receiving or acting on external feedback (Panadero et al., 2017). This point is especially relevant for AI-supported revision: if students are first required to articulate their own diagnosis of strengths, weaknesses, or uncertainties, they may be better positioned to engage with GenAI critique as a resource for comparison rather than as a substitute for judgment. From this perspective, feedback designs that require prior self-evaluation before exposure to AI-generated commentary may offer more supportive conditions for SRL and epistemic agency than designs in which AI feedback is presented immediately and directly.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.3. Scientific argumentation as disciplinary and epistemic work\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eScientific argumentation provides a particularly demanding context for studying AI-supported feedback because it is epistemically different from generic academic writing. In introductory university science courses, students must do more than produce coherent text. They are expected to formulate claims, coordinate claims with evidence and reasoning, address limitations or counterpositions, and revise explanations in light of critique (Driver et al., 2000; Osborne et al., 2004). These demands make scientific argumentation a domain in which revision quality depends on disciplinary judgment rather than only on stylistic improvement.\u003c/p\u003e\n\u003cp\u003eThis distinction has direct implications for feedback. Feedback on scientific argumentation must address the epistemic adequacy of an argument: whether evidence is relevant and sufficient, whether reasoning justifies the claim, whether uncertainty or alternative explanations are acknowledged, and whether revisions improve explanatory power. Zhu et al. (2020) demonstrated that, in the formative assessment of scientific argument writing, revisions were positively associated with learning gains and that contextualized feedback was more effective than generic feedback. Emerging work also suggests that GenAI may support disciplinary engagement when it is designed in dialogic rather than answer-providing ways. Tang and Putra (2025), for example, found that a customized chatbot could support perspective-taking, reasoning, and argumentation in science learning when it functioned as a dialogic partner rather than an authoritative answer source.\u003c/p\u003e\n\u003cp\u003eFor this reason, scientific argumentation serves as a critical case for testing whether GenAI-supported feedback amplifies or displaces students\u0026rsquo; evaluative work in higher education. If student judgment and ownership can be preserved in this context, the implications are likely to extend beyond science to other higher education settings that require evidence-based explanation, justification, and revision.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.4. Why direct, reflective, and hybrid feedback designs may differ\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe core theoretical claim of the present study is that alternative feedback designs should not be expected to generate identical learning processes because they distribute evaluative work differently across learners, peers, and AI-supported tools. From a human-centered digital learning perspective, the central issue is not AI access alone, but whether the feedback design preserves learners\u0026rsquo; responsibility for interpreting critique, making judgments, and deciding how revision should proceed. In this sense, feedback design matters because it shapes whether students engage in evaluative judgment and internal feedback processes or whether those processes are prematurely displaced by external suggestions (Tai et al., 2018; Nicol, 2021).\u003c/p\u003e\n\u003cp\u003eIn direct GenAI-supported feedback, students receive AI-generated critique without first externalizing their own diagnosis of the task. Such a design may improve efficiency and support immediate four-dimension argument-quality gain by providing rapid, criterion-referenced suggestions. At the same time, because evaluative work is front-loaded onto the system rather than the learner, direct GenAI-supported feedback may offer weaker conditions for internal comparison, self-generated judgment, and strategic revision. For this reason, it may also provide weaker conditions for feedback uptake, self-regulated revision, and delayed AI-free transfer than designs that require more active learner judgment.\u003c/p\u003e\n\u003cp\u003eIn reflective GenAI-supported feedback, students first evaluate their own work against explicit criteria or identify areas of uncertainty before engaging with AI-generated critique. Its theoretical advantage is that it requires learners to articulate an initial judgment and then compare that judgment with external feedback. This sequence is important because evaluative judgment develops through opportunities to make and test quality-related decisions rather than merely receive answers about what to change (Tai et al., 2018). It is also consistent with Nicol\u0026rsquo;s (2021) argument that learning is strengthened when feedback triggers internal comparison processes. Because students are not positioned as passive recipients of commentary, reflective GenAI-supported feedback may be expected to strengthen feedback uptake, planning, monitoring, and strategy adjustment during revision, thereby supporting not only four-dimension argument-quality gain but also more durable conceptual learning and AI-independent transfer.\u003c/p\u003e\n\u003cp\u003eIn hybrid feedback, self-evaluation, peer feedback, and GenAI critique are combined within the same instructional sequence. This design is theoretically important because these feedback sources appear to offer partially complementary rather than identical affordances. Self-evaluation foregrounds internal criteria and initial judgment, peer feedback highlights reader-oriented comprehensibility and ambiguity, and GenAI critique may identify rubric-aligned omissions, weaknesses, or alternative revision moves (Banihashem et al., 2024; Chang et al., 2026). Taken together, these sources may create the strongest comparison ecology for active judgment and revision: students are required not only to receive multiple forms of critique, but also to compare them, select among them, and justify how revision decisions should proceed. In this sense, hybrid feedback may provide especially strong support for internal feedback processes, evaluative judgment, and ownership of revision, although it also entails greater coordination demands.\u003c/p\u003e\n\u003cp\u003eBased on this framework, the present study proposes a process model in which feedback design is expected to affect learning outcomes both directly and indirectly. Direct effects remain plausible because alternative designs may immediately alter the quality, timing, and diversity of critique available to learners. Indirect effects are expected to operate through feedback uptake, through self-regulated learning during revision, and through the serial pathway from uptake to self-regulated revision. On this account, the more agentic designs, especially reflective and hybrid feedback, should be better positioned to support stronger four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer because they require learners to compare, select, and justify revision decisions rather than merely implement external suggestions. The proposed process model is presented in Figure 1.\u003c/p\u003e\n\u003cp\u003eThe model proposes that feedback design affects four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer both directly and indirectly through feedback uptake and self-regulated learning during revision. Feedback uptake is modeled as a proximal process through which learners interpret, select, and implement critique, while self-regulated learning captures the planning, monitoring, and strategy adjustment that follow during revision. The more agentic feedback designs are expected to strengthen the serial pathway from feedback uptake to self-regulated revision. The model is situated in introductory university science courses and scientific argumentation tasks.\u003c/p\u003e"},{"header":"3. Research questions and hypotheses","content":"\u003cp\u003eFor clarity, the primary immediate outcome is defined as four-dimension argument-quality gain, that is, improvement from first draft to revised submission on the four content dimensions shared across draft stages. The strategic quality of revision is treated as a separate revised-draft indicator and is used in a supplementary interpretive role rather than as the primary gain outcome.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRQ1.\u0026nbsp;\u003c/strong\u003eHow do peer feedback, direct GenAI-supported feedback, reflective GenAI-supported feedback, and hybrid feedback differ in their effects on students’ immediate four-dimension argument-quality gain in scientific argumentation tasks?\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRQ2.\u0026nbsp;\u003c/strong\u003eHow do these four feedback conditions differ in their effects on students’ conceptual learning and delayed AI-free transfer in introductory university science courses?\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRQ3.\u0026nbsp;\u003c/strong\u003eHow do these feedback conditions differ in terms of students’ feedback uptake and self-regulated learning during the revision process?\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRQ4.\u0026nbsp;\u003c/strong\u003eTo what extent do feedback uptake and self-regulated learning mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer?\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH1a.\u0026nbsp;\u003c/strong\u003eReflective GenAI-supported feedback and hybrid feedback will yield higher levels of feedback uptake than direct GenAI-supported feedback because they require learners to make evaluative judgments before and during engagement with AI-generated critique.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH1b.\u0026nbsp;\u003c/strong\u003eNo directional hypothesis is specified between reflective GenAI-supported feedback and hybrid feedback, although the hybrid design may provide additional advantages through multi-source comparison.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH2a.\u0026nbsp;\u003c/strong\u003eReflective GenAI-supported feedback and hybrid feedback will yield higher levels of self-regulated learning during revision than direct GenAI-supported feedback.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH2b.\u0026nbsp;\u003c/strong\u003eNo directional hypothesis is specified between reflective GenAI-supported feedback and hybrid feedback for self-regulated learning during revision.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH3.\u0026nbsp;\u003c/strong\u003eReflective GenAI-supported feedback and hybrid feedback are expected to support stronger conceptual learning and delayed AI-free transfer than direct GenAI-supported feedback because they preserve learners’ evaluative judgment and promote more agentic revision.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH4.\u0026nbsp;\u003c/strong\u003eDirect GenAI-supported feedback may be expected to improve immediate four-dimension argument-quality gain relative to peer feedback by providing rapid, criterion-referenced critique; however, reflective and hybrid feedback designs are expected to yield comparable or stronger gains when deeper feedback uptake and self-regulated revision are supported.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH5a.\u0026nbsp;\u003c/strong\u003eFeedback uptake will mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH5b.\u0026nbsp;\u003c/strong\u003eSelf-regulated learning during revision will mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eH5c.\u0026nbsp;\u003c/strong\u003eFeedback uptake and self-regulated learning will jointly mediate the relationship between feedback condition and students’ four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer through a serial pathway from uptake to self-regulated revision.\u003c/p\u003e"},{"header":"4. Method","content":"\u003ch3\u003e4.1. Design\u003c/h3\u003e\n\u003cp\u003eThis study employed a multisite, cluster-randomized, longitudinal field experiment to examine how different feedback designs shaped students\u0026rsquo; scientific argumentation and learning in introductory university science courses. Cluster randomization was used because the intervention was implemented at the level of intact course sections, thereby reducing contamination across conditions during peer interaction and revision activities. The unit of randomization was the course section, whereas the primary unit of analysis was the student, with clustering handled through multilevel modeling. Four conditions were compared: peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and hybrid feedback combining self-evaluation, peer feedback, and GenAI critique.\u003c/p\u003e\n\u003cp\u003eThe study followed a pre-test, post-test, and delayed post-test structure. The primary immediate outcome was four-dimension argument-quality gain in scientific argumentation tasks, defined as improvement from the initial draft to the revised submission on the four content dimensions shared across draft stages. A separate revised-draft indicator captured the strategic quality of revision and was used as a concurrent supplementary indicator rather than as the primary gain outcome. Secondary outcomes were conceptual learning and delayed AI-free transfer. Two process variables\u0026mdash;feedback uptake and self-regulated learning during revision\u0026mdash;were specified a priori as mediators linking feedback design to learning outcomes. The design, primary outcomes, focal contrasts, and analysis plan were preregistered on the Open Science Framework (OSF) prior to post-test analysis. The preregistered focal contrasts emphasized comparisons of the more agentic feedback designs with direct GenAI-supported feedback while also retaining planned comparisons with the peer-feedback condition.\u003c/p\u003e\n\u003cp\u003eAn a priori power analysis for a four-arm cluster-randomized design indicated that, assuming an intraclass correlation of .05, an average cluster size of 26 students, a two-sided alpha of .05, and a minimum detectable standardized effect of d = .25 for the planned between-condition contrasts, 48 sections would provide statistical power of .82. Power calculations were conducted in Optimal Design 3.01 under a balanced allocation scenario. On this basis, the target sample was set at approximately 1,200\u0026ndash;1,300 students across four institutions.\u003c/p\u003e\n\u003ch3\u003e4.2. Setting and participants\u003c/h3\u003e\n\u003cp\u003eThe study was conducted across four universities offering compulsory first-year undergraduate courses in biology, chemistry, and physics. These courses were selected because they required students to complete written tasks involving explanation, evidence use, and revision of scientific arguments. Across the participating universities, the intervention was embedded in regular coursework rather than delivered as an external experimental activity.\u003c/p\u003e\n\u003cp\u003eBackground data collected at baseline included demographic and academic variables relevant to the analyses, together with prior GenAI experience and feedback-related measures; descriptive characteristics of the analytic sample are reported in the Results section. The initial sample comprised 1,248 first-year undergraduate students from 48 course sections. The section-level distribution was balanced across disciplines: 16 biology sections (n = 416), 16 chemistry sections (n = 420), and 16 physics sections (n = 412). Mean section size at enrollment was 26.0 students (SD = 2.7, range = 21\u0026ndash;31).\u003c/p\u003e\n\u003cp\u003eSections were randomly assigned to one of the four feedback conditions, resulting in 12 sections per condition. Randomization was conducted by a researcher not involved in instruction using a computer-generated random-number sequence. To preserve balance, randomization was blocked by institution and discipline, and section meeting time (morning/afternoon) was used as a secondary balancing variable. Allocation was completed after course enrollment closed and before the Week 1 pre-test.\u003c/p\u003e\n\u003cp\u003eParticipation in the instructional activities formed part of normal course requirements; participation in the research component, including questionnaire, draft, interview, and log-data analysis, was voluntary and based on informed consent. Inclusion in the analytic dataset required consent for research use of student work, survey responses, and platform data. Students who did not consent completed the same class activities but were excluded from the analytic dataset. Additional exclusions related to course withdrawal and missing outcome data are reported in the Results section. Because research consent was voluntary while instructional participation was compulsory, the analytic sample may reflect a degree of consent-based selection. To reduce institution-level heterogeneity, the study used a common intervention schedule, shared task-design procedures, and a common analytic rubric across sites.\u003c/p\u003e\n\u003ch3\u003e4.3. Instructional tasks and rubric\u003c/h3\u003e\n\u003cp\u003eThe intervention comprised three scientific argumentation cycles distributed across a 10-week period. In each cycle, students produced an initial draft, received condition-specific feedback, revised their text, and submitted a final version. Tasks were designed to require explicit coordination of claim, evidence, and reasoning and were aligned with the disciplinary content of each course.\u003c/p\u003e\n\u003cp\u003eTwo task formats were used: laboratory-based interpretive writing tasks and short socio-scientific argumentation tasks. In the former, students explained data patterns, justified conclusions, and addressed sources of uncertainty; in the latter, they defended a position using disciplinary concepts and evidence. To support comparability across sites and disciplines, tasks were co-developed by participating instructors through a shared design protocol and reviewed in cross-institution calibration meetings focusing on epistemic demands, expected difficulty, and rubric fit across biology, chemistry, and physics.\u003c/p\u003e\n\u003cp\u003eEach cycle required an initial draft of approximately 350\u0026ndash;500 words and a revised submission of 400\u0026ndash;600 words. Students had 72 hours to complete the first draft, 48 hours to review assigned feedback, and 72 hours to submit the revision. In the hybrid condition, the revision memo was limited to 100\u0026ndash;150 words.\u003c/p\u003e\n\u003cp\u003eStudents\u0026rsquo; work was evaluated using an analytic scientific argumentation framework adapted from prior research on scientific argumentation and automated feedback in scientific writing (e.g., Osborne et al., 2004; Zhu et al., 2020). The framework comprised five dimensions: claim quality, evidence relevance and sufficiency, coherence of reasoning, treatment of limitations or alternative explanations, and strategic quality of revision. The first four dimensions were scored for both initial and revised drafts and were used to estimate four-dimension argument-quality gain. The fifth dimension, strategic quality of revision, was scored only for revised submissions on the basis of paired comparison with the initial draft and was retained as a supplementary indicator of revised performance.\u003c/p\u003e\n\u003cp\u003eThe same four common content dimensions were applied across intervention and transfer tasks, with topic-specific descriptors added where necessary to preserve disciplinary appropriateness. Prior to data collection, raters completed two calibration sessions using anchor responses from all three disciplinary domains. Two trained raters independently scored all first and final drafts used in outcome assessment, and discrepancies greater than one point on any rubric dimension were resolved through discussion. Full rubric details and anchor descriptors are provided in the supplementary materials.\u003c/p\u003e\n\u003ch2\u003e4.4. Experimental conditions\u003c/h2\u003e\n\u003cp\u003eAn overview of the four intervention conditions is provided in Table 1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.4.1. Peer feedback only\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStudents in the peer-feedback condition exchanged drafts with two peers and provided rubric-based comments through a structured template covering claim, evidence, reasoning, and one revision suggestion. No AI support was available in this condition.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.4.2. Direct GenAI-supported feedback\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStudents in the direct GenAI condition uploaded their draft to the course-integrated GenAI interface immediately after drafting. The system generated criterion-referenced feedback on claim quality, evidence use, reasoning, limitations, and possible revision moves. Students could consult this feedback during revision, but they did not complete a structured self-evaluation before viewing the AI output.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.4.3. Reflective GenAI-supported feedback\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStudents in the reflective GenAI condition completed a brief structured self-evaluation aligned with the rubric before receiving AI-generated feedback. They identified strengths, weaknesses, and uncertainties in their draft, so that engagement with AI critique followed an initial evaluative judgment.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.4.4. Hybrid feedback\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStudents in the hybrid condition completed a fixed sequence of self-evaluation, peer feedback, GenAI critique, and revision memo. They first evaluated their own draft, then received comments from two peers, consulted the GenAI system for additional critique and revision suggestions, and finally completed a short memo explaining which suggestions they accepted, rejected, or modified. This condition was designed to maximize comparison across self-, peer-, and AI-based feedback.\u003c/p\u003e\n\u003cp\u003eBecause the reflective and hybrid conditions incorporated additional guided evaluative activity, observed differences may reflect not only feedback-source configuration and sequencing but also differences in reflective workload. Access to condition-specific feedback resources was monitored through platform records, and potential contamination through non-assigned resources is addressed under implementation fidelity.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1. Overview of the four feedback conditions\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCondition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eFeedback sources\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSequence\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eStudent role\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAI role\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eIntended pedagogical function\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDraft \u0026rarr; peer feedback from two peers \u0026rarr; revision \u0026rarr; final submission\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eReviewer and reviser\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eNone\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eTo support evaluative judgment through peer review and revision without AI assistance\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eGenAI feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDraft \u0026rarr; GenAI-generated feedback \u0026rarr; revision \u0026rarr; final submission\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eReviser\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eProvides criterion-referenced critique and revision suggestions\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eTo provide immediate, scalable, rubric-aligned feedback with minimal preparatory judgment by the learner\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSelf-evaluation + GenAI feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDraft \u0026rarr; structured self-evaluation \u0026rarr; GenAI-generated feedback \u0026rarr; revision \u0026rarr; final submission\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSelf-evaluator and reviser\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eProvides critique after the learner articulates an initial judgment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eTo strengthen feedback uptake by requiring prior evaluative judgment before AI-supported revision\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSelf-evaluation + peer feedback + GenAI feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDraft \u0026rarr; self-evaluation \u0026rarr; peer feedback from two peers \u0026rarr; GenAI-generated feedback \u0026rarr; revision memo \u0026rarr; final submission\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSelf-evaluator, reviewer, and reviser\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eProvides additional critique and revision suggestions after self- and peer-based judgment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eTo maximize active comparison across internal judgment, peer judgment, and AI-supported critique\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote.\u003c/em\u003e All conditions used the same scientific argumentation tasks and common analytic rubric. The intervention differed in the source and sequencing of feedback and, to some extent, in the amount of guided evaluative activity required of students.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.5. GenAI system and prompting protocol\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe AI-supported conditions used a secure, institutionally hosted large language model interface based on OpenAI\u0026rsquo;s GPT-5 model, accessed through the universities\u0026rsquo; approved application environment and integrated into the learning-management platform so that interaction logs could be captured while maintaining institutional data-governance requirements. Students were instructed not to enter personally identifying information into the system.\u003c/p\u003e\n\u003cp\u003eThe GenAI tool was configured to provide criterion-referenced feedback rather than replacement text. Prompts directed the system to identify strengths and weaknesses in the argument, flag missing or weak evidence, question unsupported claims, suggest revision moves, and pose follow-up questions intended to deepen reasoning, while explicitly prohibiting full-answer rewriting. To ensure consistency across institutions, all AI-supported conditions used the same prompt architecture, rubric anchors, interface settings, and fixed model parameters across all three intervention cycles. Instructors were not permitted to alter prompts during the intervention period. To reduce variability and prompt gaming, students were limited to a single feedback generation per draft and could not iteratively regenerate responses within the same cycle, although they could review the generated feedback during revision. Chat history was reset between cycles to avoid cross-task carryover. Full prompt materials, model settings, interface screenshots, and implementation rules are provided in the supplementary materials.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.6. Procedure\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study ran for 10 instructional weeks (Figure 2). In Week 1, students completed a pre-test assessing domain-specific prior knowledge and baseline scientific argumentation, together with a background survey covering prior achievement, prior GenAI experience, and baseline feedback literacy. In Week 2, all students received a common orientation to scientific argumentation, the analytic rubric, and constructive feedback. Students in AI-supported conditions also received additional training on ethical and effective use of the GenAI system. Although this additional training was necessary for procedural consistency, it may also have increased preparation time relative to the non-AI condition.\u003c/p\u003e\n\u003cp\u003eDuring Weeks 3\u0026ndash;8, students completed three argumentation cycles, each involving an initial draft, a condition-specific feedback sequence, and a revised submission. The two task formats were distributed across the intervention in alignment with course content, while sections within the same course followed the same task sequence. Interaction logs, timestamps, revision traces, and revision memos were collected throughout this period. In Week 9, students completed a post-test assessing conceptual learning and a novel transfer task without AI access. In Week 10, they completed a delayed AI-free transfer task and a post-intervention questionnaire assessing self-regulated learning during revision and perceived engagement with feedback. The delayed transfer task was conceptually related to, but topically distinct from, the Week 9 transfer task. After the intervention, a purposive subsample of students and instructors participated in semi-structured interviews.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7. Measures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAn overview of all measures, data sources, and timing is presented in Table 2.\u003c/p\u003e\n\u003ch2\u003e4.7.1. Four-dimension argument-quality gain\u003c/h2\u003e\n\u003cp\u003eThe primary immediate outcome was four-dimension argument-quality gain, operationalized as the change in students\u0026rsquo; common-content argument score from initial draft to revised submission. Gain was calculated on the four dimensions shared across draft stages\u0026mdash;claim quality, evidence relevance and sufficiency, coherence of reasoning, and treatment of limitations or alternative explanations\u0026mdash;each scored on a four-point scale, yielding directly comparable common-content subtotals ranging from 4 to 16 at both stages. Higher values indicated greater improvement in scientific argumentation quality from draft to revision.\u003c/p\u003e\n\u003cp\u003eTo distinguish change in common argument quality from revised-product quality, the study also retained a separate revised-draft indicator: strategic quality of revision. This fifth dimension was scored only for revised submissions and was interpreted as a concurrent supplementary indicator rather than as part of pre/post gain estimation. Revision depth was coded separately on a 3-point scale (1 = predominantly surface-level, 2 = mixed, 3 = predominantly substantive) and averaged across the three intervention cycles to support interpretation of whether observed gains reflected substantive rather than predominantly surface-level changes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7.2. Conceptual learning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptual learning was measured using discipline-specific pre/post assessments co-developed by instructors across the four institutions. Each assessment included 12\u0026ndash;15 items targeting the core concepts underlying the argumentation tasks and combined selected-response and short constructed-response formats. Raw scores were converted to percentages within discipline. Because the discipline-specific forms were not identical, post-test scores were standardized within discipline before pooled modeling, allowing conceptual-learning outcomes to be compared on a common relative metric across biology, chemistry, and physics. Internal consistency at post-test was acceptable across disciplines (McDonald\u0026rsquo;s \u0026omega; = .79 in biology, .82 in chemistry, and .80 in physics).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7.3. Delayed AI-free transfer\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDelayed AI-free transfer was assessed through a novel scientific argumentation task completed individually without access to the course-integrated GenAI interface. Although the task differed in topic from both the intervention activities and the Week 9 transfer assessment, it required the same underlying epistemic practices captured by the common-content scoring framework: formulating a claim, selecting relevant evidence, linking evidence to reasoning, and addressing uncertainty or competing explanations. Novelty was ensured by using prompts that had not appeared in course activities, practice materials, or post-test instruments.\u003c/p\u003e\n\u003cp\u003eScoring was based on the same four common content dimensions used across the intervention tasks, with topic-specific descriptors added where necessary to preserve disciplinary appropriateness. Because the delayed transfer task was completed as a single independent performance rather than as a paired draft sequence, strategic quality of revision was not applicable. In the present study, AI-free task completion meant that students had no access to the study\u0026rsquo;s GenAI tool and completed the task independently in a supervised classroom setting without internet-enabled devices.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7.4. Feedback uptake\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFeedback uptake was treated as a multi-indicator construct capturing the extent to which students attended to, interpreted, and incorporated critique during revision. It drew on four sources of evidence: students\u0026rsquo; post-revision rationales, coder-rated alignment between received feedback and implemented revisions, time spent reviewing assigned feedback resources in platform logs, and the proportion of substantive revisions traceable to feedback content. In the hybrid condition, the revision memo served as the post-revision rationale; in the other conditions, the same indicator was captured through a shorter structured rationale submitted with the final draft.\u003c/p\u003e\n\u003cp\u003eEach indicator was averaged across the three intervention cycles and z-standardized across the full analytic sample. Indicators were then examined using a one-factor measurement model, with standardized loadings ranging from .58 to .81. The final uptake index used in the multilevel models was computed as the equal-weight mean of the four z-standardized indicators, with higher values indicating stronger uptake.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7.5. Feedback literacy\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBaseline feedback literacy was assessed using the full 22-item Student Feedback Literacy Instrument (SFLI), a higher-education measure comprising the dimensions of feedback attitudes and feedback practices (Weidlich et al., 2025). The instrument was selected because it was specifically developed for use with higher-education students and builds on earlier initial scale-validation work conducted in the same context (Woitt et al., 2025). Students responded on a 5-point Likert scale ranging from 1 (strongly disagree) to 5 (strongly agree), and scores were calculated as the mean across items, with higher values indicating greater feedback literacy. Internal consistency in the present sample was high (McDonald\u0026rsquo;s \u0026omega; = .90).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7.6. Self-regulated learning during revision\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSelf-regulated learning (SRL) during revision was measured using a 12-item task-specific scale adapted from the higher-education self-regulated learning tradition represented by the Motivated Strategies for Learning Questionnaire (MSLQ; Pintrich et al., 1991) and from subsequent work emphasizing planning, monitoring, and strategic adaptation as core features of SRL (Panadero, 2017). The scale sampled four revision-relevant domains\u0026mdash;planning, monitoring, strategy adjustment, and evaluation\u0026mdash;with three items per domain. Responses were recorded on a 7-point Likert scale ranging from 1 (not at all true of me) to 7 (very true of me). After reverse-coding where necessary, the self-report component was scored as the mean across items, with higher scores indicating stronger self-regulated learning during revision. Internal consistency was satisfactory (McDonald\u0026rsquo;s \u0026omega; = .88).\u003c/p\u003e\n\u003cp\u003eTo complement self-report data, the study also incorporated digital trace indicators derived from the learning-management system, including revisiting feedback, spacing of revision activity, repeated consultation of the rubric, and sequencing of draft-comparison actions. These traces were treated as behavioral indicators of revision regulation rather than as stand-alone proxies for self-regulation. A trace-based SRL subindex was computed by z-standardizing the four indicators and averaging them with equal weight. For the primary multilevel analyses, the overall SRL construct was represented by the mean of the z-standardized self-report score and the z-standardized trace-based subindex, with higher values indicating stronger self-regulated learning during revision.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.7.7. Covariates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe models adjusted for prior knowledge, baseline scientific argumentation, prior achievement, prior GenAI experience, gender, and science domain. Prior achievement was operationalized as each student\u0026rsquo;s standardized university-entry score obtained from institutional records; because entry metrics differed across institutions, scores were standardized within institution before inclusion in the pooled analyses. Institutional variation was handled through institution fixed effects in the mixed-effects models. Exploratory moderation analyses additionally examined whether baseline feedback literacy conditioned the effects of the feedback designs.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2. Overview of measures\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eVariable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eConstruct type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eOperational definition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eInstrument/source\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTime of measurement\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eScoring/format\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePrior knowledge\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCovariate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDomain-specific understanding of core course concepts before the intervention\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eInstructor-developed pre-test\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eStandardized test score\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eBaseline scientific argumentation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCovariate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eInitial quality of written scientific argumentation before the intervention\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eScientific argumentation task scored on the four common content dimensions of the analytic framework\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFour-dimension common-content subtotal (range = 4\u0026ndash;16)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePrior achievement\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCovariate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAcademic achievement prior to university study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eInstitutional entry records\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eInstitution-standardized university-entry score\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePrior GenAI experience\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCovariate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSelf-reported familiarity and prior use of GenAI tools for study-related tasks\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eBackground questionnaire\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e5-point self-report scale\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback literacy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCovariate / exploratory moderator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eStudents\u0026rsquo; readiness to interpret, value, and act on feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e22-item SFLI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eMean scale score (1\u0026ndash;5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFour-dimension argument-quality gain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003ePrimary outcome\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eImprovement from first draft to final draft on the four content dimensions shared across draft stages\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAnalytic scientific argumentation framework; strategic quality of revision and revision-depth coding used as supplementary indicators\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeeks 3\u0026ndash;8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eChange in common-content subtotal (final minus initial; four shared dimensions only)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eConceptual learning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSecondary outcome\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eUnderstanding of disciplinary concepts addressed in the intervention tasks\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDiscipline-specific post-test\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDiscipline-standardized post-test score\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDelayed AI-free transfer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSecondary outcome\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAbility to construct a scientific argument on a novel task without AI support\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eNovel scientific argumentation task scored on the four common content dimensions of the analytic framework\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeek 10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFour-dimension common-content rubric score\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback uptake\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eProcess variable / mediator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eExtent to which feedback was attended to, interpreted, and incorporated into revision\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003ePost-revision rationales, coded alignment between feedback and revisions, platform log data, traceable substantive revisions\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeeks 3\u0026ndash;8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eEqual-weight mean of four z-standardized indicators averaged across cycles\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eSelf-regulated learning during revision\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eProcess variable / mediator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003ePlanning, monitoring, strategic adjustment, and evaluation during revision\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e12-item task-specific SRL scale plus digital trace indicators\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eWeeks 3\u0026ndash;10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eMean of z-standardized self-report score and z-standardized trace-based subindex\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eImplementation fidelity\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eQuality-control indicator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDegree to which the intervention was delivered as intended across sites and sections\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eInstructor logs, platform records, manual audits\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eThroughout intervention\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFidelity checklist / audit record\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote\u003c/em\u003e. For pooled analyses across biology, chemistry, and physics, discipline-specific conceptual-learning scores were standardized within discipline before modeling. Institutional variation in prior achievement was addressed by standardizing university-entry scores within institution before inclusion in the pooled analyses. Baseline scientific argumentation, four-dimension argument-quality gain, and delayed AI-free transfer were all based on the four common content dimensions shared across draft stages; the strategic quality of revision was scored only for revised intervention submissions and was interpreted as a concurrent supplementary indicator rather than as part of pre/post gain estimation. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.8. Fidelity of implementation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eImplementation fidelity was monitored through four procedures: use of a common intervention schedule, rubric, and instructional materials across sites; standardized prompt architecture and fixed system settings in the AI-supported conditions; platform logs verifying that students accessed only the feedback resources assigned to their condition; and manual audits of a random 15% sample of student cases across sites. Fidelity was assessed using a predefined checklist covering sequence integrity, timing compliance, access restrictions, and completion of condition-specific activities. Audits were conducted by trained research assistants who were not involved in instruction, and 20% of audited cases were double-coded, with high agreement on checklist items (Cohen\u0026rsquo;s \u0026kappa; = .89). Detailed audit criteria and procedures are reported in the supplementary materials.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.9. Data analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBecause students were nested within course sections and only four institutions participated, the primary quantitative analyses were estimated using mixed-effects models with students at Level 1, sections at Level 2, and institution entered as a fixed effect. This approach was preferred to specifying institutions as a random third level because the number of higher-level units was too small to support stable estimation of institution-level random effects. Analyses were conducted in R 4.3.2.\u003c/p\u003e\n\u003cp\u003eFor RQ1 and RQ2, separate mixed-effects models were estimated for four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer. Experimental condition was treated as the principal fixed effect, section was specified as a random intercept, and baseline scores and background variables were included as covariates. Planned pairwise contrasts were estimated from the fitted models using estimated marginal means and Holm-adjusted \u003cem\u003ep\u003c/em\u003e values. In line with the preregistered comparison strategy, outcomes collected across the three intervention cycles were averaged at the student level before modeling. The primary immediate outcome was modeled as a gain score, consistent with both the instructional design and the preregistered outcome definition.\u003c/p\u003e\n\u003cp\u003eFor RQ3, between-condition differences in feedback uptake and self-regulated learning during revision were examined using the same fixed- and random-effects structure. For RQ4, 2-1-1 multilevel mediation models were estimated with feedback condition at the section level and both mediators and outcomes at the student level. Indirect effects were estimated separately through feedback uptake, through self-regulated learning during revision, and through the serial pathway from uptake to self-regulated revision.\u003c/p\u003e\n\u003cp\u003eAll primary analyses followed an intention-to-treat approach. Missing data were handled using multiple imputation by chained equations with 20 imputations under a missing-at-random assumption. For mediation analyses, indirect effects were estimated using Monte Carlo confidence intervals based on 20,000 simulations. To assess robustness, the primary models were also re-estimated using per-protocol analyses, models excluding low-fidelity sections, and an alternative missing-data specification based on full-information maximum likelihood. Additional analytic details and robustness results are reported in the supplementary materials.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.10. Qualitative analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo complement the quantitative analyses, semi-structured interviews were conducted with a purposive subsample of 32 students and 12 instructors. The student subsample was balanced across conditions and approximately balanced across disciplines, while instructor interviews captured site-level variation in feasibility and implementation. Interviews focused on how participants interpreted feedback, made revision decisions, and experienced ownership in AI-supported revision. All interviews were audio-recorded, transcribed verbatim, and analyzed using reflexive thematic analysis following Braun and Clarke (2006).\u003c/p\u003e\n\u003cp\u003eThe qualitative component was interpretive rather than reliability-seeking. Themes were developed iteratively through memoing and team discussion, with reflexive attention to how researchers\u0026rsquo; assumptions informed coding. No inter-coder reliability coefficient was calculated, in keeping with the reflexive thematic analysis approach. The qualitative findings were used in an explanatory role to clarify how participants experienced the feedback designs and to interpret the mechanisms underlying the quantitative patterns.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.11. Ethical considerations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEthical approval was obtained from the institutional review boards of the participating universities before data collection. Participation in the research component was voluntary, and students were informed that non-participation would not affect course grades. All data were de-identified prior to analysis, and identifiable platform data were stored on institutionally approved encrypted servers with access restricted to the research team in accordance with local ethics procedures.\u003c/p\u003e\n\u003cp\u003eParticipants were informed of their right to withdraw their research data up to the point at which the dataset had been de-identified and merged for analysis. Because the intervention involved GenAI tools, students received explicit guidance on ethical AI use, data-entry limits, and the requirement that submitted work remain their own. Monitoring of AI use was limited to interactions within the study platform. To reduce the risk of unequal access across conditions, all groups received the same core instruction in scientific argumentation, rubric use, and feedback principles, and students in the non-AI condition were provided with access to the GenAI orientation materials and a demonstration version of the interface after the intervention.\u003c/p\u003e"},{"header":"5. Results","content":"\u003cp\u003e\u003cstrong\u003e5.1. Preliminary analyses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe final analytic sample comprised 1,176 students nested within 48 course sections across four institutions. Of the 1,248 students initially enrolled, 72 were excluded because of non-consent for research use of their data, course withdrawal before the first intervention cycle, or absence from both the post-test and delayed transfer sessions. Final sample sizes were balanced across conditions, and attrition did not differ significantly by condition, \u0026chi;\u0026sup2;(3) = 0.33, \u003cem\u003ep\u003c/em\u003e = .954. Missing data on the retained sample ranged from 0.6% to 3.4% across variables and were handled using multiple imputation with 20 imputations.\u003c/p\u003e\n\u003cp\u003eBaseline equivalence analyses indicated that the four conditions did not differ significantly on age, gender distribution, prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, or prior GenAI experience (all \u003cem\u003ep\u003c/em\u003e \u0026ge; .472), suggesting that the randomization procedure produced broadly comparable groups prior to the intervention. Descriptive statistics and baseline-equivalence tests are reported in Table 3.\u003c/p\u003e\n\u003cp\u003eThe internal consistency of the self-report measures was satisfactory, and inter-rater reliability for scientific argumentation scoring and revision-depth coding was high. Implementation fidelity was also high across sites: based on instructor logs, platform records, and checklist-based audits of 176 randomly selected student cases, 95.8% of applicable instructional steps were delivered as intended, with high agreement on fidelity checklist items (Cohen\u0026rsquo;s \u0026kappa; = .89). Cross-condition contamination was minimal. Taken together, these findings support the adequacy of the dataset and the integrity of the intervention. Section-level variance components and intraclass correlation coefficients are reported in Supplementary Table S1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3. Sample characteristics and baseline equivalence across conditions\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eVariable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ePeer feedback only (n = 293)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eDirect GenAI-supported feedback (n = 295)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eReflective GenAI-supported feedback (n = 294)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eHybrid feedback (n = 294)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTotal (N = 1,176)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTest statistic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eAge, years, M (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e18.92 (1.31)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e18.88 (1.27)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e18.95 (1.29)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e18.90 (1.34)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e18.91 (1.30)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 1172) = 0.14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.936\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eWomen, n (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e167 (57.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e171 (58.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e166 (56.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e169 (57.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e673 (57.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026chi;\u0026sup2;(3) = 0.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.981\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePrior knowledge (0\u0026ndash;100), M (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e58.63 (10.84)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e59.11 (10.52)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e60.02 (10.41)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e59.38 (10.76)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e59.29 (10.63)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 1172) = 0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.472\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eBaseline scientific argumentation (4\u0026ndash;16), M (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e10.32 (2.08)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e10.51 (2.03)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e10.64 (2.10)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e10.57 (2.07)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e10.51 (2.07)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 1172) = 1.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.349\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePrior achievement, M (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.05 (0.98)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.03 (1.01)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.04 (0.96)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.01 (1.00)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.00 (0.99)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 1172) = 0.71\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.548\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback literacy (1\u0026ndash;5), M (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3.41 (0.52)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3.44 (0.50)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3.47 (0.49)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3.45 (0.51)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3.44 (0.51)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 1172) = 0.67\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.571\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePrior GenAI experience (1\u0026ndash;5), M (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2.78 (0.89)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2.93 (0.92)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2.88 (0.90)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2.99 (0.91)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2.89 (0.91)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 1172) = 1.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.236\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eNote.\u0026nbsp;Baseline-equivalence tests are based on the final analytic sample. Prior achievement scores were standardized within institution before pooled comparison. No significant between-condition differences were observed at baseline. Baseline scientific argumentation scores reflect the four common content dimensions of the analytic framework and therefore ranged from 4 to 16 rather than including the revised-draft strategic-quality-of-revision indicator.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5.2.\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eEffects of feedback condition on four-dimension argument-quality gain\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo address RQ1, a linear mixed-effects model was estimated using students\u0026rsquo; mean four-dimension argument-quality gain across the three intervention cycles as the dependent variable, with section included as a random intercept and institution entered as a fixed effect. Feedback condition significantly predicted four-dimension argument-quality gain, \u003cem\u003eF\u003c/em\u003e(3, 43.8) = 18.76, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001.\u003c/p\u003e\n\u003cp\u003eAs shown in Figure 3 and Table 4, argument-quality gain was highest in the hybrid condition, followed by the reflective GenAI-supported feedback condition, the direct GenAI-supported feedback condition, and the peer-feedback condition. Relative to peer feedback, direct GenAI-supported feedback improved four-dimension argument-quality gain, and both reflective and hybrid feedback produced still larger gains. Both reflective and hybrid conditions also significantly outperformed direct GenAI-supported feedback, whereas the difference between the hybrid and reflective conditions was not statistically significant.\u003c/p\u003e\n\u003cp\u003eSupplementary analyses of coder-rated revision depth showed a parallel pattern. Relative to peer feedback, revision depth was higher in the direct GenAI-supported feedback condition and higher still in the reflective and hybrid conditions; both reflective and hybrid feedback also exceeded direct GenAI-supported feedback, with no significant difference between the reflective and hybrid conditions. This supplementary pattern indicates that the advantages of the more agentic conditions were accompanied by more substantive rather than merely surface-level revision. Detailed estimates are reported in Table 4 and Supplementary Table S2.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable 4. Multilevel model predicting four-dimension argument-quality gain\u003c/p\u003e\n\u003ch3\u003ePanel A. Estimated marginal means by condition\u003c/h3\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCondition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAdjusted M\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2.84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[2.55, 3.13]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3.41\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[3.14, 3.68]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e4.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[3.80, 4.34]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e4.31\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[4.05, 4.57]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003ePanel B. Pairwise contrasts\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eContrast\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eb\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCohen\u0026rsquo;s d\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI for b\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.57\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.002\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.22, 0.92]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.54\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.88, 1.58]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.47\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.64\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[1.12, 1.82]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.66\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.33, 0.99]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.57, 1.23]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Reflective GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.16\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.138\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.11\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.08, 0.56]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003ePanel C. Overall model summary\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eEffect\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTest statistic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback condition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 43.8) = 18.76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote.\u003c/em\u003e The model was adjusted for prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, prior GenAI experience, gender, science domain, and institution fixed effects, with section included as a random intercept. Four-dimension argument-quality gain scores were computed as change scores from initial draft to revised submission on the four common content dimensions shared across draft stages and were averaged across the three intervention cycles. Higher scores indicate greater improvement in common-content scientific argument quality from draft to revision. The separate strategic-quality-of-revision indicator was not included in pre/post gain estimation. Holm-adjusted p values are reported for pairwise contrasts.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5.3. Effects on conceptual learning and delayed AI-free transfer\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo address RQ2, separate linear mixed-effects models were estimated for conceptual learning and delayed AI-free transfer, with section included as a random intercept and institution entered as a fixed effect. Feedback condition significantly predicted both conceptual learning, \u003cem\u003eF\u003c/em\u003e(3, 43.5) = 7.62, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001, and delayed AI-free transfer, \u003cem\u003eF\u003c/em\u003e(3, 43.9) = 12.03, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001.\u003c/p\u003e\n\u003cp\u003eFor conceptual learning, the pattern favored the more agentic conditions. Direct GenAI-supported feedback did not differ significantly from peer feedback, whereas reflective and hybrid feedback outperformed peer feedback. Hybrid feedback also significantly outperformed direct GenAI-supported feedback, while the reflective\u0026ndash;direct contrast was not statistically significant after adjustment, and the hybrid\u0026ndash;reflective contrast was also not significant.\u003c/p\u003e\n\u003cp\u003eFor delayed AI-free transfer, the pattern was clearer. Direct GenAI-supported feedback did not outperform peer feedback, whereas both reflective and hybrid feedback significantly outperformed both peer feedback and direct GenAI-supported feedback. The difference between reflective and hybrid feedback was not statistically significant. Taken together, these findings indicate that the more agentic feedback designs were associated with stronger distal learning outcomes, particularly AI-independent transfer. Detailed estimates and pairwise contrasts are reported in Table 5.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 5. Mixed-effects models predicting conceptual learning and delayed AI-free transfer\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePanel A. Conceptual learning\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCondition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAdjusted M\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.20, 0.04]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.01\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.11, 0.13]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.07, 0.31]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.14, 0.38]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eContrast\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eb\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCohen\u0026rsquo;s d\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI for b\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.314\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.09, 0.27]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.006\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.08, 0.46]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.32\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.16, 0.52]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.072\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.02, 0.38]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.014\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.05, 0.45]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Reflective GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.438\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.11, 0.25]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eEffect\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTest statistic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback condition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 43.5) = 7.62\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003ePanel B. Delayed AI-free transfer\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCondition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAdjusted M\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e11.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[11.41, 12.23]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e11.46\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[11.07, 11.85]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e12.61\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[12.24, 12.98]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e12.88\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[12.49, 13.27]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eContrast\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eb\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCohen\u0026rsquo;s d\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI for b\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.36\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.149\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.85, 0.13]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.79\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.003\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.28, 1.30]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.36\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.57, 1.55]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.64, 1.66]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.48\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.89, 1.95]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Reflective GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.28\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.327\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.28, 0.82]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eEffect\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTest statistic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback condition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 43.9) = 12.03\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote.\u003c/em\u003e Models were adjusted for prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, prior GenAI experience, gender, science domain, and institution fixed effects, with section included as a random intercept. Conceptual learning scores were standardized within discipline before pooled analysis. Higher delayed-transfer scores indicate stronger AI-independent scientific argumentation on the novel task. Holm-adjusted p values are reported for pairwise contrasts.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5.4. Effects on feedback uptake and self-regulated learning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo address RQ3, separate linear mixed-effects models were estimated for feedback uptake and self-regulated learning during revision, with section included as a random intercept and institution entered as a fixed effect. Feedback condition significantly predicted both feedback uptake, \u003cem\u003eF\u003c/em\u003e(3, 43.6) = 16.28, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001, and self-regulated learning, \u003cem\u003eF\u003c/em\u003e(3, 43.2) = 11.91, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001.\u003c/p\u003e\n\u003cp\u003eFor both process variables, the same overall pattern emerged. Direct GenAI-supported feedback did not differ significantly from peer feedback, whereas reflective and hybrid feedback yielded significantly higher feedback uptake and self-regulated learning than both peer feedback and direct GenAI-supported feedback. No significant difference emerged between reflective and hybrid feedback on either process variable. These results are consistent with H1a and H2a, and with the nondirectional expectations stated in H1b and H2b. Full estimates are presented in Table 6.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 6. Mixed-effects models predicting feedback uptake and self-regulated learning during revision\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePanel A. Feedback uptake\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCondition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAdjusted M\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.16, 0.04]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.02\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.12, 0.08]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.14, 0.34]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.31\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.21, 0.41]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eContrast\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eb\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.04\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.563\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.30\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Reflective GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.314\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eEffect\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTest statistic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback condition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 43.6) = 16.28\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003ePanel B. Self-regulated learning during revision\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCondition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAdjusted M\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003ePeer feedback only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.03\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.13, 0.07]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[-0.19, 0.01]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.07, 0.27]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e[0.16, 0.36]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eContrast\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eb\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eDirect GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e-0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.401\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.015\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Peer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.002\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid \u0026ndash; Reflective GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e.226\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eEffect\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTest statistic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eFeedback condition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eF(3, 43.2) = 11.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote.\u003c/em\u003e Both process variables were standardized composite indices. Models were adjusted for prior knowledge, baseline scientific argumentation, prior achievement, feedback literacy, prior GenAI experience, gender, science domain, and institution fixed effects, with section included as a random intercept. Higher scores indicate stronger engagement with feedback and greater self-regulation during revision. Holm-adjusted p values are reported for pairwise contrasts.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5.5. Mediation analyses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo address RQ4, 2-1-1 multilevel mediation models were estimated to examine whether feedback uptake and self-regulated learning during revision mediated the relationship between feedback condition and four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer. Because the theoretically focal contrasts concerned the more agentic feedback designs, the mediation analyses compared the reflective and hybrid conditions primarily against direct GenAI-supported feedback. Standardized coefficients are reported in Table 7.\u003c/p\u003e\n\u003cp\u003eAcross outcomes, the mediation results showed that the advantages of reflective and hybrid feedback relative to direct GenAI-supported feedback were explained in part by stronger feedback uptake and self-regulated learning. For four-dimension argument-quality gain, both reflective and hybrid feedback showed significant indirect effects through feedback uptake, self-regulated learning, and the serial pathway from uptake to self-regulated revision. For conceptual learning, the strongest indirect effects operated through self-regulated learning and the serial pathway, whereas the pathway through feedback uptake alone did not reach significance. For delayed AI-free transfer, all three indirect pathways were statistically significant for both reflective and hybrid feedback. Taken together, these findings indicate that the more agentic feedback designs improved learning outcomes not only by changing the feedback resources available to students, but also by changing how students engaged with and acted on those resources. Detailed direct, indirect, and total effects are reported in Table 7.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 7. Direct, indirect, and total effects of feedback condition on outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eContrast\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eOutcome\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eDirect effect, \u0026beta; [95% CI]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTotal indirect effect, \u0026beta; [95% CI]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eVia feedback uptake, \u0026beta; [95% CI]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eVia SRL, \u0026beta; [95% CI]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eVia uptake \u0026rarr; SRL, \u0026beta; [95% CI]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTotal effect, \u0026beta; [95% CI]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI vs. Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFour-dimension argument-quality gain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.36 [0.11, 0.61]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.30 [0.18, 0.44]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18 [0.08, 0.31]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07 [0.02, 0.14]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05 [0.01, 0.10]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.66 [0.41, 0.91]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid vs. Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFour-dimension argument-quality gain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.52 [0.27, 0.77]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.38 [0.25, 0.52]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.22 [0.11, 0.35]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.09 [0.03, 0.16]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.07 [0.02, 0.13]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.90 [0.65, 1.15]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI vs. Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eConceptual learning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05 [-0.04, 0.14]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.13 [0.06, 0.21]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.03 [-0.01, 0.08]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.06 [0.02, 0.12]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.04 [0.01, 0.08]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18 [0.08, 0.28]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid vs. Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eConceptual learning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08 [-0.02, 0.18]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.17 [0.09, 0.25]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.04 [-0.00, 0.09]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.08 [0.03, 0.14]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.05 [0.02, 0.10]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.25 [0.15, 0.35]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective GenAI vs. Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDelayed AI-free transfer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.65 [0.30, 1.00]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.50 [0.31, 0.72]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.21 [0.09, 0.36]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.18 [0.07, 0.32]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.11 [0.04, 0.19]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.15 [0.80, 1.50]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid vs. Direct GenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDelayed AI-free transfer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.80 [0.44, 1.16]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.62 [0.42, 0.84]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.26 [0.12, 0.41]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.22 [0.10, 0.36]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.14 [0.06, 0.23]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1.42 [1.06, 1.78]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote.\u003c/em\u003e Standardized coefficients are reported for all focal contrasts. Indirect-effect confidence intervals were derived from Monte Carlo simulation; direct- and total-effect confidence intervals were derived from model-based standard errors. Confidence intervals that do not cross zero indicate statistically reliable effects. For the primary immediate outcome, mediation models used four-dimension argument-quality gain rather than the separate revised-draft strategic-quality-of-revision indicator, because gain estimation required directly comparable scoring across draft stages.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5.6. Qualitative findings on feedback use and epistemic agency\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe qualitative findings were used to clarify how students and instructors experienced the feedback designs and to interpret the mechanisms underlying the quantitative patterns. Although both student and instructor interviews informed the thematic analysis, the themes reported below focus primarily on student accounts of engaging with different feedback designs. Instructor interviews largely corroborated these patterns, particularly with respect to the trade-off between efficiency and ownership and the greater deliberation observed in the reflective and hybrid conditions. Representative quotations are provided in Table 8.\u003c/p\u003e\n\u003cp\u003eThe first theme, AI as a fast but sometimes over-directive feedback source, captured students\u0026rsquo; descriptions of GenAI feedback as immediate, accessible, and efficient for identifying obvious weaknesses in their drafts. Students in the direct GenAI-supported feedback condition frequently valued the speed and specificity of the feedback but also described a risk of following AI-suggested structures too readily, particularly when they had not first articulated their own judgment about the task.\u003c/p\u003e\n\u003cp\u003eThe second theme, reflective prompting as support for evaluative judgment, captured students\u0026rsquo; reports that the self-evaluation step changed how they read AI-generated feedback. Rather than treating the feedback as an answer key, students described using it as a comparison point against their own diagnosis of strengths, weaknesses, and uncertainties. This pattern helps explain why the reflective condition was associated with stronger feedback uptake and self-regulated learning than direct GenAI-supported feedback alone.\u003c/p\u003e\n\u003cp\u003eThe third theme, hybrid feedback as comparison across perspectives, reflected students\u0026rsquo; descriptions of the combined value of self-, peer-, and AI-based feedback. Participants often noted that these sources highlighted different weaknesses in the same draft, thereby supporting more deliberate judgment about which revisions to prioritize. Students in the hybrid condition also reported that the revision memo encouraged them to justify rather than simply enact revision decisions.\u003c/p\u003e\n\u003cp\u003eThe fourth theme, tensions between efficiency and ownership of revision, described students\u0026rsquo; sense that AI-enabled feedback reduced the effort required to generate revision ideas while sometimes weakening their sense of authorship over the final text. This tension was most visible in AI-supported conditions and was described by some instructors as a pedagogical challenge: the same features that made AI feedback attractive could also reduce the extent to which students saw revision as their own evaluative work.\u003c/p\u003e\n\u003cp\u003eTaken together, the qualitative findings help explain why the reflective and hybrid conditions were associated with stronger feedback uptake and self-regulated learning than the direct GenAI-supported feedback condition. They also reinforce the broader interpretation that more agentic feedback designs preserve student ownership over revision by requiring comparison, selection, and justification rather than simple implementation of external critique.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 8. Qualitative themes and representative quotations\u003c/strong\u003e\u003c/p\u003e\n\u003ctable\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTheme\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eDescription\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eRepresentative quotation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eTypical conditions\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eAI as a fast but sometimes over-directive feedback source\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eStudents valued the speed and clarity of AI-generated critique but sometimes felt that it encouraged formulaic revision.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026ldquo;It helped me see what was missing almost immediately, but sometimes it felt as though I was borrowing its structure instead of testing my own reasoning.\u0026rdquo; (S07, direct GenAI-supported feedback, biology)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eMostly direct GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eReflective prompting as support for evaluative judgment\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCompleting self-evaluation before reading AI feedback prompted comparison, verification, and more active judgment.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026ldquo;Because I had already written down what I thought was weak, I was not just accepting the AI\u0026rsquo;s comments. I was checking where it agreed with me and where it made me rethink the argument.\u0026rdquo; (S18, reflective GenAI-supported feedback, chemistry)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eMostly reflective GenAI-supported feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHybrid feedback as comparison across perspectives\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eStudents described hybrid feedback as helpful because self, peer, and AI input highlighted different weaknesses in the same draft.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026ldquo;The peer comments showed me what was unclear to a reader, and the AI pointed out where my evidence still was not doing enough work. Seeing both made it easier to decide what actually needed revision.\u0026rdquo; (S24, hybrid feedback, physics)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eMostly hybrid feedback\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eTensions between efficiency and ownership of revision\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eStudents reported a trade-off between faster revision and maintaining a sense of authorship and control over argument quality.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026ldquo;It definitely saved time, but I had to stop and ask whether this was still my explanation or whether I was just polishing the version the tool seemed to prefer.\u0026rdquo; (S31, hybrid feedback, chemistry)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAcross AI-supported conditions\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e5.7. Sensitivity analyses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSeveral sensitivity analyses were conducted to examine the robustness of the main findings. First, the primary models were re-estimated using a per-protocol sample excluding students who completed fewer than two of the three intervention cycles (n = 1,102). The overall pattern of results remained substantively unchanged. Second, models excluding the three lowest-fidelity sections yielded parameter estimates that differed by less than 0.04 standard deviations from the main intention-to-treat analyses. Third, alternative missing-data specifications using full-information maximum likelihood produced conclusions identical to those obtained from multiple imputation.\u003c/p\u003e\n\u003cp\u003eExploratory moderation analyses examined whether the effects of feedback condition varied as a function of baseline feedback literacy and science domain. The interaction between condition and baseline feedback literacy was significant for delayed AI-free transfer, F(3, 1158) = 2.83, p = .037, indicating that the disadvantage of the direct GenAI-supported feedback condition relative to the reflective and hybrid conditions was more pronounced among students with lower initial feedback literacy. Because these moderation tests were exploratory and not the focus of the preregistered confirmatory analyses, they should be interpreted cautiously. By contrast, the interaction between condition and science domain was not significant for any primary or secondary outcome (all p \u0026gt; .10), indicating that the overall pattern of effects was largely consistent across biology, chemistry, and physics. Detailed sensitivity results are reported in the Supplementary Materials.\u003c/p\u003e\n\u003cp\u003eThese additional analyses support the robustness of the main findings while also suggesting that feedback literacy may condition the extent to which students benefit from more direct forms of AI-supported feedback.\u003c/p\u003e"},{"header":"6. Discussion","content":"\u003cdiv id=\"Sec30\" class=\"Section2\"\u003e \u003ch2\u003e6.1. Interpretation of the principal findings\u003c/h2\u003e \u003cp\u003eThe present study examined how alternative GenAI-supported feedback designs shaped revision, learning, and transfer in introductory university science courses. The findings indicate that the educational value of GenAI feedback depended less on AI access per se than on how students were positioned within the feedback process. Direct GenAI-supported feedback improved immediate four-dimension argument-quality gain relative to peer feedback, but reflective and hybrid feedback designs produced stronger outcomes on feedback uptake, self-regulated learning, conceptual learning, and delayed AI-free transfer. The hybrid condition yielded the highest adjusted mean for four-dimension argument-quality gain, whereas both reflective and hybrid designs outperformed direct GenAI-supported feedback on delayed AI-free transfer. Taken together, these findings suggest that GenAI can support rapid improvement in common-content argument quality, but more agentic feedback designs are more effective when the goal is durable learning and AI-independent transfer.\u003c/p\u003e \u003cp\u003eThe overall pattern is important because it differentiates between short-term improvement in argument quality and broader educational value. Direct GenAI-supported feedback appears to be effective when students need rapid, criterion-referenced critique to strengthen a draft in the moment. However, the stronger outcomes associated with reflective and hybrid designs indicate that longer-term conceptual gains and transfer are more likely when feedback processes require students to engage in their own evaluative work. In this respect, the study supports the view that immediate improvement and durable learning should not be treated as interchangeable outcomes in AI-mediated feedback environments.\u003c/p\u003e \u003cp\u003eThe supplementary indicators help sharpen this interpretation. Although the primary immediate outcome was defined as four-dimension argument-quality gain, the more agentic conditions were also associated with more substantive revision and stronger revised-draft strategic quality of revision. This pattern suggests that the advantages of reflective and hybrid feedback were not limited to larger gain scores alone; they were also linked to revision that appeared more purposeful, better integrated, and more epistemically substantive. The distinction matters because it indicates that stronger performance in the more agentic conditions was accompanied by qualitatively stronger revision activity rather than by superficial editing alone.\u003c/p\u003e \u003cp\u003eThe absence of statistically significant differences between reflective and hybrid conditions on several outcomes should also be interpreted cautiously. Rather than demonstrating equivalence in any strong sense, the pattern is consistent with the possibility that requiring learners to externalize an initial evaluative judgment before receiving AI critique is itself a particularly important ingredient of effective GenAI-supported feedback. The additional components included in the hybrid condition may still offer practical and pedagogical benefits, but these advantages were not always statistically separable from those associated with reflective GenAI-supported feedback alone in the present sample.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec31\" class=\"Section2\"\u003e \u003ch2\u003e6.2. Theoretical and empirical implications\u003c/h2\u003e \u003cp\u003eThese findings support the claim that feedback design is a more meaningful analytic unit than feedback source alone. In the present study, direct GenAI-supported feedback provided rapid, criterion-referenced critique, but it did so without requiring learners to externalize an initial judgment about the quality of their own work. By contrast, reflective and hybrid designs required learners to compare, interpret, and justify revision decisions before or alongside engagement with AI critique. The stronger outcomes associated with these more agentic designs are therefore consistent with process-oriented accounts of feedback, which emphasize that learning depends not merely on receiving comments, but on learners\u0026rsquo; active interpretation and use of them (Carless \u0026amp; Boud, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Nicol \u0026amp; Macfarlane-Dick, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2006\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe findings also extend work on self-regulated learning in AI-mediated environments. The mediation analyses suggest that reflective and hybrid designs outperformed direct GenAI-supported feedback partly because they strengthened feedback uptake and self-regulated learning during revision. This pattern is consistent with SRL-oriented perspectives that view feedback as educationally powerful when it becomes part of learners\u0026rsquo; own regulatory activity rather than remaining external advice (Nicol \u0026amp; Macfarlane-Dick, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2006\u003c/span\u003e; Panadero, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). The distinction between four-dimension argument-quality gain and the strategic quality of revision sharpens this interpretation further: the more agentic designs were associated not only with larger gains in common-content argument quality, but also with stronger supplementary indicators of revision substance and revised-draft strategic quality. Their advantages therefore appear to reflect a more productive organization of evaluative activity rather than short-term score improvement alone (Carless \u0026amp; Boud, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Buckingham Shum et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe study also contributes to emerging discussions of epistemic agency in AI-supported learning. Recent work has suggested that the educational use of automated feedback depends on how human and machine roles are distributed within feedback processes, and that AI assistance may weaken student agency when critique is accepted too readily or when difficult judgments are outsourced to the system (Buckingham Shum et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Darvishi et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). The present results are consistent with those concerns: students in the direct GenAI-supported feedback condition described AI as fast and useful, but sometimes overly directive, whereas students in the reflective and hybrid conditions described self-evaluation, peer input, and revision justification as preserving a stronger sense of ownership. In this respect, the findings add empirical weight to the argument that the pedagogical value of GenAI depends not only on what the tool can generate, but also on whether surrounding feedback designs amplify or displace students\u0026rsquo; evaluative responsibility (Buckingham Shum et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Darvishi et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). They also align with prior work showing that AI- and peer-generated feedback offer partially complementary affordances (Banihashem et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), that AI-supported peer feedback is most effective when collaborative argumentation is scaffolded (Chang et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2026\u003c/span\u003e), and that GenAI in science learning is most educationally useful when it functions as a dialogic partner rather than an authoritative answer source (Tang \u0026amp; Putra, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). More broadly, the study adds nuance to recent syntheses of GenAI in higher education by suggesting that apparent benefits are highly design-sensitive and that concerns about authorship, evidence of learning, and learner responsibility are not only matters of policy or integrity, but also matters of feedback design (Wang \u0026amp; Fan, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2025\u003c/span\u003e; Xia et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec32\" class=\"Section2\"\u003e \u003ch2\u003e6.3. Practical implications for higher education digital learning\u003c/h2\u003e \u003cp\u003eThe findings have several practical implications for higher education digital learning. First, GenAI feedback should not be treated as pedagogically neutral or self-sufficient. When the goal is quick draft improvement, direct GenAI-supported feedback may be useful because it provides rapid, criterion-referenced critique with relatively low orchestration demands. However, when the goal is deeper learning, stronger reasoning, or AI-independent transfer, instructors should not present AI critique as the first or only evaluative input. The results instead suggest that AI critique should be preceded by a brief, structured self-evaluation that asks students to identify strengths, weaknesses, uncertainties, and revision priorities before viewing external feedback.\u003c/p\u003e \u003cp\u003eSecond, hybrid feedback designs appear particularly well suited to complex reasoning tasks in which students must compare perspectives, weigh evidence, and justify revision decisions. In such contexts, peer feedback can be scaffolded through a structured template, while AI critique can be positioned as a third perspective rather than a final answer. A short revision rationale or memo may further help students explain why particular comments were accepted, rejected, or modified. More broadly, the results imply that GenAI feedback systems should be designed around comparison, selection, and justification rather than passive implementation of suggestions. In practical terms, this means constraining regeneration, making rubric criteria visible, and prompting students to record at least one revision decision in their own words. At the institutional level, the study suggests that decisions about GenAI adoption should not be made only in terms of scalability or efficiency, but also in terms of whether digital-learning designs preserve student agency, evaluative judgment, and ownership of learning when AI tools are embedded in routine coursework.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec33\" class=\"Section2\"\u003e \u003ch2\u003e6.4. Limitations\u003c/h2\u003e \u003cp\u003eSeveral limitations should be considered when interpreting the findings. First, the reflective and hybrid conditions involved greater amounts of guided evaluative activity, meaning that feedback-source configuration and reflective workload were not fully separable. Second, students in AI-supported conditions received additional system training, which may have increased condition-specific preparation time relative to the peer-feedback condition. Third, the study was conducted in first-year university science courses, and the extent to which the findings generalize to other disciplines, levels of study, or institutional contexts remains uncertain.\u003c/p\u003e \u003cp\u003eFourth, although AI-free transfer tasks were completed without access to the study\u0026rsquo;s GenAI system in supervised settings, the design cannot eliminate broader concerns about students\u0026rsquo; external familiarity with AI beyond the intervention context. Fifth, the analytic sample may reflect a degree of consent-based selection, because instructional participation was compulsory whereas research participation was voluntary. Sixth, although peer feedback was structured through a common template, peer-comment quality was not independently scored as a separate explanatory variable, making it difficult to determine the extent to which condition differences may partly reflect variation in the quality of peer input. Seventh, the participating institutions were modeled as fixed effects because only four institutions were included. Finally, the moderation analyses were exploratory and should therefore be interpreted with caution.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec34\" class=\"Section2\"\u003e \u003ch2\u003e6.5. Future research directions\u003c/h2\u003e \u003cp\u003eFuture research should extend this work in several directions. First, studies should examine the minimum effective amount of self-evaluation required before AI critique, since the present findings suggest that prior self-evaluation matters but do not establish whether shorter or lighter reflective prompts would produce comparable benefits. Second, future work should compare teacher-plus-AI, peer-plus-AI, and hybrid self\u0026ndash;peer\u0026ndash;AI designs more directly in order to clarify which combinations are most effective for different types of higher education tasks.\u003c/p\u003e \u003cp\u003eThird, longer-duration studies are needed to determine whether the advantages of reflective and hybrid designs persist across a semester or academic year and whether repeated cycles of AI-supported revision produce cumulative gains in feedback literacy, self-regulated learning, or AI-independent transfer. Finally, further qualitative and mixed-method research should examine how students experience responsibility, authorship, and evaluative judgment in AI-supported feedback environments across disciplines and institutional contexts.\u003c/p\u003e \u003c/div\u003e"},{"header":"7. Conclusion","content":"\u003cp\u003eThis study shows that the pedagogical value of GenAI feedback in higher education depends less on the availability of AI-generated critique than on how feedback processes are designed. In the present study, direct GenAI-supported feedback supported immediate improvement in four-dimension argument-quality gain, but reflective and hybrid designs were more consistently associated with stronger feedback uptake, self-regulated learning during revision, conceptual learning, and AI-independent transfer. These findings reinforce the view that feedback is educationally valuable not simply when it is delivered efficiently, but when learners are positioned to interpret, compare, and act on critique as active evaluative agents (Carless \u0026amp; Boud, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Nicol \u0026amp; Macfarlane-Dick, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2006\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eMore broadly, the findings suggest that the benefits of GenAI in higher education are highly design-sensitive. When AI feedback is organized primarily around speed and efficiency, it may support short-term improvement without necessarily strengthening the broader learning processes that underlie durable understanding and transfer. By contrast, when AI-supported feedback is embedded in reflective and comparative designs that preserve student judgment, self-regulation, and ownership, it is more likely to support meaningful learning (Buckingham Shum et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Darvishi et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Wang \u0026amp; Fan, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). For higher education digital learning, the central implication is therefore not whether students can access GenAI feedback, but how GenAI feedback is pedagogically organized so that human evaluative work remains at the center (Buckingham Shum et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Xia et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analyzed during the current study are not publicly available because they include de-identified student performance, survey, interview, and platform-trace data collected under institutional ethics approvals and data-governance restrictions. De-identified data are available from the corresponding author on reasonable request, subject to institutional and ethical requirements.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author declares that there are no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors’ contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author was solely responsible for the conception and design of the study, the development of the intervention materials, data collection, data analysis, interpretation of the findings, and preparation of the manuscript. The author read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBanihashem, S. K., Kerman, N. T., Noroozi, O., Moon, J., \u0026amp; Drachsler, H. (2024). Feedback sources in essay writing: Peer-generated or AI-generated feedback? \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s41239-024-00455-4\u003c/span\u003e\u003cspan address=\"10.1186/s41239-024-00455-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Article 23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBraun, V., \u0026amp; Clarke, V. (2006). Using thematic analysis in psychology. \u003cem\u003eQualitative Research in Psychology\u003c/em\u003e, \u003cem\u003e3\u003c/em\u003e(2), 77\u0026ndash;101. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1191/1478088706qp063oa\u003c/span\u003e\u003cspan address=\"10.1191/1478088706qp063oa\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuckingham Shum, S., Lim, L. A., Boud, D., Bearman, M., \u0026amp; Dawson, P. (2023). A comparative analysis of the skilled use of automated feedback tools through the lens of teacher feedback literacy. \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e., Article 40. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s41239-023-00410-9\u003c/span\u003e\u003cspan address=\"10.1186/s41239-023-00410-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarless, D., \u0026amp; Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. \u003cem\u003eAssessment \u0026amp; Evaluation in Higher Education\u003c/em\u003e, \u003cem\u003e43\u003c/em\u003e(8), 1315\u0026ndash;1325. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/02602938.2018.1463354\u003c/span\u003e\u003cspan address=\"10.1080/02602938.2018.1463354\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang, Y., Liu, Q., Lu, Y., \u0026amp; Miao, E. (2026). Leveraging generative AI to facilitate peer feedback in collaborative argumentation learning. \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e23\u003c/em\u003e., Article 10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s41239-026-00586-w\u003c/span\u003e\u003cspan address=\"10.1186/s41239-026-00586-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDarvishi, A., Khosravi, H., Sadiq, S. W., Gašević, D., \u0026amp; Siemens, G. (2024). Impact of AI assistance on student agency. \u003cem\u003eComputers \u0026amp; Education\u003c/em\u003e, \u003cem\u003e210\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.compedu.2023.104967\u003c/span\u003e\u003cspan address=\"10.1016/j.compedu.2023.104967\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Article 104967.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDriver, R., Newton, P., \u0026amp; Osborne, J. (2000). Establishing the norms of scientific argumentation in classrooms. \u003cem\u003eScience Education\u003c/em\u003e, \u003cem\u003e84\u003c/em\u003e(3), 287\u0026ndash;312. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/(SICI)1098-237X(200005)84:3%3C287::AID-SCE1%3E3.0.CO;2-A\u003c/span\u003e\u003cspan address=\"10.1002/(SICI)1098-237X(200005)84:3%3C287::AID-SCE1%3E3.0.CO;2-A\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHattie, J., \u0026amp; Timperley, H. (2007). The power of feedback. \u003cem\u003eReview of Educational Research\u003c/em\u003e, \u003cem\u003e77\u003c/em\u003e(1), 81\u0026ndash;112. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3102/003465430298487\u003c/span\u003e\u003cspan address=\"10.3102/003465430298487\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMolloy, E., Boud, D., \u0026amp; Henderson, M. (2020). Developing a learning-centred framework for feedback literacy. \u003cem\u003eAssessment \u0026amp; Evaluation in Higher Education\u003c/em\u003e, \u003cem\u003e45\u003c/em\u003e(4), 527\u0026ndash;540. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/02602938.2019.1667955\u003c/span\u003e\u003cspan address=\"10.1080/02602938.2019.1667955\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNicol, D. (2021). The power of internal feedback: Exploiting natural comparison processes. \u003cem\u003eAssessment \u0026amp; Evaluation in Higher Education\u003c/em\u003e, \u003cem\u003e46\u003c/em\u003e(5), 756\u0026ndash;778. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/02602938.2020.1823314\u003c/span\u003e\u003cspan address=\"10.1080/02602938.2020.1823314\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNicol, D. J., \u0026amp; Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. \u003cem\u003eStudies in Higher Education\u003c/em\u003e, \u003cem\u003e31\u003c/em\u003e(2), 199\u0026ndash;218. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/03075070600572090\u003c/span\u003e\u003cspan address=\"10.1080/03075070600572090\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOsborne, J., Erduran, S., \u0026amp; Simon, S. (2004). Enhancing the quality of argumentation in school science. \u003cem\u003eJournal of Research in Science Teaching\u003c/em\u003e, \u003cem\u003e41\u003c/em\u003e(10), 994\u0026ndash;1020. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/tea.20035\u003c/span\u003e\u003cspan address=\"10.1002/tea.20035\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePanadero, E. (2017). A review of self-regulated learning: Six models and four directions for research. \u003cem\u003eFrontiers in Psychology\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e., Article 422. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpsyg.2017.00422\u003c/span\u003e\u003cspan address=\"10.3389/fpsyg.2017.00422\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePanadero, E., Jonsson, A., \u0026amp; Botella, J. (2017). Effects of self-assessment on self-regulated learning and self-efficacy: Four meta-analyses. \u003cem\u003eEducational Research Review\u003c/em\u003e, \u003cem\u003e22\u003c/em\u003e, 74\u0026ndash;98. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.edurev.2017.08.004\u003c/span\u003e\u003cspan address=\"10.1016/j.edurev.2017.08.004\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePintrich, P. R., Smith, D. A. F., Garcia, T., \u0026amp; McKeachie, W. J. (1991). \u003cem\u003eA manual for the use of the Motivated Strategies for Learning Questionnaire (MSLQ)\u003c/em\u003e (NCRIPTAL Technical Report No. 91-B-004). National Center for Research to Improve Postsecondary Teaching and Learning, University of Michigan. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://files.eric.ed.gov/fulltext/ED338122.pdf\u003c/span\u003e\u003cspan address=\"https://files.eric.ed.gov/fulltext/ED338122.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTai, J., Ajjawi, R., Boud, D., Dawson, P., \u0026amp; Panadero, E. (2018). Developing evaluative judgement: Enabling students to make decisions about the quality of work. \u003cem\u003eHigher Education\u003c/em\u003e, \u003cem\u003e76\u003c/em\u003e(3), 467\u0026ndash;481. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10734-017-0220-3\u003c/span\u003e\u003cspan address=\"10.1007/s10734-017-0220-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang, K. S., \u0026amp; Putra, G. B. S. (2025). Generative AI as a dialogic partner: Enhancing multiple perspectives, reasoning, and argumentation in science education with customized chatbots. \u003cem\u003eJournal of Science Education and Technology\u003c/em\u003e, \u003cem\u003e35\u003c/em\u003e, 128\u0026ndash;140. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10956-025-10240-1\u003c/span\u003e\u003cspan address=\"10.1007/s10956-025-10240-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, J., \u0026amp; Fan, W. (2025). The effect of ChatGPT on students\u0026rsquo; learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. \u003cem\u003eHumanities and Social Sciences Communications\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e, 621. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1057/s41599-025-04787-y\u003c/span\u003e\u003cspan address=\"10.1057/s41599-025-04787-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeidlich, J., Jivet, I., Woitt, S., Orhan G\u0026ouml;ks\u0026uuml;n, D., Kraus, J., \u0026amp; Drachsler, H. (2025). The student feedback literacy instrument (SFLI): Multilingual validation and introduction of a short-form version. \u003cem\u003eAssessment \u0026amp; Evaluation in Higher Education\u003c/em\u003e, \u003cem\u003e50\u003c/em\u003e(5), 677\u0026ndash;693. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/02602938.2025.2451729\u003c/span\u003e\u003cspan address=\"10.1080/02602938.2025.2451729\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWoitt, S., Weidlich, J., Jivet, I., Orhan G\u0026ouml;ks\u0026uuml;n, D., Drachsler, H., \u0026amp; Kalz, M. (2025). Students\u0026rsquo; feedback literacy in higher education: An initial scale validation study. \u003cem\u003eTeaching in Higher Education\u003c/em\u003e, \u003cem\u003e30\u003c/em\u003e(1), 257\u0026ndash;276. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/13562517.2023.2263838\u003c/span\u003e\u003cspan address=\"10.1080/13562517.2023.2263838\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXia, Q., Weng, X., Ouyang, F., Lin, T. J., \u0026amp; Chiu, T. K. F. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. \u003cem\u003eInternational Journal of Educational Technology in Higher Education\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e., Article 40. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s41239-024-00468-z\u003c/span\u003e\u003cspan address=\"10.1186/s41239-024-00468-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu, M., Liu, O. L., \u0026amp; Lee, H. S. (2020). The effect of automated feedback on revision behavior and learning gains in formative assessment of scientific argument writing. \u003cem\u003eComputers \u0026amp; Education\u003c/em\u003e, \u003cem\u003e143\u003c/em\u003e, 103668. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.compedu.2019.103668\u003c/span\u003e\u003cspan address=\"10.1016/j.compedu.2019.103668\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Generative artificial intelligence, feedback design, higher education, digital learning, scientific argumentation, self-regulated learning, epistemic agency","lastPublishedDoi":"10.21203/rs.3.rs-9396658/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9396658/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eGenerative artificial intelligence (GenAI) is increasingly used for feedback in higher education, yet evidence remains limited on how alternative human\u0026ndash;AI feedback designs shape learning processes and durable outcomes. This study addresses that gap through a multisite, cluster-randomized, longitudinal field experiment comparing four feedback designs in introductory university science courses: peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and a hybrid design combining self-evaluation, peer feedback, and GenAI critique. The analytic sample comprised 1,176 first-year undergraduate students from 48 course sections across four universities and three science domains. Primary and secondary outcomes were four-dimension argument-quality gain, conceptual learning, and delayed AI-free transfer; feedback uptake and self-regulated learning during revision were modeled as process mediators. Direct GenAI-supported feedback improved immediate argument-quality gain relative to peer feedback, but reflective and hybrid designs produced stronger outcomes on feedback uptake, self-regulated learning, conceptual learning, and delayed AI-free transfer. The hybrid condition yielded the highest adjusted mean for immediate argument-quality gain, whereas both reflective and hybrid conditions outperformed direct GenAI-supported feedback on delayed AI-free transfer. Multilevel mediation analyses indicated that feedback uptake and self-regulated learning partially explained these advantages. By combining four feedback designs, process mediation, and delayed AI-free transfer in a multisite field experiment, the study shows that the educational value of GenAI in higher education depends less on AI access per se than on whether feedback environments preserve student agency, evaluative judgment, and ownership during revision.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e","manuscriptTitle":"Human-Centered GenAI Feedback Design in Higher Education: A Multisite Experiment on Direct, Reflective, and Hybrid Approaches to Scientific Argumentation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-29 17:22:11","doi":"10.21203/rs.3.rs-9396658/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9e6cad58-482b-46ba-95a0-9ae35fc9fda4","owner":[],"postedDate":"April 29th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-19T03:08:24+00:00","index":15,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-01T09:51:40+00:00","index":14,"fulltext":""},{"type":"reviewerAgreed","content":"280979404554521192700454847334885361072","date":"2026-04-30T16:14:16+00:00","index":13,"fulltext":""},{"type":"reviewerAgreed","content":"138740899996747455357259925836672177648","date":"2026-04-30T09:37:13+00:00","index":12,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-29T17:22:12+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-29 17:22:11","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9396658","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9396658","identity":"rs-9396658","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00