Uncovering Immersive Competence as a Hidden Bias in VR-Based Clinical Assessment – A Randomized Controlled Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Uncovering Immersive Competence as a Hidden Bias in VR-Based Clinical Assessment – A Randomized Controlled Study Jan Schaal, Tobias Leutritz, Marco Lindner, Alexander Zamzow, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7660457/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 09 Mar, 2026 Read the published version in npj Digital Medicine → Version 1 posted 9 You are reading this latest preprint version Abstract Virtual reality (VR) is increasingly used for assessment in educational and clinical settings. However, users’ immersive competence (IC) - the ability to navigate and operate VR systems - may introduce bias unrelated to clinical skills or patient functioning, a relationship that remains unexplored. In this randomized controlled trial, 88 medical students received either general IC training, general plus specific IC training, or no structured training before completing a VR-based assessment scenario. Multimodal outcome data were collected, including physiological stress markers (electrodermal activity), cognitive-load ratings (NASA-TLX), procedural efficiency, and self-reported usability barriers. Specific IC training improved performance compared with control (28.3%±10.3% vs. 21.2%±10.8%, p = .010, d = 0.67), partly mediated by procedural efficiency (η 2 = 0.26) and increased cognitive load (η 2 = 0.15). Prior experience with 3D applications was unrelated to performance in the control group (ρ = 0.165, p = .383) but significantly associated with higher performance in the specific training group (ρ = 0.387, p = .034). Participants reported high enjoyment, although interface barriers distracted untrained users. These findings indicate that IC is a causal, modifiable factor in VR-based assessments and should be accounted for to ensure fair and valid evaluations. Health sciences/Health care Health sciences/Medical research Biological sciences/Neuroscience Biological sciences/Psychology Social science/Psychology Figures Figure 1 Figure 2 Figure 3 Introduction Virtual Reality (VR) is increasingly used not only in medical education 1 but also in clinical diagnostics and therapy 2 . In immersive environments, users can engage in realistic and standardized simulations to practice or assess professional skills 1 , 3 , monitor cognitive and motor functions 4 – 6 , or receive psychological or rehabilitative interventions 7 – 11 . However, as these applications gain momentum in high-stakes contexts, questions of validity and equitable access become critical. VR systems are neither standardized nor always intuitive, leading to substantial variations in functionality and usability 12 , 13 . Learners and patients alike must master abstract interaction metaphors, interpret spatial feedback, and operate handheld controllers 12 – 15 . This creates a hidden source of variance, where prior informal digital experience (e.g., video gaming) can unfairly advantage certain users 16 – 19 . Yet, the underlying ability that mediates this effect has remained largely underdefined. Recent work has introduced the concept of Immersive Competence (IC) - the user’s familiarity and skill in operating VR systems 17 , 18 , 20 - as a potentially construct-irrelevant factor 21 influencing performance. From a psychometric perspective, this threatens the validity of VR-based assessments, particularly the extrapolation inference 22 , by conflating technological fluency with domain-specific ability (e.g. clinical performance). Despite this, none of the 26 VR-based assessments in health professions education listed in a recent systematic review accounted for IC as a potential confounder 3 . Likewise, clinical VR assessments in neurology 4 , 23 , psychiatry 5 , 6 , 9 , 24 , or geriatrics 25 have almost exclusively focused on task outcomes, while largely overlooking patients’ ability to operate these immersive systems. However, without accounting for IC, VR-based evaluations risk perpetuating unseen performance biases by disadvantaging less digitally experienced learners or patients and undermining fairness and validity. To address this gap, recent tools have been developed to measure and train IC in a standardized manner, covering general interaction categories such as manipulation, navigation, and system control 17 , 18 and to train and assess domain-specific IC directly within a target VR environment 18 . Yet, empirical studies demonstrating whether IC causally influences clinical performance - and whether training can mitigate this effect - remain lacking. The present study therefore investigates whether training of IC systematically affects performance in a VR-based clinical assessment. We hypothesized that (H1) targeted IC training would enhance clinical performance compared to no training, (H2) this effect would be mediated by increases in procedural efficiency (H2a) and/or reductions in cognitive load (H2b), (H3) IC training would mitigate performance disparities associated with prior 3D or VR application experience, and (H4) trained participants would report less distraction by usability barriers and a more positive evaluation of the VR-based assessment. Findings aim to inform the design of equitable and valid VR-based tools in both education and healthcare – ensuring that performance reflects actual clinical competence in examinees and functional ability in patients rather than technological familiarity. Methods Study design and setting This single-center, randomized controlled trial was conducted between January and August 2025 at the clinical skills laboratory of a German medical faculty. The aim was to test whether IC training influences performance in a VR-based clinical assessment. Eligible participants were advanced undergraduate medical students who had completed the internal medicine written exam (end of 7th semester). Exclusion criteria were known epilepsy or severe simulator sickness. Recruitment occurred via semester mailing lists and enrollment was online. Sessions took place individually, outside curricular teaching hours in a standardized VR laboratory. Participants received a €25 book voucher upon study completion. Randomization and blinding Participants were randomly allocated (1:1:1) to three groups using a computer-generated sequence with fixed block size (n = 9). Due to visible differences in training, blinding for participants was not feasible. However, performance was recorded automatically and rated from anonymized videos, ensuring rater blinding (single-blind design). Sample size Power analysis (one-way ANOVA, α = 0.05, power = 0.80) indicated that a minimum of 84 participants would be needed to detect a moderate-to-large effect size (f = 0.25). To account for potential attrition, 94 were enrolled. Interventions Participants were assigned to one of three study groups: Intervention group 1 (I1): General IC training (abstract VR interactions). Intervention group 2 (I2): General + specific IC training (in-examination environment). Control group (CO): No structured IC training. After providing informed consent and a basic familiarization with VR controls, I1 and I2 completed their respective IC training sessions. All groups then underwent the same VR-based clinical assessment. General IC training General IC training was delivered using the VR Competence App 17 , which includes eight modular interaction challenges (e.g., object selection, teleportation, system control) unrelated to clinical content. Participants received immediate performance feedback and repeated tasks if their score fell below the 75th percentile of a pilot cohort. Training duration was capped at 25 minutes. A full overview of challenges and scoring thresholds is provided in Supplementary Table 1. Specific IC training Specific IC training was conducted in the STEP-VR emergency simulation environment 18 , 26 , identical to the assessment setting. Participants followed an audio-guided sequence of seven procedural tasks (e.g., blood sampling, IV access setup), each comprising multiple time steps. Tasks required no medical knowledge but replicated complex interaction patterns. As in the general IC training, only unsuccessful tasks were repeated. Training was limited to 25 minutes. Task content and time limits are detailed in Supplementary Table 1. VR-based clinical assessment Participants completed a 9-minute VR-based simulation scenario of septic shock due to abscess-forming Crohn’s disease. The scenario focused on implementing the "Sepsis One-Hour Bundle" and initiating guideline-based surgical or interventional measures 27 . Content validity was confirmed by medical educators and the scenario had been used previously in a VR-based OSCE 28 . Before the assessment, participants received a 5-minute orientation to the virtual emergency room and structured case information equivalent to standard OSCE pre-station instructions. The scenario was presented in first-person VR, mirrored to an external screen for supervision and video recording. Outcome measures Primary outcome: Clinical performance The primary outcome of the study was clinical performance in the VR-based assessment. Performance was measured with a previously validated OSCE checklist 28 , adapted to the knowledge level of students in the 7th to 10th semester (Supplementary Table 2). The adapted version included nine items, scored either binarily (1 or 0) or ternarily (1, 0.5, or 0). To ensure that the assessment captured genuine clinical competence rather than short-term memory effects, none of the checklist items were rehearsed during the IC training. Ratings were based on anonymized video recordings and independently conducted by two blinded raters. All items were equally weighted, and a total performance score was calculated for each participant. Secondary outcomes: Procedural efficiency, cognitive load, as well as self-reported usability barriers and acceptance Procedural efficiency was defined as the standardized procedural time (SPT), measured as the interval between the observable initiation and the completion of a procedure. SPT was determined for one previously trained task (intravenous cannulation) and four untrained tasks (blood culture sampling, temperature measurement, abdominal ultrasound, and pneumatic tube use). Since these procedures differed in their inherent time requirements, raw durations were standardized within the sample to Z-scores (mean = 0, SD = 1). For untrained tasks, an aggregate score was calculated as the mean of standardized times; lower SPT values indicated higher procedural efficiency. Cognitive load was assessed both subjectively and objectively. Subjective ratings were obtained using the German version of the unweighted (raw) NASA-TLX 29 , 30 at two time points: after 3 minutes (TLX1, Mental Demand subscale) and immediately after the simulation (TLX2, full scale). Ratings ranged from 0 = lowest to 100 = highest. Objective data were derived from electrodermal activity (EDA) recordings using the Empatica Embrace Pro wristband. Valid recordings were defined as skin conductance levels between 0.05 and 60 µS and skin temperature between 30 and 40°C 31 . Data were processed in R (v4.3.1, wearables package 32 ) following best-practice standards 33 , with median per minute values used for group-comparisons. Finally, self-reported usability barriers and acceptance were evaluated with a post-assessment questionnaire. Items addressed distraction by usability barriers, perceptions of fairness and enjoyment. An open-ended item invited participants to comment on prerequisites for future VR examinations use and responses were subjected to “thematic analysis” 34 . Participant characteristics and experience Participant characteristics and prior experiences were collected, including gender, age and exam performance, operationalized as the mean grade across Internal Medicine, Emergency Medicine, and Anesthesiology according to the German grading system (1 = best, 5 = worst). In addition, participants reported their experience with 3D (computer-based) and VR applications, quantified as the number of hours of use during the past six months. Software and hardware General IC training was conducted using the VR Competence App 17 , developed by the Chair for Human-Computer Interaction at the University of Würzburg. Specific training and assessment scenarios were implemented on the STEP-VR platform 26 , an immersive simulation built on a commercial game engine (Unreal Engine 5.1), developed in collaboration with ThreeDee GmbH (Munich, Germany). Interaction logic of both applications was implemented using native controller input and physics-based object handling: object manipulation used a direct-grab paradigm with collision physics and snap-to-socket behavior for connectors (e.g., infusion lines), while UI interactions employed a menu with raycast selection and confirmation buttons, both supported by haptic feedback and contextual audio cues. Locomotion combined smooth walking for short distances with teleportation for longer distances. The VR setup included an OMEN by HP 17-ck0075ng laptop (Intel Core i7-11800H, NVIDIA GeForce RTX 3070 [8 GB]) and an Oculus Quest III headset with handheld controllers. Applications were rendered at > 90 frames per second with high-quality graphic settings. Statistical analysis Normality of residuals was assessed using the Shapiro-Wilk test. For group comparisons, one-way ANOVA was applied when assumption of normality was met. In cases of deviation, the non-parametric Kruskal-Wallis test was used. Post hoc comparisons of non-normally distributed data were performed using the Mann-Whitney test. For correlation analyses, Pearson’s correlation coefficient was calculated for normally distributed variables, while Spearman’s rank correlation was applied otherwise. To explore interaction effects of covariates, regression analyses were conducted in R (v.4.5.1 with RStudio v. 2024.04.2). Effect sizes were expressed as partial eta-squared ( η 2 ) with corresponding confidence intervals, with thresholds of 0.01, 0.06, and 0.14 interpreted as small, medium and large. Regression equations were reported in the form Y = a + bX, to indicate the magnitude and direction of effects. For cell sizes smaller than five, regression coefficients were not computed 35 . Interrater reliability of checklist scoring was calculated using linear weighted Cohen κ, and group differences in gender distribution were tested using chi-square. EDA data were processed in R (v4.3.1; RRID:SCR_001905) using a specific wearables package 32 . Median per-minute values were computed for each participant and used for group-level comparisons. All other analyses were conducted with GraphPad Prism (v10.1.2; RRID:SCR_002798). Reporting was conducted in accordance with the CONSORT-EHEALTH recommendations (for a completed reporting checklist, see supplementary files) 36 . Ethics and data protection This study was reviewed and approved by the Ethics Committee of the University of Würzburg (Proposal No. 2024-316-ka). Written informed consent was obtained from all participants. All data were processed anonymously and handled in accordance with local data protection regulations. VR interaction data and physiological recordings were stored on encrypted institutional servers and anonymized before analysis. First-person video recordings used for performance rating were de-identified (audio track removed) to ensure participant confidentiality. Participation in the study had no academic consequences for students. Results Participant characteristics A total of 94 medical students were enrolled and randomized across the three study arms (Fig. 1 ). Due to technical problems with the VR hardware/software, missing video recordings and incomplete questionnaires, six datasets were excluded from the final analysis. The remaining 88 participants were representative of the overall population of advanced undergraduate medical students (Table 1 ). Academic performance, measured as the exam score, was slightly better in CO compared to I1 and I2 (p = .020). Overall, participants reported limited prior exposure to 3D applications (mean: 9.36 ± 37.0 hours in the past six months) and minimal prior VR experience (mean: 0.07 ± 0.23 hours in the past six months), consistent with findings from comparable cohorts in medical educational. Table 1 Participant characteristics and experiences across groups. ¹Exam score = mean grade across Internal Medicine, Emergency Medicine, and Anesthesiology. German grading system: 1 = best, 5 = worst. Characteristics and experiences Total (n = 88) I1 (n = 28) I2 (n = 30) CO (n = 30) p-value Gender, n (%): Female Male Diverse 63 (72%) 25 (28%) 0 (0%) 18 (64%) 10 (36%) 0 (0%) 22 (73%) 8 (27%) 0 (0%) 23 (77%) 7 (23%) 0 (0%) .56 Age (years), mean ± SD 24.5 ± 2.8 25.0 ± 3.6 24.9 ± 2.6 23.6 ± 2.0 .169 Exam score¹, mean ± SD 2.14 ± 0.73 2.34 ± 0.83 2.29 ± 0.66 1.83 ± 0.60 .020 3D experience (hours)², mean ± SD 9.36 ± 37.0 12.0 ± 31.6 12.2 ± 54.8 4.0 ± 11.6 .630 VR experience (hours) 2 , mean ± SD 0.07 ± 0.23 0.11 ± 0.28 0.05 ± 0.20 0.07 ± 0.22 .641 ² Refers to the past six months. General and specific IC training effectiveness Both general and specific training modalities resulted in significant gains of IC (Fig. 2 a). Participants who received general IC training (I1 and I2) improved their general IC score from 54.9% to 69.1% (p < .001), mean training time 24.1 ± 2.3 minutes. Participants in the specific IC training group (I2) increased their specific IC score from 44.7% to 86.2% (p < .001), mean training time 21.7 ± 3.3 minutes. Impact of IC training on clinical performance (H1) As the primary outcome, clinical performance in the VR-based emergency scenario differed significantly between groups (Fig. 2 b). Mean performance scores were 19.9% ± 10.6% in I1 (general training), 28.3% ± 10.3% in I2 (general + specific training), and 21.2% ± 10.8% in CO. While performance in I1 did not substantially differ from CO, participants in I2 achieved markedly higher scores than both other groups. This pattern was confirmed by Kruskal-Wallis test (p = .010) and post hoc analyses, indicating that the combined training condition (I2) was the only intervention associated with a clear performance advantage corresponding to a medium-to-large effect size (Cohen’s d = 0.67) compared to CO. Interrater reliability of checklist scoring was consistently high, with a linear weighted Cohen κ of 0.895 across all items. Procedural efficiency as a mediator (H2a) To evaluate the effect of training on procedural efficiency, SPT was analyzed for one trained procedure (intravenous cannulation) and four untrained transfer tasks (blood culture sampling, temperature measurement, abdominal ultrasound, pneumatic tube operation) (Fig. 2 d). For the trained procedure, mean SPT differed significantly between groups (Kruskal–Wallis: p < .001). Participants in I2 demonstrated the shortest SPT (-0.683 ± 0.332), followed by CO (0.260 ± 0.853) and group I1 (0.443 ± 1.181). For the aggregate SPT of the untrained transfer tasks, group differences were less pronounced but still significant (Kruskal-Wallis: p = .017). SPT was − 0.368 ± 0.424 in I2, 0.184 ± 0.662 in I1, and 0.214 ± 0.806 in CO. Across groups, SPT for untrained tasks showed a significant main effect on clinical performance in interaction analysis (p < .01), whereas SPT for trained tasks did not. Students with low SPT values (greater efficiency) in untrained tasks achieved higher clinical performance with an effect size of η 2 = 0.26 [0.08,1.00]. Expressed as regression equation, clinical performance decreased on average by Y = 32.6 – 8.70 for low values of untrained SPT and by Y = 32.6–17.3 for high values. Thus, high SPT values (indicating low efficiency) in untrained tasks were associated with approximately a 50% reduction in clinical performance. Cognitive load as a mediator (H2b) Objective cognitive load, assessed via EDA, did not differ significantly across groups during the VR assessment (Fig. 2 e, left). While there were group differences immediately before the assessment (minutes − 3 to -1) with I1 showing the highest median EDA levels, followed by I2 and CO (Kruskal-Wallis: all p < .002), median EDA values converged during the assessment (minutes 0–10). Transient differences at isolated time points did not remain significant after correction for multiple testing (Kruskal-Wallis after FDR adjustment). However, a trend toward higher EDA in I1 compared to the other groups persisted throughout the assessment period. During the VR assessment (TLX1), subjective NASA-TLX ratings showed a small but significant group difference. Surprisingly, CO reported the lowest subjective cognitive load (I1: 69.4 ± 18.9, I2: 68.0 ± 15.6, CO: 58.6 ± 15.6; Kruskal-Wallis: p = .010). Immediately after the assessment (TLX2), however, subjective cognitive load was similar across groups (I1: 75.8 ± 17.8, I2: 80.9 ± 13.9, CO: 77.3 ± 10.6; Kruskal-Wallis: p = .229) (Fig. 2 e, right). Regarding interaction, EDA did not show significant main effects within or between groups. In contrast, subjective ratings of cognitive load (TLX1) demonstrated a strong main effect on performance (p < .001): Across groups, students with medium NASA-TLX values achieved the best clinical performance, followed by those with high values ( η 2 = 0.15 [0.04,1.00]). Expressed as regression equations, clinical performance increased by Y = 7.80 + 18.8 for medium NASA-TLX values and by Y = 7.80 + 15.1 for high values. A similar, though non-significant pattern was observed for the TLX2 when comparing I2 to CO. Again, students with medium NASA-TLX values showed the highest performance, followed by those with high values ( η 2 = 0.04 [0.00,1.00], regression equation: medium Y = 11 + 16.0, high: Y = 11 + 13.9). Correlation between 3D experience and clinical performance (H3) Values of prior 3D experience were not normally distributed across groups (Shapiro-Wilk test, all p < .001) and overall very low, with 75% of participants reporting a value of 0. Spearman correlations indicated small, non-significant associations in CO (ρ = 0.165, p = .383) and I1 (ρ = 0.162, p = .409), but a moderate, significant correlation in I2 (ρ = 0.387, p = .034) (Fig. 2 c). As the participants had virtually no prior VR experience (Table 1 ), no correlation with clinical performance was observed for this variable (all ρ < 0.100). Self-reported usability barriers and acceptance (H4) Across all groups, enjoyment ratings for the VR scenarios were high (I1 = 4.39 ± 0.69, I2 = 4.53 ± 0.63, CO = 4.14 ± 0.87; Kruskal-Wallis p = .725), and perception of future applicability of VR-based examinations were predominantly positive (I1 = 3.11 ± 1.13, I2 = 3.57 ± 1.07, CO = 3.21 ± 1.24; Kruskal-Wallis: p = .284). However, participants in I1 and CO reported significantly more distraction from clinical content due to interface barriers compared to I2 (I1 = 3.00 ± 1.25, I2 = 2.47 ± 1.17, CO = 3.38 ± 1.37; Kruskal-Wallis: p = .024). Interestingly, distraction scores in CO were correlated negatively with prior 3D experience (r = -0.39, p = .035). This association was weaker and no longer statistically significant in I2 (r = -0.28, p = .136), suggesting that specific IC training attenuated the subjective effect of digital background. A total of 91 responses to the open-ended question on prerequisites for future acceptance of VR-based examinations were coded for thematic analysis (Supplementary Table 3). Three main priorities emerged: (1) Most participants emphasized the need for prior preparation with the VR environment, either theoretically (23/91, 25%) or practically (66/91, 73%), with 78% mentioning at least one of these forms and some additionally suggesting early curricular integration (16/91, 18%). Extensive training was viewed as essential to avoid inequities in operating the simulation. (2) A considerable number suggested program-related improvements to enhance accessibility (17/91, 19%), such as increased fonts or facilitated grasping interactions. (3) Finally, several participants also recommended adjustments to the examination conditions (11/91, 12%) such as extending the time limit or modifying evaluation criteria. Discussion This randomized controlled trial provides the first causal evidence that IC - defined as a user’s ability to navigate and interact with VR systems - significantly influences clinical performance in VR-based assessments (confirming H1). Participants who received specific IC training achieved higher scores than untrained peers in a complex, time-sensitive medical simulation. From a validity perspective, IC represents construct-irrelevant variance that may compromise the extrapolation inference in assessment design 21 , 22 . Analogous to challenges encountered in early computer-based testing 37 , 38 , our data indicate that VR-based performance metrics may conflate true clinical competence with interface proficiency. Interestingly, training general interaction techniques alone was insufficient - instead, context-specific training within the target application (specific IC training) proved necessary to achieve measurable performance benefits. Importantly, our findings extend beyond education. Comparable dynamics are likely in clinical VR applications - such as diagnostics in neurology, psychiatry or geriatrics 5 , 6 , 23 – 25 - where patients with limited digital abilities may underperform for non-clinical reasons, leading to misclassification or inappropriate therapeutic decisions. Such disparities mirror second-order digital divides in digital health, where user capability - rather than hardware access - limits participation 39 – 41 . Especially patient populations with little exposure to digital tools, including older adults, people with disabilities, and socioeconomically disadvantaged groups, may therefore experience recurrent disadvantages in VR-based diagnostics and interventions. While such issues are already well recognized in the context of algorithmic bias research 42 , 43 , they remain underexplored in immersive systems for educational and clinical assessments 32 , 8 , 10 . Procedural efficiency, measured as SPT, emerged as a plausible covariate (confirming H2a). Group effects appeared for both trained and untrained tasks with I2 participants completing more clinically relevant actions within the limited timeframe, which is consistent with effects from other VR trainings 44 , 45 . However, only SPT for untrained tasks contributed significantly as a moderator. Thus, clinical outcomes depended less on executing rehearsed steps quickly than on managing novel scenario components efficiently. Similar to clinical performance, general IC training (I1) had little effect on SPT, likely due to the higher cognitive and motor demands of high-fidelity VR environments 13 , 46 , for which general engagement with VR interaction techniques was likely insufficient. Comparable patterns have been documented in studies on internet skills and motor tasks, where general abilities show limited transfer to strategic or context-specific applications 47 , 48 . Altogether, these findings highlight procedural efficiency as both a sensitive proximal outcome and a potential mechanism for broader skill transfer in immersive simulations 49 , 50 . Although cognitive load was also hypothesized as an interacting variable, results were unexpected. Both NASA-TLX and EDA indicated lower load in the CO, yet these participants achieved the weakest clinical outcomes, and no linear mediating role was observed (refuting H2b). Added to that, subjective cognitive load showed an interaction effect, with medium values yielding the best performance 51 . Taken together, these observations challenge the assumption that IC training primarily reduces extraneous cognitive burden 52 , at least when training and assessment occur back-to-back. One explanation is that training itself induced substantial activation, creating a carry-over effect masking downstream reductions in cognitive load 53 . The high pre-assessment EDA in I1, who trained under the greatest time pressure due to rapid task repetition, support this interpretation. However, cognitive load theory offers an additional explanation 54 : low TLX1 ratings in CO might also reflect disengagement with task-relevant actions due to difficulties in VR interaction. In contrast, the higher load reported by trained participants might have been primarily germane - arising from active problem solving and procedural engagement 55 , 56 - and thus conducive to superior performance. Although we assessed cognitive load not only retrospectively, but also during the simulation using multiple measures 57 , especially EDA may lack the specificity required to disentangle cognitive from emotional arousal 58 , 59 . Thus, more precise methods (e.g., eye-tracking) - which were not available in combination with VR headsets at the time of the study - might provide a more detailed picture in the future. Non-parametric correlation analyses revealed a pattern opposite to our initial hypothesis that IC training would mitigate digital advantages (refuting H3). Unlike previous studies 16 , 18 , 19 , in our population unexposed to IC training (CO), no relevant association between prior 3D experience and clinical performance was detectable. In contrast, after specific training (I2) higher levels of prior 3D exposure were moderately related to superior performance, just exceeding the significance level. As participants in our groups reported uniformly low 3D experience (mean < 10 hours during the past 6 months with 75% reporting a value of 0), leaving little variance to reveal larger systematic effects 60 , these findings should generally be interpreted with caution. If they represent a genuine effect, training may not function solely as an equalizer, but rather as a facilitator that allows participants to leverage pre-existing competencies once basic interaction barriers are removed. This dynamic parallels the “Matthew effect” (magnification theory) described in cognitive 61 and educational research 62 , 63 , where training may amplify rather than level initial differences, for instance because a certain baseline skill is necessary to benefit. It is also possible that the training dose (25 minutes per module) was insufficient to compensate for substantial disparities in prior digital experience, as the duration of training may determine whether effects are compensatory or amplifying 64 . These considerations underscore the need for future studies to investigate extended or repeated IC as a potential levelling intervention. High enjoyment ratings and generally positive outlook on the fairness of VR-based examinations align with prior reports that immersive technologies are well accepted in medical education when they are perceived as engaging and safe learning environments 65 . At the same time, the pronounced distraction caused by interface barriers in untrained participants (confirming H4) reflects well-documented usability challenges of VR systems, including navigation difficulties, limited haptic feedback, and inconsistent interaction metaphors 12 , 13 , 66 . Consistent with previous findings, prior 3D experience appeared to buffer these effects subjectively 16 – 19 , underscoring that digital background may inadvertently bias assessment outcomes. Importantly, participants themselves emphasized the need for curricular integration, structured training access, and usability refinements as prerequisites for fair implementation, which converges with broader recommendations for the design of inclusive VR systems in clinical and educational contexts 67 . The current study demonstrates that failure to account for IC may reinforce digital inequities while creating a misleading appearance of objectivity. To address this risk, we propose three actionable mitigation strategies (Fig. 3 ): Design for inclusive interaction Developers should adhere to principles of intuitive interface design - reducing unnecessary complexity, standardizing interaction metaphors, and applying universal design elements 65,67 . These measures help minimize dependence on IC and enable participation by users with diverse digital backgrounds. Monitor for bias during implementation Performance data should be continuously monitored for correlations with digital background characteristics. Any unexpected disparities may indicate hidden dependencies on IC and require adjustments of the tool or its deployment protocol. Address IC through training or statistical control Before implementing immersive assessments, it is important to clarify how differences in IC will be addressed. Where feasible, standardized IC training modules can be provided, while monitoring their potential to amplify pre-existing advantages. When training is not feasible or its effects are uncertain, IC should be measured and accounted for in analyses - either as a covariate or as a stratification factor. Future research should first clarify the role of IC training. It remains uncertain whether training primarily reduces disparities or amplifies pre-existing advantages, and which doses, formats, or contexts promote levelling effects. Identifying reliable training protocols is therefore central to ensuring fair educational and clinical use. Longitudinal studies are also needed to trace how IC develops across educational stages and patient populations, and whether its influence on performance endures over time, particularly as extended reality becomes part of daily life. In parallel, cross-cultural and demographic research should examine how socioeconomic status, age, and technology access shape IC, linking it to broader issues of digital health equity. Finally, studies must determine how IC biases diagnostic conclusions in clinical domains such as neurology, psychiatry, and geriatrics, and how these biases can be mitigated through design improvements, targeted training, or calibration strategies. Limitations This study has several limitations. First, the sample consisted of medical students from a single German institution, which may restrict the generalizability of findings to other educational or clinical populations. Digital fluency and prior exposure to immersive systems likely vary across regions, age groups, and socioeconomic backgrounds. Second, our study tested a single, relatively brief IC training protocol. It therefore remains unclear whether different training doses, formats, or contexts might yield distinct effects - for example, acting as levelers in some cases and amplifiers in others. This uncertainty limits the generalizability of our findings and highlights the need for systematic variation of IC training designs in future work. Third, the measures of cognitive load have inherent constraints: EDA lacks specificity to distinguish cognitive from emotional or physical arousal, while NASA-TLX is subjective and susceptible to individual response tendencies. Finally, the VR scenarios focused exclusively on internal medicine emergencies. Although valid for this context, the findings may not be directly transferable to other specialties or patient populations. Conclusion Immersive competence may be a critical yet overlooked determinant of user performance in VR-based healthcare applications. Our findings demonstrate that targeted IC training can significantly enhance medical task execution and may help mitigate performance disparities, although its effects may depend on baseline user abilities and training conditions. Without accounting for IC, immersive technologies risk introducing unintended bias, compromising both the fairness and validity of high-stakes assessments and digital interventions. As VR continues to expand into clinical diagnostics, rehabilitation, and professional education, systematic strategies for measuring, training, or mitigating IC will be essential. Recognizing immersive competence as a modifiable, equity-relevant factor can help ensure that VR-based systems serve as enablers - rather than barriers - in advancing accessible and just digital health. Declarations Clinical trial number Not applicable. Funding This study received no funding. Acknowledgements None. Data availability The datasets generated and analyzed in this study are provided in Supplementary Data 1. The video recordings used for performance ratings can be obtained from the authors upon reasonable request. Author contribution statement J.S. conducted the experiments, collected the data, contributed to the analysis, and participated in writing the manuscript. T.L. processed and statistically analyzed the electrodermal activity (EDA) data and prepared the corresponding figures. M.L. supported the experiments and the handling of the EDA devices. A.Z. assisted in the processing of the EDA data. J.B. contributed to the study design and performed the interaction analysis. S.K. contributed to data presentation and manuscript writing. T.M. conceived and designed the study, supervised its conduct, performed the primary data analyses (except for EDA data and interaction analysis), and wrote the manuscript. All authors reviewed and approved the final version. Conflicts of interest TM was involved in the software development of the STEP-VR software. All other authors declare no conflicts of interest. References Liu, J. Y. W. et al. The Effects of Immersive Virtual Reality Applications on Enhancing the Learning Outcomes of Undergraduate Health Care Students: Systematic Review With Meta-synthesis. Journal of medical Internet research 25, e39989; 10.2196/39989 (2023). Cushnan, J., McCafferty, P. & Best, P. Clinicians' perspectives of immersive tools in clinical mental health settings: a systematic scoping review. BMC health services research 24, 1091; 10.1186/s12913-024-11481-3 (2024). Neher, A. N. et al. Virtual reality for assessment in undergraduate nursing and medical education - a systematic review. BMC medical education 25, 292; 10.1186/s12909-025-06867-8 (2025). Schiza, E., Matsangidou, M., Neokleous, K. & Pattichis, C. S. Virtual Reality Applications for Neurological Disease: A Review. Frontiers in robotics and AI 6, 100; 10.3389/frobt.2019.00100 (2019). Geraets, C. N. W., Wallinius, M. & Sygel, K. Use of Virtual Reality in Psychiatric Diagnostic Assessments: A Systematic Review. Frontiers in psychiatry 13, 828410; 10.3389/fpsyt.2022.828410 (2022). Hørlyck, L. D., Obenhausen, K., Jansari, A., Ullum, H. & Miskowiak, K. W. Virtual reality assessment of daily life executive functions in mood disorders: associations with neuropsychological and functional measures. Journal of affective disorders 280, 478–487; 10.1016/j.jad.2020.11.084 (2021). Freeman, D. et al . Automated psychological therapy using immersive virtual reality for treatment of fear of heights: a single-blind, parallel-group, randomised controlled trial. The lancet. Psychiatry 5, 625–632; 10.1016/S2215-0366(18)30226-8 (2018). Howard, M. C. A meta-analysis and systematic literature review of virtual reality rehabilitation programs. Computers in Human Behavior 70, 317–327; 10.1016/j.chb.2017.01.013 (2017). Spiegel, B. M. R. et al. Feasibility of combining spatial computing and AI for mental health support in anxiety and depression. NPJ digital medicine 7, 22; 10.1038/s41746-024-01011-0 (2024). Bargeri, S. et al. Effectiveness and safety of virtual reality rehabilitation after stroke: an overview of systematic reviews. eClinicalMedicine 64; 10.1016/j.eclinm.2023.102220 (2023). Wiederhold, B. K. & Wiederhold, M. D. Virtual reality therapy combined with physiological monitoring provides effective treatment, with objective metrics, for post-traumatic stress disorder. Expert review of medical devices 22, 117–119; 10.1080/17434440.2025.2454930 (2025). Lie, S. S., Helle, N., Sletteland, N. V., Vikman, M. D. & Bonsaksen, T. Implementation of Virtual Reality in Health Professions Education: Scoping Review. JMIR medical education 9, e41589; 10.2196/41589 (2023). Tusher, H. M., Mallam, S. & Nazir, S. A Systematic Review of Virtual Reality Features for Skill Training. Tech Know Learn 29, 843–878; 10.1007/s10758-023-09713-2 (2024). LaViola, J. J., Kruijff, E., McMahan, R. P., Bowman, D. & Poupyrev, I. P. 3D User Interfaces: Theory and Practice . 2nd ed. (Addison-Wesley, Boston, 2017). Tuena, C. et al. Usability Issues of Clinical and Research Applications of Virtual Reality in Older People: A Systematic Review. Frontiers in human neuroscience 14, 93; 10.3389/fnhum.2020.00093 (2020). Katz, D. et al. Relationship between demographic and social variables and performance in virtual reality among healthcare personnel: an observational study. BMC medical education 24, 227; 10.1186/s12909-024-05180-0 (2024). Oberdörfer, S. et al. Ready for VR? Assessing VR Competence and Exploring the Role of Human Abilities and Characteristics. (Preprint) (2025). Schreiner, V. et al. Specific Immersive Competence in VR-Based Assessments: Development, Psychometric Evaluation and Associations with Medical Performance (Preprint) (2025). Schlickum, M. K., Hedman, L., Enochsson, L., Kjellin, A. & Felländer-Tsai, L. Systematic video game training in surgical novices improves performance in virtual reality endoscopic surgical simulators: a prospective randomized study. World journal of surgery 33, 2360–2367; 10.1007/s00268-009-0151-y (2009). Steed, A. et al. Immersive competence and immersive literacy: Exploring how users learn about immersive experiences. Front. Virtual Real. 4; 10.3389/frvir.2023.1129242 (2023). Messick, S. Standards of Validity and the Validity of Standards in Performance Asessment. Educational Measurement 14, 5–8; 10.1111/j.1745-3992.1995.tb00881.x (1995). Kane, M. T. Validating the Interpretations and Uses of Test Scores. J Educational Measurement 50, 1–73; 10.1111/jedm.12000 (2013). Du, K., Benavides, L. R., Isenstein, E. L., Tadin, D. & Busza, A. C. Virtual reality assessment of reaching accuracy in patients with recent cerebellar stroke. BMC Digit Health 2, 50; 10.1186/s44247-024-00107-7 (2024). Bell, I. H., Nicholas, J., Alvarez-Jimenez, M., Thompson, A. & Valmaggia, L. Virtual reality as a clinical tool in mental health research and practice. Dialogues in clinical neuroscience 22, 169–177; 10.31887/DCNS.2020.22.2/lvalmaggia (2020). Atkins, A. S. et al. Assessment of Age-Related Differences in Functional Capacity Using the Virtual Reality Functional Capacity Assessment Tool (VRFCAT). The journal of prevention of Alzheimer's disease 2, 121–127; 10.14283/jpad.2015.61 (2015). Mühling, T. et al. Virtual reality in medical emergencies training: benefits, perceived stress, and learning success. Multimedia Systems 29, 2239–2252; 10.1007/s00530-023-01102-0 (2023). Weiss, S. L. et al . Surviving Sepsis Campaign International Guidelines for the Management of Septic Shock and Sepsis-Associated Organ Dysfunction in Children. Pediatric critical care medicine: a journal of the Society of Critical Care Medicine and the World Federation of Pediatric Intensive and Critical Care Societies 21, e52-e106; 10.1097/PCC.0000000000002198 (2020). Mühling, T., Schreiner, V., Appel, M., Leutritz, T. & König, S. Comparing Virtual Reality-Based and Traditional Physical Objective Structured Clinical Examination (OSCE) Stations for Clinical Competency Assessments: Randomized Controlled Trial. Journal of medical Internet research 27, e55066; 10.2196/55066 (2025). Hart, S. G. & Staveland, L. E. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Advances in Psychology: Human Mental Workload , edited by P. A. Hancock & N. Meshkati (North-Holland1988), Vol. 52, pp. 139–183. Flägel, K., Galler, B., Steinhäuser, J. & Götz, K. Der „National Aeronautics and Space Administration-Task Load Index“ (NASA-TLX) – ein Instrument zur Erfassung der Arbeitsbelastung in der hausärztlichen Sprechstunde: Bestimmung der psychometrischen Eigenschaften. Zeitschrift fur Evidenz, Fortbildung und Qualitat im Gesundheitswesen 147–148, 90–96; 10.1016/j.zefq.2019.10.003 (2019). Kleckner, I. R. et al. Simple, Transparent, and Flexible Automated Quality Assessment Procedures for Ambulatory Electrodermal Activity Data. IEEE transactions on bio-medical engineering 65, 1460–1467; 10.1109/TBME.2017.2758643 (2018). Looff, P. de et al. Wearables: An R Package With Accompanying Shiny Application for Signal Analysis of a Wearable Device Targeted at Clinicians and Researchers. Frontiers in behavioral neuroscience 16, 856544; 10.3389/fnbeh.2022.856544 (2022). Posada-Quintero, H. F. & Chon, K. H. Innovations in Electrodermal Activity Data Collection and Signal Processing: A Systematic Review. Sensors (Basel, Switzerland) 20; 10.3390/s20020479 (2020). Braun, V. & Clarke, V. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 77–101; 10.1191/1478088706qp063oa (2006). Hanley, J. A. Simple and multiple linear regression: sample size considerations. Journal of clinical epidemiology 79, 112–119; 10.1016/j.jclinepi.2016.05.014 (2016). Eysenbach, G. CONSORT-EHEALTH: improving and standardizing evaluation reports of Web-based and mobile health interventions. Journal of medical Internet research 13, e126; 10.2196/jmir.1923 (2011). Valentine, A., Vrbik, P. & Thomas, R. A systematic review of paper-based versus computer-based testing in engineering and computing education, 364–372; 10.1109/EDUCON52537.2022.9766469 (2022). Khoshsima, H., Hosseini, M. & Toroujeni, S. M. H. Cross-Mode Comparability of Computer-Based Testing (CBT) Versus Paper-Pencil Based Testing (PPT): An Investigation of Testing Administration Mode among Iranian Intermediate EFL Learners. ELT 10, 23; 10.5539/elt.v10n2p23 (2017). Adedinsewo, D. et al. Health Disparities, Clinical Trials, and the Digital Divide. Mayo Clinic proceedings 98, 1875–1887; 10.1016/j.mayocp.2023.05.003 (2023). Shaw, J., Brewer, L. C. & Veinot, T. Recommendations for Health Equity and Virtual Care Arising From the COVID-19 Pandemic: Narrative Review. JMIR formative research 5, e23233; 10.2196/23233 (2021). Riggins, F. & Dewan, S. The Digital Divide: Current and Future Research Directions. JAIS 6, 298–337; 10.17705/1jais.00074 (2005). Vayena, E., Blasimme, A. & Cohen, I. G. Machine learning in medicine: Addressing ethical challenges. PLoS medicine 15, e1002689; 10.1371/journal.pmed.1002689 (2018). Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science (New York, N.Y.) 366, 447–453; 10.1126/science.aax2342 (2019). Binstadt, E., Donner, S., Nelson, J., Flottemesch, T. & Hegarty, C. Simulator training improves fiber-optic intubation proficiency among emergency medicine residents. Academic emergency medicine: official journal of the Society for Academic Emergency Medicine 15, 1211–1214; 10.1111/j.1553-2712.2008.00199.x (2008). Lohre, R. et al. Effectiveness of Immersive Virtual Reality on Orthopedic Surgical Skills and Knowledge Acquisition Among Senior Surgical Residents: A Randomized Clinical Trial. JAMA network open 3, e2031217; 10.1001/jamanetworkopen.2020.31217 (2020). Arthur, T. et al. Examining the validity and fidelity of a virtual reality simulator for basic life support training. BMC Digit Health 1; 10.1186/s44247-023-00016-1 (2023). van Deursen, A. & van Dijk, J. Internet skills and the digital divide. New Media & Society 13, 893–911; 10.1177/1461444810386774 (2011). Wulf, G. & Shea, C. H. Principles derived from the study of simple skills do not generalize to complex skill learning. Psychonomic bulletin & review 9, 185–211; 10.3758/BF03196276 . (2002). Torkington, J., Smith, S. G., Rees, B. I. & Darzi, A. Skill transfer from virtual reality to a real laparoscopic task. Surgical endoscopy 15, 1076–1079; 10.1007/s004640000233 (2001). Clarke, D. B. et al. Knowledge transfer and retention of simulation-based learning for neurosurgical instruments: a randomised trial of perioperative nurses. BMJ simulation & technology enhanced learning 7, 146–153; 10.1136/bmjstel-2019-000576 (2021). Pyke, W., Lunau, J. & Javadi, A.-H. Does difficulty moderate learning? A comparative analysis of the desirable difficulties framework and cognitive load theory. Quarterly journal of experimental psychology (2006) , 17470218241308143; 10.1177/17470218241308143 (2024). Marsh, W. E., Hantel, T., Zetzsche, C. & Schill, K. Is the user trained? Assessing performance and cognitive resource demands in the Virtusphere, 15–22; 10.1109/3DUI.2013.6550191 (2013). Rölfing, J. D., Nørskov, J. K., Paltved, C., Konge, L. & Andersen, S. A. W. Failure affects subjective estimates of cognitive load through a negative carry-over effect in virtual reality simulation of hip fracture surgery. Advances in simulation (London, England) 4, 26; 10.1186/s41077-019-0114-9 (2019). Chandler, P. & Sweller, J. Cognitive Load Theory and the Format of Instruction. Cognition and Instruction 8, 293–332; 10.1207/s1532690xci0804_2 (1991). Leppink, J., Paas, F., van Gog, T., van der Vleuten, C. P. & van Merriënboer, J. J. Effects of pairs of problems and examples on task performance and different types of cognitive load. Learning and Instruction 30, 32–42; 10.1016/j.learninstruc.2013.12.001 (2014). Young, J. Q., van Merrienboer, J., Durning, S. & Cate, O. ten. Cognitive Load Theory: implications for medical education: AMEE Guide No. 86. Medical teacher 36, 371–384; 10.3109/0142159X.2014.889290 (2014). Naismith, L. M. & Cavalcanti, R. B. Validity of Cognitive Load Measures in Simulation-Based Training: A Systematic Review. Academic medicine: journal of the Association of American Medical Colleges 90, S24-35; 10.1097/ACM.0000000000000893 (2015). Ayres, P., Lee, J. Y., Paas, F. & van Merriënboer, J. J. G. The Validity of Physiological Measures to Identify Differences in Intrinsic Cognitive Load. Frontiers in psychology 12, 702538; 10.3389/fpsyg.2021.702538 (2021). Halbig, A. & Latoschik, M. E. A Systematic Review of Physiological Measurements, Factors, Methods, and Applications in Virtual Reality. Front. Virtual Real. 2; 10.3389/frvir.2021.694567 (2021). Glass, G. V. & Hopkins, K. D. Statistical Methods in Education and Psychology . 3rd ed. (Allyn and Bacon, 1996). Lövdén, M., Brehmer, Y., Li, S.-C. & Lindenberger, U. Training-induced compensation versus magnification of individual differences in memory performance. Frontiers in human neuroscience 6, 141; 10.3389/fnhum.2012.00141 (2012). Bast, J. & Reitsma, P. Analyzing the development of individual differences in terms of Matthew effects in reading: results from a Dutch Longitudinal study. Developmental psychology 34, 1373–1399; 10.1037/0012-1649.34.6.1373 (1998). McVey, R. et al. Baseline Laparoscopic Skill May Predict Baseline Robotic Skill and Early Robotic Surgery Learning Curve. Journal of endourology 30, 588–592; 10.1089/end.2015.0774 (2016). Liu, L. et al. Dose-response relationship between computerized cognitive training and cognitive improvement. NPJ digital medicine 7, 214; 10.1038/s41746-024-01210-9 (2024). Radianti, J., Majchrzak, T. A., Fromm, J. & Wohlgenannt, I. A systematic review of immersive virtual reality applications for higher education: Design elements, lessons learned, and research agenda. Computers & Education 147, 103778; 10.1016/j.compedu.2019.103778 (2020). Mühling, T., Backhaus, J., Demmler, L. & König, S. How Personality and Affective Responses Are Associated with Skepticism Towards Virtual Reality in Medical Training-A Pre-Post Intervention Study. Cyberpsychology, behavior and social networking 28, 335–341; 10.1089/cyber.2024.0567 (2025). Chamusca, I. L. et al. Evaluating Design Guidelines for Intuitive, Therefore Sustainable, Virtual Reality Authoring Tools. Sustainability 16, 1744; 10.3390/su16051744 (2024). Additional Declarations Competing interest reported. TM was involved in the software development of the STEP-VR software. All other authors declare no conflicts of interest. Supplementary Files SchaaletalSupplementaryData1.xlsx SchaaletalSupplementaryInformation.docx SchaaletalCONSORTEHEALTHV1.61.docx Cite Share Download PDF Status: Published Journal Publication published 09 Mar, 2026 Read the published version in npj Digital Medicine → Version 1 posted Editorial decision: Revision requested 13 Nov, 2025 Reviews received at journal 10 Nov, 2025 Reviews received at journal 27 Oct, 2025 Reviewers agreed at journal 19 Oct, 2025 Reviewers agreed at journal 03 Oct, 2025 Reviewers invited by journal 03 Oct, 2025 Editor assigned by journal 27 Sep, 2025 Submission checks completed at journal 26 Sep, 2025 First submitted to journal 19 Sep, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7660457","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":529625351,"identity":"a9e28684-e000-4dfd-9ffa-8d6b7d6e83c8","order_by":0,"name":"Jan Schaal","email":"","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":false,"prefix":"","firstName":"Jan","middleName":"","lastName":"Schaal","suffix":""},{"id":529625352,"identity":"f88bce8e-f34e-4736-9ff0-cf75ecaf92a6","order_by":1,"name":"Tobias Leutritz","email":"","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":false,"prefix":"","firstName":"Tobias","middleName":"","lastName":"Leutritz","suffix":""},{"id":529625353,"identity":"23e65bec-5d7a-4eef-8bba-bc2e2f2ec136","order_by":2,"name":"Marco Lindner","email":"","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":false,"prefix":"","firstName":"Marco","middleName":"","lastName":"Lindner","suffix":""},{"id":529625354,"identity":"7d1cd1d7-d488-489d-a230-78917d7d0866","order_by":3,"name":"Alexander Zamzow","email":"","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":false,"prefix":"","firstName":"Alexander","middleName":"","lastName":"Zamzow","suffix":""},{"id":529625355,"identity":"66cda513-80de-4902-bfd4-e0f4251714b3","order_by":4,"name":"Joy Backhaus","email":"","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":false,"prefix":"","firstName":"Joy","middleName":"","lastName":"Backhaus","suffix":""},{"id":529625356,"identity":"b15fa0b0-5f44-4aea-b612-ad459007ff2b","order_by":5,"name":"Sarah König","email":"","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":false,"prefix":"","firstName":"Sarah","middleName":"","lastName":"König","suffix":""},{"id":529625357,"identity":"7d9ee3f7-0b4a-421e-9ff6-4950daa7ce2c","order_by":6,"name":"Tobias Mühling","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxklEQVRIiWNgGAWjYHACNoYHBiCa+QBjAwMDCBOhJQGkhY0tgRQtYIrHgDgtujOSjz1IKLhj1yDf801yZhuDbD8hLWY30tINEgyeJTew8W6T3NjGYDyTkDVmN3LMJBIMDiczgLQ83MaQuOEAQS3536BaeJ6BtewnrCWHDaTFDqiFTXIjyBaCfjnzDOywBAa2NGPLmf8kjGcQtOV48jOJD38O2zMwH354s+eMjWx/AyFroADmBQki1QOBPfFKR8EoGAWjYMQBAAymQjG2THATAAAAAElFTkSuQmCC","orcid":"","institution":"University Hospital Würzburg","correspondingAuthor":true,"prefix":"","firstName":"Tobias","middleName":"","lastName":"Mühling","suffix":""}],"badges":[],"createdAt":"2025-09-19 16:53:24","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7660457/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7660457/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41746-026-02482-z","type":"published","date":"2026-03-09T15:59:40+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":93731037,"identity":"7328ae1d-a8dc-4722-9081-7bf1b52c48ba","added_by":"auto","created_at":"2025-10-17 02:22:44","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":779157,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalManuscriptAmendment1.docx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/0bbd70d65fd1c0e4c1d07abd.docx"},{"id":93731036,"identity":"a10b1e7b-c48c-412c-a3e6-4e6c7a062c9a","added_by":"auto","created_at":"2025-10-17 02:22:44","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8591,"visible":true,"origin":"","legend":"","description":"","filename":"13d469c1e2e242cdb15e915a2f005ce2.json","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/c90316ba02f716b73dff0740.json"},{"id":93732379,"identity":"a3baa082-06b8-4f5d-a348-b75f1a17d195","added_by":"auto","created_at":"2025-10-17 02:30:44","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18726,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalCONSORTEHEALTHV1.61.docx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/1be58f97eb6f58d0075b08b8.docx"},{"id":93732380,"identity":"782e618c-0204-4032-a028-675618f0631f","added_by":"auto","created_at":"2025-10-17 02:30:44","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":41025,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalSupplementaryData1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/91e107c056d917a8a93e635b.xlsx"},{"id":93729349,"identity":"b8e79758-f7fe-41d8-a9de-b5a3cd0991ec","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"docx","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":24671,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalSupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/f42ab4f90bf2328ec95d3527.docx"},{"id":93729357,"identity":"cea64980-89f1-44f4-9e0a-c2928bf5c293","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"xml","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":155659,"visible":true,"origin":"","legend":"","description":"","filename":"13d469c1e2e242cdb15e915a2f005ce21enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/7cb465ca94ed443baa866a2b.xml"},{"id":93731042,"identity":"91db7eee-1775-41d5-912c-2b811a0866d2","added_by":"auto","created_at":"2025-10-17 02:22:44","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":260133,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/6dfd2f929ee3f378d3b5a679.png"},{"id":93731041,"identity":"a675231e-0458-4d86-ad98-0e1e242fa726","added_by":"auto","created_at":"2025-10-17 02:22:44","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":50911,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/5cd86f44ce2cac9c1510093f.png"},{"id":93729361,"identity":"169d03ea-9382-4079-8dd5-a13e3847bb4b","added_by":"auto","created_at":"2025-10-17 02:14:45","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":79820,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/d61bca733e17203efb628349.png"},{"id":93729359,"identity":"942a4c3b-15af-47eb-a2e8-ef8cf88d30c1","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45042,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/56f8e4a26735ce7939256846.png"},{"id":93729358,"identity":"76592bd1-cc01-4c47-804f-3eedf5affe30","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"xml","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":155892,"visible":true,"origin":"","legend":"","description":"","filename":"13d469c1e2e242cdb15e915a2f005ce21structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/3d407baf29404b65eb832b8b.xml"},{"id":93729362,"identity":"387e0f0c-c53b-46d0-ac3b-748b4e6bee66","added_by":"auto","created_at":"2025-10-17 02:14:45","extension":"html","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":170678,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/80b0d6c2c2674fc1da69b2a8.html"},{"id":93729354,"identity":"8784f862-0266-4b7f-8c55-161d8bbfe19b","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":144190,"visible":true,"origin":"","legend":"\u003cp\u003eStudy design and data collection process for IC training and VR-based assessment of clinical performance. EDA: electrodermal activity, IC: immersive competence, TLX: NASA task load index, SPT: standardized procedural time, VR: virtual reality\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/0842fcc18f2f23778bb1d0cc.png"},{"id":93729347,"identity":"3557da58-86e8-4870-aaaa-9a423d3bc2c2","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":249023,"visible":true,"origin":"","legend":"\u003cp\u003ePrimary and secondary outcomes. (a) Effectiveness of general and specific IC training. Bar graphs depict group means, with overlaid dots representing individual participants in the pre–post comparison. (b) Impact of IC on clinical performance in the VR assessment, illustrated by violin plots with medians (black horizontal lines) and means ± SD (square symbols with error bars). (c) Correlation between 3D experience (expressed as rank) and clinical performance across groups. Individual data points are shown together with their respective regression lines. Ranks reflect the relative position of each participant’s 3D experience within the combined sample. (d) Procedural efficiency comparing previously trained (left) and untrained (right) measures. (e) Cognitive load, including progression of EDA during VR assessment and group differences for NASA-TLX scores at time points 1 (during assessment) and 2 (immediately after assessment). (f) Effect size of the interaction analysis: influence of IC on clinical performance through cognitive load and procedural efficiency. Abbreviations. CO = control group, EDA = electrodermal activity, I1/I2 = intervention group 1/2, IC = immersive competence, SPT = standardized procedural time, TLX = task load index, VR = virtual reality.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/902593c39f683abe285d7a4c.png"},{"id":93729353,"identity":"44f246f4-b94c-4970-919a-d42d63bd0e2c","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":126126,"visible":true,"origin":"","legend":"\u003cp\u003eConceptual model of IC bias in VR-based assessments and potential mitigation strategies. System factors (e.g., low interaction fidelity, unfavorable control mapping) and user factors (varying levels of IC) may introduce bias into VR-based clinical or functional assessments. In the main assessment flow (horizontal arrow), the intended target constructs - clinical competence and cognitive/motor function - become contaminated by IC-related variance. Mitigation measures are shown on the right: intuitive and inclusive design addresses system factors, while IC training or statistical control addresses user factors. Monitoring performance bias during implementation feeds back into both domains, enabling continuous detection and reduction of bias. IC: immersive competence, VR: virtual reality\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/1572188d8e19ccb441e5bd16.png"},{"id":104739565,"identity":"3d506c2e-d54e-4e6c-b9d7-83b339f0bc79","added_by":"auto","created_at":"2026-03-16 16:09:19","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1497669,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/2cdd5cba-3847-4d68-8bef-fb9a610eed9e.pdf"},{"id":93729345,"identity":"079b0667-95ee-4bb6-80f9-97ff1311ddd5","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":41025,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalSupplementaryData1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/7646854dbd0027a592640d91.xlsx"},{"id":93729344,"identity":"7a5b4fe5-f24a-4489-b253-237806c950ce","added_by":"auto","created_at":"2025-10-17 02:14:44","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":24671,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalSupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/3fac5239e54dc2297b5b1361.docx"},{"id":93731040,"identity":"ab862e49-d596-4b47-8e8c-db6217fc4f42","added_by":"auto","created_at":"2025-10-17 02:22:44","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":18726,"visible":true,"origin":"","legend":"","description":"","filename":"SchaaletalCONSORTEHEALTHV1.61.docx","url":"https://assets-eu.researchsquare.com/files/rs-7660457/v1/de83f2a757281bb90ceee759.docx"}],"financialInterests":"Competing interest reported. TM was involved in the software development of the STEP-VR software. All other authors declare no conflicts of interest.","formattedTitle":"Uncovering Immersive Competence as a Hidden Bias in VR-Based Clinical Assessment – A Randomized Controlled Study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eVirtual Reality (VR) is increasingly used not only in medical education\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e but also in clinical diagnostics and therapy\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. In immersive environments, users can engage in realistic and standardized simulations to practice or assess professional skills\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, monitor cognitive and motor functions\u003csup\u003e\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, or receive psychological or rehabilitative interventions\u003csup\u003e\u003cspan additionalcitationids=\"CR8 CR9 CR10\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. However, as these applications gain momentum in high-stakes contexts, questions of validity and equitable access become critical. VR systems are neither standardized nor always intuitive, leading to substantial variations in functionality and usability\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Learners and patients alike must master abstract interaction metaphors, interpret spatial feedback, and operate handheld controllers\u003csup\u003e\u003cspan additionalcitationids=\"CR13 CR14\" citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. This creates a hidden source of variance, where prior informal digital experience (e.g., video gaming) can unfairly advantage certain users\u003csup\u003e\u003cspan additionalcitationids=\"CR17 CR18\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Yet, the underlying ability that mediates this effect has remained largely underdefined.\u003c/p\u003e\u003cp\u003eRecent work has introduced the concept of Immersive Competence (IC) - the user\u0026rsquo;s familiarity and skill in operating VR systems\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e - as a potentially construct-irrelevant factor\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e influencing performance. From a psychometric perspective, this threatens the validity of VR-based assessments, particularly the extrapolation inference\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e, by conflating technological fluency with domain-specific ability (e.g. clinical performance). Despite this, none of the 26 VR-based assessments in health professions education listed in a recent systematic review accounted for IC as a potential confounder\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Likewise, clinical VR assessments in neurology\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, psychiatry\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e, or geriatrics\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e have almost exclusively focused on task outcomes, while largely overlooking patients\u0026rsquo; ability to operate these immersive systems. However, without accounting for IC, VR-based evaluations risk perpetuating unseen performance biases by disadvantaging less digitally experienced learners or patients and undermining fairness and validity.\u003c/p\u003e\u003cp\u003eTo address this gap, recent tools have been developed to measure and train IC in a standardized manner, covering general interaction categories such as manipulation, navigation, and system control\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e and to train and assess domain-specific IC directly within a target VR environment\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. Yet, empirical studies demonstrating whether IC causally influences clinical performance - and whether training can mitigate this effect - remain lacking. The present study therefore investigates whether training of IC systematically affects performance in a VR-based clinical assessment. We hypothesized that (H1) targeted IC training would enhance clinical performance compared to no training, (H2) this effect would be mediated by increases in procedural efficiency (H2a) and/or reductions in cognitive load (H2b), (H3) IC training would mitigate performance disparities associated with prior 3D or VR application experience, and (H4) trained participants would report less distraction by usability barriers and a more positive evaluation of the VR-based assessment. Findings aim to inform the design of equitable and valid VR-based tools in both education and healthcare \u0026ndash; ensuring that performance reflects actual clinical competence in examinees and functional ability in patients rather than technological familiarity.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eStudy design and setting\u003c/h2\u003e\u003cp\u003eThis single-center, randomized controlled trial was conducted between January and August 2025 at the clinical skills laboratory of a German medical faculty. The aim was to test whether IC training influences performance in a VR-based clinical assessment. Eligible participants were advanced undergraduate medical students who had completed the internal medicine written exam (end of 7th semester). Exclusion criteria were known epilepsy or severe simulator sickness. Recruitment occurred via semester mailing lists and enrollment was online. Sessions took place individually, outside curricular teaching hours in a standardized VR laboratory. Participants received a \u0026euro;25 book voucher upon study completion.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eRandomization and blinding\u003c/h3\u003e\n\u003cp\u003eParticipants were randomly allocated (1:1:1) to three groups using a computer-generated sequence with fixed block size (n\u0026thinsp;=\u0026thinsp;9). Due to visible differences in training, blinding for participants was not feasible. However, performance was recorded automatically and rated from anonymized videos, ensuring rater blinding (single-blind design).\u003c/p\u003e\n\u003ch3\u003eSample size\u003c/h3\u003e\n\u003cp\u003ePower analysis (one-way ANOVA, α\u0026thinsp;=\u0026thinsp;0.05, power\u0026thinsp;=\u0026thinsp;0.80) indicated that a minimum of 84 participants would be needed to detect a moderate-to-large effect size (f\u0026thinsp;=\u0026thinsp;0.25). To account for potential attrition, 94 were enrolled.\u003c/p\u003e\n\u003ch3\u003eInterventions\u003c/h3\u003e\n\u003cp\u003eParticipants were assigned to one of three study groups:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eIntervention group 1 (I1): General IC training (abstract VR interactions).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIntervention group 2 (I2): General\u0026thinsp;+\u0026thinsp;specific IC training (in-examination environment).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eControl group (CO): No structured IC training.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eAfter providing informed consent and a basic familiarization with VR controls, I1 and I2 completed their respective IC training sessions. All groups then underwent the same VR-based clinical assessment.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eGeneral IC training\u003c/h3\u003e\n\u003cp\u003eGeneral IC training was delivered using the VR Competence App\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, which includes eight modular interaction challenges (e.g., object selection, teleportation, system control) unrelated to clinical content. Participants received immediate performance feedback and repeated tasks if their score fell below the 75th percentile of a pilot cohort. Training duration was capped at 25 minutes. A full overview of challenges and scoring thresholds is provided in Supplementary Table\u0026nbsp;1.\u003c/p\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003eSpecific IC training\u003c/h2\u003e\u003cp\u003eSpecific IC training was conducted in the STEP-VR emergency simulation environment\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e, identical to the assessment setting. Participants followed an audio-guided sequence of seven procedural tasks (e.g., blood sampling, IV access setup), each comprising multiple time steps. Tasks required no medical knowledge but replicated complex interaction patterns. As in the general IC training, only unsuccessful tasks were repeated. Training was limited to 25 minutes. Task content and time limits are detailed in Supplementary Table\u0026nbsp;1.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eVR-based clinical assessment\u003c/h3\u003e\n\u003cp\u003eParticipants completed a 9-minute VR-based simulation scenario of septic shock due to abscess-forming Crohn\u0026rsquo;s disease. The scenario focused on implementing the \"Sepsis One-Hour Bundle\" and initiating guideline-based surgical or interventional measures\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. Content validity was confirmed by medical educators and the scenario had been used previously in a VR-based OSCE\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Before the assessment, participants received a 5-minute orientation to the virtual emergency room and structured case information equivalent to standard OSCE pre-station instructions. The scenario was presented in first-person VR, mirrored to an external screen for supervision and video recording.\u003c/p\u003e\n\u003ch3\u003eOutcome measures\u003c/h3\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003ePrimary outcome: Clinical performance\u003c/h2\u003e\u003cp\u003eThe primary outcome of the study was clinical performance in the VR-based assessment. Performance was measured with a previously validated OSCE checklist\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e, adapted to the knowledge level of students in the 7th to 10th semester (Supplementary Table\u0026nbsp;2). The adapted version included nine items, scored either binarily (1 or 0) or ternarily (1, 0.5, or 0). To ensure that the assessment captured genuine clinical competence rather than short-term memory effects, none of the checklist items were rehearsed during the IC training. Ratings were based on anonymized video recordings and independently conducted by two blinded raters. All items were equally weighted, and a total performance score was calculated for each participant.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003eSecondary outcomes: Procedural efficiency, cognitive load, as well as self-reported usability barriers and acceptance\u003c/h2\u003e\u003cp\u003eProcedural efficiency was defined as the standardized procedural time (SPT), measured as the interval between the observable initiation and the completion of a procedure. SPT was determined for one previously trained task (intravenous cannulation) and four untrained tasks (blood culture sampling, temperature measurement, abdominal ultrasound, and pneumatic tube use). Since these procedures differed in their inherent time requirements, raw durations were standardized within the sample to Z-scores (mean\u0026thinsp;=\u0026thinsp;0, SD\u0026thinsp;=\u0026thinsp;1). For untrained tasks, an aggregate score was calculated as the mean of standardized times; lower SPT values indicated higher procedural efficiency.\u003c/p\u003e\u003cp\u003eCognitive load was assessed both subjectively and objectively. Subjective ratings were obtained using the German version of the unweighted (raw) NASA-TLX\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e,\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e at two time points: after 3 minutes (TLX1, Mental Demand subscale) and immediately after the simulation (TLX2, full scale). Ratings ranged from 0\u0026thinsp;=\u0026thinsp;lowest to 100\u0026thinsp;=\u0026thinsp;highest.\u003c/p\u003e\u003cp\u003eObjective data were derived from electrodermal activity (EDA) recordings using the Empatica Embrace Pro wristband. Valid recordings were defined as skin conductance levels between 0.05 and 60 \u0026micro;S and skin temperature between 30 and 40\u0026deg;C\u003csup\u003e31\u003c/sup\u003e. Data were processed in R (v4.3.1, wearables package\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e) following best-practice standards\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e, with median per minute values used for group-comparisons.\u003c/p\u003e\u003cp\u003eFinally, self-reported usability barriers and acceptance were evaluated with a post-assessment questionnaire. Items addressed distraction by usability barriers, perceptions of fairness and enjoyment. An open-ended item invited participants to comment on prerequisites for future VR examinations use and responses were subjected to \u0026ldquo;thematic analysis\u0026rdquo;\u003csup\u003e34\u003c/sup\u003e.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003eParticipant characteristics and experience\u003c/h2\u003e\u003cp\u003eParticipant characteristics and prior experiences were collected, including gender, age and exam performance, operationalized as the mean grade across Internal Medicine, Emergency Medicine, and Anesthesiology according to the German grading system (1\u0026thinsp;=\u0026thinsp;best, 5\u0026thinsp;=\u0026thinsp;worst). In addition, participants reported their experience with 3D (computer-based) and VR applications, quantified as the number of hours of use during the past six months.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003eSoftware and hardware\u003c/h2\u003e\u003cp\u003eGeneral IC training was conducted using the VR Competence App\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, developed by the Chair for Human-Computer Interaction at the University of W\u0026uuml;rzburg. Specific training and assessment scenarios were implemented on the STEP-VR platform\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e, an immersive simulation built on a commercial game engine (Unreal Engine 5.1), developed in collaboration with ThreeDee GmbH (Munich, Germany). Interaction logic of both applications was implemented using native controller input and physics-based object handling: object manipulation used a direct-grab paradigm with collision physics and snap-to-socket behavior for connectors (e.g., infusion lines), while UI interactions employed a menu with raycast selection and confirmation buttons, both supported by haptic feedback and contextual audio cues. Locomotion combined smooth walking for short distances with teleportation for longer distances.\u003c/p\u003e\u003cp\u003eThe VR setup included an OMEN by HP 17-ck0075ng laptop (Intel Core i7-11800H, NVIDIA GeForce RTX 3070 [8 GB]) and an Oculus Quest III headset with handheld controllers. Applications were rendered at \u0026gt;\u0026thinsp;90 frames per second with high-quality graphic settings.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003eStatistical analysis\u003c/h2\u003e\u003cp\u003eNormality of residuals was assessed using the Shapiro-Wilk test. For group comparisons, one-way ANOVA was applied when assumption of normality was met. In cases of deviation, the non-parametric Kruskal-Wallis test was used. Post hoc comparisons of non-normally distributed data were performed using the Mann-Whitney test. For correlation analyses, Pearson\u0026rsquo;s correlation coefficient was calculated for normally distributed variables, while Spearman\u0026rsquo;s rank correlation was applied otherwise.\u003c/p\u003e\u003cp\u003eTo explore interaction effects of covariates, regression analyses were conducted in R (v.4.5.1 with RStudio v. 2024.04.2). Effect sizes were expressed as partial eta-squared (\u003cem\u003eη\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e) with corresponding confidence intervals, with thresholds of 0.01, 0.06, and 0.14 interpreted as small, medium and large. Regression equations were reported in the form Y\u0026thinsp;=\u0026thinsp;a\u0026thinsp;+\u0026thinsp;bX, to indicate the magnitude and direction of effects. For cell sizes smaller than five, regression coefficients were not computed\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eInterrater reliability of checklist scoring was calculated using linear weighted Cohen κ, and group differences in gender distribution were tested using chi-square.\u003c/p\u003e\u003cp\u003eEDA data were processed in R (v4.3.1; RRID:SCR_001905) using a specific wearables package\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. Median per-minute values were computed for each participant and used for group-level comparisons.\u003c/p\u003e\u003cp\u003eAll other analyses were conducted with GraphPad Prism (v10.1.2; RRID:SCR_002798).\u003c/p\u003e\u003cp\u003eReporting was conducted in accordance with the CONSORT-EHEALTH recommendations (for a completed reporting checklist, see supplementary files)\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\u003ch2\u003eEthics and data protection\u003c/h2\u003e\u003cp\u003eThis study was reviewed and approved by the Ethics Committee of the University of W\u0026uuml;rzburg (Proposal No. 2024-316-ka). Written informed consent was obtained from all participants. All data were processed anonymously and handled in accordance with local data protection regulations. VR interaction data and physiological recordings were stored on encrypted institutional servers and anonymized before analysis. First-person video recordings used for performance rating were de-identified (audio track removed) to ensure participant confidentiality. Participation in the study had no academic consequences for students.\u003c/p\u003e\u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003eParticipant characteristics\u003c/h2\u003e\u003cp\u003eA total of 94 medical students were enrolled and randomized across the three study arms (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Due to technical problems with the VR hardware/software, missing video recordings and incomplete questionnaires, six datasets were excluded from the final analysis. The remaining 88 participants were representative of the overall population of advanced undergraduate medical students (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Academic performance, measured as the exam score, was slightly better in CO compared to I1 and I2 (p\u0026thinsp;=\u0026thinsp;.020). Overall, participants reported limited prior exposure to 3D applications (mean: 9.36\u0026thinsp;\u0026plusmn;\u0026thinsp;37.0 hours in the past six months) and minimal prior VR experience (mean: 0.07\u0026thinsp;\u0026plusmn;\u0026thinsp;0.23 hours in the past six months), consistent with findings from comparable cohorts in medical educational.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eParticipant characteristics and experiences across groups. \u0026sup1;Exam score\u0026thinsp;=\u0026thinsp;mean grade across Internal Medicine, Emergency Medicine, and Anesthesiology. German grading system: 1\u0026thinsp;=\u0026thinsp;best, 5\u0026thinsp;=\u0026thinsp;worst.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003eCharacteristics and experiences\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTotal \u003c/p\u003e\u003cp\u003e(n\u0026thinsp;=\u0026thinsp;88)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eI1 \u003c/p\u003e\u003cp\u003e(n\u0026thinsp;=\u0026thinsp;28)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eI2 \u003c/p\u003e\u003cp\u003e(n\u0026thinsp;=\u0026thinsp;30)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eCO \u003c/p\u003e\u003cp\u003e(n\u0026thinsp;=\u0026thinsp;30)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003ep-value\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGender, n (%):\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFemale\u003c/p\u003e\u003cp\u003eMale\u003c/p\u003e\u003cp\u003eDiverse\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e63 (72%) \u003c/p\u003e\u003cp\u003e25 (28%) \u003c/p\u003e\u003cp\u003e0 (0%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e18 (64%) \u003c/p\u003e\u003cp\u003e10 (36%) \u003c/p\u003e\u003cp\u003e0 (0%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e22 (73%)\u003c/p\u003e\u003cp\u003e8 (27%)\u003c/p\u003e\u003cp\u003e0 (0%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e23 (77%)\u003c/p\u003e\u003cp\u003e7 (23%)\u003c/p\u003e\u003cp\u003e0 (0%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e.56\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003eAge (years), mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e24.5\u0026thinsp;\u0026plusmn;\u0026thinsp;2.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e25.0\u0026thinsp;\u0026plusmn;\u0026thinsp;3.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e24.9\u0026thinsp;\u0026plusmn;\u0026thinsp;2.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e23.6\u0026thinsp;\u0026plusmn;\u0026thinsp;2.0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e.169\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003eExam score\u0026sup1;, mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.14\u0026thinsp;\u0026plusmn;\u0026thinsp;0.73\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2.34\u0026thinsp;\u0026plusmn;\u0026thinsp;0.83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2.29\u0026thinsp;\u0026plusmn;\u0026thinsp;0.66\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.83\u0026thinsp;\u0026plusmn;\u0026thinsp;0.60\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e.020\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003e3D experience (hours)\u0026sup2;, mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e9.36\u0026thinsp;\u0026plusmn;\u0026thinsp;37.0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e12.0\u0026thinsp;\u0026plusmn;\u0026thinsp;31.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e12.2\u0026thinsp;\u0026plusmn;\u0026thinsp;54.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e4.0\u0026thinsp;\u0026plusmn;\u0026thinsp;11.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e.630\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003eVR experience (hours)\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e, mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.07\u0026thinsp;\u0026plusmn;\u0026thinsp;0.23\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.11\u0026thinsp;\u0026plusmn;\u0026thinsp;0.28\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.05\u0026thinsp;\u0026plusmn;\u0026thinsp;0.20\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.07\u0026thinsp;\u0026plusmn;\u0026thinsp;0.22\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e.641\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cem\u003e\u0026sup2; Refers to the past six months.\u003c/em\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\u003ch2\u003eGeneral and specific IC training effectiveness\u003c/h2\u003e\u003cp\u003eBoth general and specific training modalities resulted in significant gains of IC (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). Participants who received general IC training (I1 and I2) improved their general IC score from 54.9% to 69.1% (p\u0026thinsp;\u0026lt;\u0026thinsp;.001), mean training time 24.1\u0026thinsp;\u0026plusmn;\u0026thinsp;2.3 minutes. Participants in the specific IC training group (I2) increased their specific IC score from 44.7% to 86.2% (p\u0026thinsp;\u0026lt;\u0026thinsp;.001), mean training time 21.7\u0026thinsp;\u0026plusmn;\u0026thinsp;3.3 minutes.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003eImpact of IC training on clinical performance (H1)\u003c/h2\u003e\u003cp\u003eAs the primary outcome, clinical performance in the VR-based emergency scenario differed significantly between groups (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb). Mean performance scores were 19.9% \u0026plusmn; 10.6% in I1 (general training), 28.3% \u0026plusmn; 10.3% in I2 (general\u0026thinsp;+\u0026thinsp;specific training), and 21.2% \u0026plusmn; 10.8% in CO. While performance in I1 did not substantially differ from CO, participants in I2 achieved markedly higher scores than both other groups. This pattern was confirmed by Kruskal-Wallis test (p\u0026thinsp;=\u0026thinsp;.010) and post hoc analyses, indicating that the combined training condition (I2) was the only intervention associated with a clear performance advantage corresponding to a medium-to-large effect size (Cohen\u0026rsquo;s d\u0026thinsp;=\u0026thinsp;0.67) compared to CO. Interrater reliability of checklist scoring was consistently high, with a linear weighted Cohen κ of 0.895 across all items.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003eProcedural efficiency as a mediator (H2a)\u003c/h2\u003e\u003cp\u003eTo evaluate the effect of training on procedural efficiency, SPT was analyzed for one trained procedure (intravenous cannulation) and four untrained transfer tasks (blood culture sampling, temperature measurement, abdominal ultrasound, pneumatic tube operation) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed). For the trained procedure, mean SPT differed significantly between groups (Kruskal\u0026ndash;Wallis: p\u0026thinsp;\u0026lt;\u0026thinsp;.001). Participants in I2 demonstrated the shortest SPT (-0.683\u0026thinsp;\u0026plusmn;\u0026thinsp;0.332), followed by CO (0.260\u0026thinsp;\u0026plusmn;\u0026thinsp;0.853) and group I1 (0.443\u0026thinsp;\u0026plusmn;\u0026thinsp;1.181). For the aggregate SPT of the untrained transfer tasks, group differences were less pronounced but still significant (Kruskal-Wallis: p\u0026thinsp;=\u0026thinsp;.017). SPT was \u0026minus;\u0026thinsp;0.368\u0026thinsp;\u0026plusmn;\u0026thinsp;0.424 in I2, 0.184\u0026thinsp;\u0026plusmn;\u0026thinsp;0.662 in I1, and 0.214\u0026thinsp;\u0026plusmn;\u0026thinsp;0.806 in CO.\u003c/p\u003e\u003cp\u003eAcross groups, SPT for untrained tasks showed a significant main effect on clinical performance in interaction analysis (p\u0026thinsp;\u0026lt;\u0026thinsp;.01), whereas SPT for trained tasks did not. Students with low SPT values (greater efficiency) in untrained tasks achieved higher clinical performance with an effect size of \u003cem\u003eη\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.26 [0.08,1.00]. Expressed as regression equation, clinical performance decreased on average by Y\u0026thinsp;=\u0026thinsp;32.6\u003cb\u003e\u0026ndash;\u003c/b\u003e8.70 for low values of untrained SPT and by Y\u0026thinsp;=\u0026thinsp;32.6\u0026ndash;17.3 for high values. Thus, high SPT values (indicating low efficiency) in untrained tasks were associated with approximately a 50% reduction in clinical performance.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec22\" class=\"Section2\"\u003e\u003ch2\u003eCognitive load as a mediator (H2b)\u003c/h2\u003e\u003cp\u003eObjective cognitive load, assessed via EDA, did not differ significantly across groups during the VR assessment (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ee, left). While there were group differences immediately before the assessment (minutes \u0026minus;\u0026thinsp;3 to -1) with I1 showing the highest median EDA levels, followed by I2 and CO (Kruskal-Wallis: all p\u0026thinsp;\u0026lt;\u0026thinsp;.002), median EDA values converged during the assessment (minutes 0\u0026ndash;10). Transient differences at isolated time points did not remain significant after correction for multiple testing (Kruskal-Wallis after FDR adjustment). However, a trend toward higher EDA in I1 compared to the other groups persisted throughout the assessment period.\u003c/p\u003e\u003cp\u003eDuring the VR assessment (TLX1), subjective NASA-TLX ratings showed a small but significant group difference. Surprisingly, CO reported the lowest subjective cognitive load (I1: 69.4\u0026thinsp;\u0026plusmn;\u0026thinsp;18.9, I2: 68.0\u0026thinsp;\u0026plusmn;\u0026thinsp;15.6, CO: 58.6\u0026thinsp;\u0026plusmn;\u0026thinsp;15.6; Kruskal-Wallis: p\u0026thinsp;=\u0026thinsp;.010). Immediately after the assessment (TLX2), however, subjective cognitive load was similar across groups (I1: 75.8\u0026thinsp;\u0026plusmn;\u0026thinsp;17.8, I2: 80.9\u0026thinsp;\u0026plusmn;\u0026thinsp;13.9, CO: 77.3\u0026thinsp;\u0026plusmn;\u0026thinsp;10.6; Kruskal-Wallis: p\u0026thinsp;=\u0026thinsp;.229) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ee, right).\u003c/p\u003e\u003cp\u003eRegarding interaction, EDA did not show significant main effects within or between groups. In contrast, subjective ratings of cognitive load (TLX1) demonstrated a strong main effect on performance (p\u0026thinsp;\u0026lt;\u0026thinsp;.001): Across groups, students with medium NASA-TLX values achieved the best clinical performance, followed by those with high values (\u003cem\u003eη\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.15 [0.04,1.00]). Expressed as regression equations, clinical performance increased by Y\u0026thinsp;=\u0026thinsp;7.80\u0026thinsp;+\u0026thinsp;18.8 for medium NASA-TLX values and by Y\u0026thinsp;=\u0026thinsp;7.80\u0026thinsp;+\u0026thinsp;15.1 for high values. A similar, though non-significant pattern was observed for the TLX2 when comparing I2 to CO. Again, students with medium NASA-TLX values showed the highest performance, followed by those with high values (\u003cem\u003eη\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.04 [0.00,1.00], regression equation: medium Y\u0026thinsp;=\u0026thinsp;11\u0026thinsp;+\u0026thinsp;16.0, high: Y\u0026thinsp;=\u0026thinsp;11\u0026thinsp;+\u0026thinsp;13.9).\u003c/p\u003e\u003cdiv id=\"Sec23\" class=\"Section3\"\u003e\u003ch2\u003eCorrelation between 3D experience and clinical performance (H3)\u003c/h2\u003e\u003cp\u003eValues of prior 3D experience were not normally distributed across groups (Shapiro-Wilk test, all p\u0026thinsp;\u0026lt;\u0026thinsp;.001) and overall very low, with 75% of participants reporting a value of 0. Spearman correlations indicated small, non-significant associations in CO (ρ\u0026thinsp;=\u0026thinsp;0.165, p\u0026thinsp;=\u0026thinsp;.383) and I1 (ρ\u0026thinsp;=\u0026thinsp;0.162, p\u0026thinsp;=\u0026thinsp;.409), but a moderate, significant correlation in I2 (ρ\u0026thinsp;=\u0026thinsp;0.387, p\u0026thinsp;=\u0026thinsp;.034) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). As the participants had virtually no prior VR experience (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), no correlation with clinical performance was observed for this variable (all ρ\u0026thinsp;\u0026lt;\u0026thinsp;0.100).\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec24\" class=\"Section2\"\u003e\u003ch2\u003eSelf-reported usability barriers and acceptance (H4)\u003c/h2\u003e\u003cp\u003eAcross all groups, enjoyment ratings for the VR scenarios were high (I1\u0026thinsp;=\u0026thinsp;4.39\u0026thinsp;\u0026plusmn;\u0026thinsp;0.69, I2\u0026thinsp;=\u0026thinsp;4.53\u0026thinsp;\u0026plusmn;\u0026thinsp;0.63, CO\u0026thinsp;=\u0026thinsp;4.14\u0026thinsp;\u0026plusmn;\u0026thinsp;0.87; Kruskal-Wallis p\u0026thinsp;=\u0026thinsp;.725), and perception of future applicability of VR-based examinations were predominantly positive (I1\u0026thinsp;=\u0026thinsp;3.11\u0026thinsp;\u0026plusmn;\u0026thinsp;1.13, I2\u0026thinsp;=\u0026thinsp;3.57\u0026thinsp;\u0026plusmn;\u0026thinsp;1.07, CO\u0026thinsp;=\u0026thinsp;3.21\u0026thinsp;\u0026plusmn;\u0026thinsp;1.24; Kruskal-Wallis: p\u0026thinsp;=\u0026thinsp;.284). However, participants in I1 and CO reported significantly more distraction from clinical content due to interface barriers compared to I2 (I1\u0026thinsp;=\u0026thinsp;3.00\u0026thinsp;\u0026plusmn;\u0026thinsp;1.25, I2\u0026thinsp;=\u0026thinsp;2.47\u0026thinsp;\u0026plusmn;\u0026thinsp;1.17, CO\u0026thinsp;=\u0026thinsp;3.38\u0026thinsp;\u0026plusmn;\u0026thinsp;1.37; Kruskal-Wallis: p\u0026thinsp;=\u0026thinsp;.024). Interestingly, distraction scores in CO were correlated negatively with prior 3D experience (r = -0.39, p\u0026thinsp;=\u0026thinsp;.035). This association was weaker and no longer statistically significant in I2 (r = -0.28, p\u0026thinsp;=\u0026thinsp;.136), suggesting that specific IC training attenuated the subjective effect of digital background.\u003c/p\u003e\u003cp\u003eA total of 91 responses to the open-ended question on prerequisites for future acceptance of VR-based examinations were coded for thematic analysis (Supplementary Table\u0026nbsp;3). Three main priorities emerged: (1) Most participants emphasized the need for prior preparation with the VR environment, either theoretically (23/91, 25%) or practically (66/91, 73%), with 78% mentioning at least one of these forms and some additionally suggesting early curricular integration (16/91, 18%). Extensive training was viewed as essential to avoid inequities in operating the simulation. (2) A considerable number suggested program-related improvements to enhance accessibility (17/91, 19%), such as increased fonts or facilitated grasping interactions. (3) Finally, several participants also recommended adjustments to the examination conditions (11/91, 12%) such as extending the time limit or modifying evaluation criteria.\u003c/p\u003e\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis randomized controlled trial provides the first causal evidence that IC - defined as a user\u0026rsquo;s ability to navigate and interact with VR systems - significantly influences clinical performance in VR-based assessments (confirming H1). Participants who received specific IC training achieved higher scores than untrained peers in a complex, time-sensitive medical simulation. From a validity perspective, IC represents construct-irrelevant variance that may compromise the extrapolation inference in assessment design\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. Analogous to challenges encountered in early computer-based testing\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e,\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e, our data indicate that VR-based performance metrics may conflate true clinical competence with interface proficiency. Interestingly, training general interaction techniques alone was insufficient - instead, context-specific training within the target application (specific IC training) proved necessary to achieve measurable performance benefits.\u003c/p\u003e\u003cp\u003eImportantly, our findings extend beyond education. Comparable dynamics are likely in clinical VR applications - such as diagnostics in neurology, psychiatry or geriatrics\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan additionalcitationids=\"CR24\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e - where patients with limited digital abilities may underperform for non-clinical reasons, leading to misclassification or inappropriate therapeutic decisions. Such disparities mirror second-order digital divides in digital health, where user capability - rather than hardware access - limits participation\u003csup\u003e\u003cspan additionalcitationids=\"CR40\" citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. Especially patient populations with little exposure to digital tools, including older adults, people with disabilities, and socioeconomically disadvantaged groups, may therefore experience recurrent disadvantages in VR-based diagnostics and interventions. While such issues are already well recognized in the context of algorithmic bias research\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e,\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e, they remain underexplored in immersive systems for educational and clinical assessments\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e ,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eProcedural efficiency, measured as SPT, emerged as a plausible covariate (confirming H2a). Group effects appeared for both trained and untrained tasks with I2 participants completing more clinically relevant actions within the limited timeframe, which is consistent with effects from other VR trainings\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e,\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. However, only SPT for untrained tasks contributed significantly as a moderator. Thus, clinical outcomes depended less on executing rehearsed steps quickly than on managing novel scenario components efficiently. Similar to clinical performance, general IC training (I1) had little effect on SPT, likely due to the higher cognitive and motor demands of high-fidelity VR environments\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e, for which general engagement with VR interaction techniques was likely insufficient. Comparable patterns have been documented in studies on internet skills and motor tasks, where general abilities show limited transfer to strategic or context-specific applications\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e,\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. Altogether, these findings highlight procedural efficiency as both a sensitive proximal outcome and a potential mechanism for broader skill transfer in immersive simulations\u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e,\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eAlthough cognitive load was also hypothesized as an interacting variable, results were unexpected. Both NASA-TLX and EDA indicated lower load in the CO, yet these participants achieved the weakest clinical outcomes, and no linear mediating role was observed (refuting H2b). Added to that, subjective cognitive load showed an interaction effect, with medium values yielding the best performance\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. Taken together, these observations challenge the assumption that IC training primarily reduces extraneous cognitive burden\u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e, at least when training and assessment occur back-to-back. One explanation is that training itself induced substantial activation, creating a carry-over effect masking downstream reductions in cognitive load\u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e. The high pre-assessment EDA in I1, who trained under the greatest time pressure due to rapid task repetition, support this interpretation. However, cognitive load theory offers an additional explanation\u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e: low TLX1 ratings in CO might also reflect disengagement with task-relevant actions due to difficulties in VR interaction. In contrast, the higher load reported by trained participants might have been primarily germane - arising from active problem solving and procedural engagement\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e,\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e - and thus conducive to superior performance. Although we assessed cognitive load not only retrospectively, but also during the simulation using multiple measures\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e, especially EDA may lack the specificity required to disentangle cognitive from emotional arousal\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e. Thus, more precise methods (e.g., eye-tracking) - which were not available in combination with VR headsets at the time of the study - might provide a more detailed picture in the future.\u003c/p\u003e\u003cp\u003eNon-parametric correlation analyses revealed a pattern opposite to our initial hypothesis that IC training would mitigate digital advantages (refuting H3). Unlike previous studies\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, in our population unexposed to IC training (CO), no relevant association between prior 3D experience and clinical performance was detectable. In contrast, after specific training (I2) higher levels of prior 3D exposure were moderately related to superior performance, just exceeding the significance level. As participants in our groups reported uniformly low 3D experience (mean\u0026thinsp;\u0026lt;\u0026thinsp;10 hours during the past 6 months with 75% reporting a value of 0), leaving little variance to reveal larger systematic effects\u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e, these findings should generally be interpreted with caution. If they represent a genuine effect, training may not function solely as an equalizer, but rather as a facilitator that allows participants to leverage pre-existing competencies once basic interaction barriers are removed. This dynamic parallels the \u0026ldquo;Matthew effect\u0026rdquo; (magnification theory) described in cognitive\u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e and educational research\u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e,\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e, where training may amplify rather than level initial differences, for instance because a certain baseline skill is necessary to benefit. It is also possible that the training dose (25 minutes per module) was insufficient to compensate for substantial disparities in prior digital experience, as the duration of training may determine whether effects are compensatory or amplifying\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e. These considerations underscore the need for future studies to investigate extended or repeated IC as a potential levelling intervention.\u003c/p\u003e\u003cp\u003eHigh enjoyment ratings and generally positive outlook on the fairness of VR-based examinations align with prior reports that immersive technologies are well accepted in medical education when they are perceived as engaging and safe learning environments\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e. At the same time, the pronounced distraction caused by interface barriers in untrained participants (confirming H4) reflects well-documented usability challenges of VR systems, including navigation difficulties, limited haptic feedback, and inconsistent interaction metaphors\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e. Consistent with previous findings, prior 3D experience appeared to buffer these effects subjectively\u003csup\u003e\u003cspan additionalcitationids=\"CR17 CR18\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, underscoring that digital background may inadvertently bias assessment outcomes. Importantly, participants themselves emphasized the need for curricular integration, structured training access, and usability refinements as prerequisites for fair implementation, which converges with broader recommendations for the design of inclusive VR systems in clinical and educational contexts\u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe current study demonstrates that failure to account for IC may reinforce digital inequities while creating a misleading appearance of objectivity. To address this risk, we propose three actionable mitigation strategies (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e):\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003eDesign for inclusive interaction\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eDevelopers should adhere to principles of intuitive interface design - reducing unnecessary complexity, standardizing interaction metaphors, and applying universal design elements\u003csup\u003e65,67\u003c/sup\u003e. These measures help minimize dependence on IC and enable participation by users with diverse digital backgrounds.\u0026nbsp;\u003c/p\u003e\n\u003col start=\"2\"\u003e\n \u003cli\u003eMonitor for bias during implementation\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003ePerformance data should be continuously monitored for correlations with digital background characteristics. Any unexpected disparities may indicate hidden dependencies on IC and require adjustments of the tool or its deployment protocol.\u003c/p\u003e\n\u003col start=\"3\"\u003e\n \u003cli\u003eAddress IC through training or statistical control\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eBefore implementing immersive assessments, it is important to clarify how differences in IC will be addressed. Where feasible, standardized IC training modules can be provided, while monitoring their potential to amplify pre-existing advantages. When training is not feasible or its effects are uncertain, IC should be measured and accounted for in analyses - either as a covariate or as a stratification factor.\u003c/p\u003e\n\u003cp\u003eFuture research should first clarify the role of IC training. It remains uncertain whether training primarily reduces disparities or amplifies pre-existing advantages, and which doses, formats, or contexts promote levelling effects. Identifying reliable training protocols is therefore central to ensuring fair educational and clinical use. Longitudinal studies are also needed to trace how IC develops across educational stages and patient populations, and whether its influence on performance endures over time, particularly as extended reality becomes part of daily life. In parallel, cross-cultural and demographic research should examine how socioeconomic status, age, and technology access shape IC, linking it to broader issues of digital health equity. Finally, studies must determine how IC biases diagnostic conclusions in clinical domains such as neurology, psychiatry, and geriatrics, and how these biases can be mitigated through design improvements, targeted training, or calibration strategies.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLimitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study has several limitations. First, the sample consisted of medical students from a single German institution, which may restrict the generalizability of findings to other educational or clinical populations. Digital fluency and prior exposure to immersive systems likely vary across regions, age groups, and socioeconomic backgrounds.\u003c/p\u003e\n\u003cp\u003eSecond, our study tested a single, relatively brief IC training protocol. It therefore remains unclear whether different training doses, formats, or contexts might yield distinct effects - for example, acting as levelers in some cases and amplifiers in others. This uncertainty limits the generalizability of our findings and highlights the need for systematic variation of IC training designs in future work.\u003c/p\u003e\n\u003cp\u003eThird, the measures of cognitive load have inherent constraints: EDA lacks specificity to distinguish cognitive from emotional or physical arousal, while NASA-TLX is subjective and susceptible to individual response tendencies.\u003c/p\u003e\n\u003cp\u003eFinally, the VR scenarios focused exclusively on internal medicine emergencies. Although valid for this context, the findings may not be directly transferable to other specialties or patient populations.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eImmersive competence may be a critical yet overlooked determinant of user performance in VR-based healthcare applications. Our findings demonstrate that targeted IC training can significantly enhance medical task execution and may help mitigate performance disparities, although its effects may depend on baseline user abilities and training conditions.\u003c/p\u003e\u003cp\u003eWithout accounting for IC, immersive technologies risk introducing unintended bias, compromising both the fairness and validity of high-stakes assessments and digital interventions. As VR continues to expand into clinical diagnostics, rehabilitation, and professional education, systematic strategies for measuring, training, or mitigating IC will be essential. Recognizing immersive competence as a modifiable, equity-relevant factor can help ensure that VR-based systems serve as enablers - rather than barriers - in advancing accessible and just digital health.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eClinical trial number\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study received no funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and analyzed in this study are provided in Supplementary Data 1. The video recordings used for performance ratings can be obtained from the authors upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contribution statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJ.S. conducted the experiments, collected the data, contributed to the analysis, and participated in writing the manuscript. T.L. processed and statistically analyzed the electrodermal activity (EDA) data and prepared the corresponding figures. M.L. supported the experiments and the handling of the EDA devices. A.Z. assisted in the processing of the EDA data. J.B. contributed to the study design and performed the interaction analysis. S.K. contributed to data presentation and manuscript writing. T.M. conceived and designed the study, supervised its conduct, performed the primary data analyses (except for EDA data and interaction analysis), and wrote the manuscript. All authors reviewed and approved the final version.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTM was involved in the software development of the STEP-VR software. All other authors declare no conflicts of interest.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eLiu, J. Y. W. \u003cem\u003eet al.\u003c/em\u003e The Effects of Immersive Virtual Reality Applications on Enhancing the Learning Outcomes of Undergraduate Health Care Students: Systematic Review With Meta-synthesis. \u003cem\u003eJournal of medical Internet research\u003c/em\u003e 25, e39989; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/39989\u003c/span\u003e\u003cspan address=\"10.2196/39989\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCushnan, J., McCafferty, P. \u0026amp; Best, P. Clinicians' perspectives of immersive tools in clinical mental health settings: a systematic scoping review. \u003cem\u003eBMC health services research\u003c/em\u003e 24, 1091; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12913-024-11481-3\u003c/span\u003e\u003cspan address=\"10.1186/s12913-024-11481-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNeher, A. N. \u003cem\u003eet al.\u003c/em\u003e Virtual reality for assessment in undergraduate nursing and medical education - a systematic review. \u003cem\u003eBMC medical education\u003c/em\u003e 25, 292; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12909-025-06867-8\u003c/span\u003e\u003cspan address=\"10.1186/s12909-025-06867-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSchiza, E., Matsangidou, M., Neokleous, K. \u0026amp; Pattichis, C. S. Virtual Reality Applications for Neurological Disease: A Review. \u003cem\u003eFrontiers in robotics and AI\u003c/em\u003e 6, 100; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/frobt.2019.00100\u003c/span\u003e\u003cspan address=\"10.3389/frobt.2019.00100\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGeraets, C. N. W., Wallinius, M. \u0026amp; Sygel, K. Use of Virtual Reality in Psychiatric Diagnostic Assessments: A Systematic Review. \u003cem\u003eFrontiers in psychiatry\u003c/em\u003e 13, 828410; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fpsyt.2022.828410\u003c/span\u003e\u003cspan address=\"10.3389/fpsyt.2022.828410\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eH\u0026oslash;rlyck, L. D., Obenhausen, K., Jansari, A., Ullum, H. \u0026amp; Miskowiak, K. W. Virtual reality assessment of daily life executive functions in mood disorders: associations with neuropsychological and functional measures. \u003cem\u003eJournal of affective disorders\u003c/em\u003e 280, 478\u0026ndash;487; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jad.2020.11.084\u003c/span\u003e\u003cspan address=\"10.1016/j.jad.2020.11.084\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFreeman, D. \u003cem\u003eet al\u003c/em\u003e. Automated psychological therapy using immersive virtual reality for treatment of fear of heights: a single-blind, parallel-group, randomised controlled trial. \u003cem\u003eThe lancet. Psychiatry\u003c/em\u003e 5, 625\u0026ndash;632; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S2215-0366(18)30226-8\u003c/span\u003e\u003cspan address=\"10.1016/S2215-0366(18)30226-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHoward, M. C. A meta-analysis and systematic literature review of virtual reality rehabilitation programs. \u003cem\u003eComputers in Human Behavior\u003c/em\u003e 70, 317\u0026ndash;327; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.chb.2017.01.013\u003c/span\u003e\u003cspan address=\"10.1016/j.chb.2017.01.013\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSpiegel, B. M. R. \u003cem\u003eet al.\u003c/em\u003e Feasibility of combining spatial computing and AI for mental health support in anxiety and depression. \u003cem\u003eNPJ digital medicine\u003c/em\u003e 7, 22; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-024-01011-0\u003c/span\u003e\u003cspan address=\"10.1038/s41746-024-01011-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBargeri, S. \u003cem\u003eet al.\u003c/em\u003e Effectiveness and safety of virtual reality rehabilitation after stroke: an overview of systematic reviews. \u003cem\u003eeClinicalMedicine\u003c/em\u003e 64; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.eclinm.2023.102220\u003c/span\u003e\u003cspan address=\"10.1016/j.eclinm.2023.102220\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWiederhold, B. K. \u0026amp; Wiederhold, M. D. Virtual reality therapy combined with physiological monitoring provides effective treatment, with objective metrics, for post-traumatic stress disorder. \u003cem\u003eExpert review of medical devices\u003c/em\u003e 22, 117\u0026ndash;119; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/17434440.2025.2454930\u003c/span\u003e\u003cspan address=\"10.1080/17434440.2025.2454930\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLie, S. S., Helle, N., Sletteland, N. V., Vikman, M. D. \u0026amp; Bonsaksen, T. Implementation of Virtual Reality in Health Professions Education: Scoping Review. \u003cem\u003eJMIR medical education\u003c/em\u003e 9, e41589; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/41589\u003c/span\u003e\u003cspan address=\"10.2196/41589\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTusher, H. M., Mallam, S. \u0026amp; Nazir, S. A Systematic Review of Virtual Reality Features for Skill Training. \u003cem\u003eTech Know Learn\u003c/em\u003e 29, 843\u0026ndash;878; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10758-023-09713-2\u003c/span\u003e\u003cspan address=\"10.1007/s10758-023-09713-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLaViola, J. J., Kruijff, E., McMahan, R. P., Bowman, D. \u0026amp; Poupyrev, I. P. \u003cem\u003e3D User Interfaces: Theory and Practice\u003c/em\u003e. 2nd ed. (Addison-Wesley, Boston, 2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTuena, C. \u003cem\u003eet al.\u003c/em\u003e Usability Issues of Clinical and Research Applications of Virtual Reality in Older People: A Systematic Review. \u003cem\u003eFrontiers in human neuroscience\u003c/em\u003e 14, 93; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fnhum.2020.00093\u003c/span\u003e\u003cspan address=\"10.3389/fnhum.2020.00093\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKatz, D. \u003cem\u003eet al.\u003c/em\u003e Relationship between demographic and social variables and performance in virtual reality among healthcare personnel: an observational study. \u003cem\u003eBMC medical education\u003c/em\u003e 24, 227; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12909-024-05180-0\u003c/span\u003e\u003cspan address=\"10.1186/s12909-024-05180-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOberd\u0026ouml;rfer, S. \u003cem\u003eet al.\u003c/em\u003e Ready for VR? Assessing VR Competence and Exploring the Role of Human Abilities and Characteristics. (Preprint) (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSchreiner, V. \u003cem\u003eet al. Specific Immersive Competence in VR-Based Assessments: Development, Psychometric Evaluation and Associations with Medical Performance (Preprint)\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSchlickum, M. K., Hedman, L., Enochsson, L., Kjellin, A. \u0026amp; Fell\u0026auml;nder-Tsai, L. Systematic video game training in surgical novices improves performance in virtual reality endoscopic surgical simulators: a prospective randomized study. \u003cem\u003eWorld journal of surgery\u003c/em\u003e 33, 2360\u0026ndash;2367; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00268-009-0151-y\u003c/span\u003e\u003cspan address=\"10.1007/s00268-009-0151-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2009).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSteed, A. \u003cem\u003eet al.\u003c/em\u003e Immersive competence and immersive literacy: Exploring how users learn about immersive experiences. \u003cem\u003eFront. Virtual Real.\u003c/em\u003e 4; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/frvir.2023.1129242\u003c/span\u003e\u003cspan address=\"10.3389/frvir.2023.1129242\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMessick, S. Standards of Validity and the Validity of Standards in Performance Asessment. \u003cem\u003eEducational Measurement\u003c/em\u003e 14, 5\u0026ndash;8; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/j.1745-3992.1995.tb00881.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1745-3992.1995.tb00881.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (1995).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKane, M. T. Validating the Interpretations and Uses of Test Scores. \u003cem\u003eJ Educational Measurement\u003c/em\u003e 50, 1\u0026ndash;73; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/jedm.12000\u003c/span\u003e\u003cspan address=\"10.1111/jedm.12000\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2013).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDu, K., Benavides, L. R., Isenstein, E. L., Tadin, D. \u0026amp; Busza, A. C. Virtual reality assessment of reaching accuracy in patients with recent cerebellar stroke. \u003cem\u003eBMC Digit Health\u003c/em\u003e 2, 50; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s44247-024-00107-7\u003c/span\u003e\u003cspan address=\"10.1186/s44247-024-00107-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBell, I. H., Nicholas, J., Alvarez-Jimenez, M., Thompson, A. \u0026amp; Valmaggia, L. Virtual reality as a clinical tool in mental health research and practice. \u003cem\u003eDialogues in clinical neuroscience\u003c/em\u003e 22, 169\u0026ndash;177; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.31887/DCNS.2020.22.2/lvalmaggia\u003c/span\u003e\u003cspan address=\"10.31887/DCNS.2020.22.2/lvalmaggia\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAtkins, A. S. \u003cem\u003eet al.\u003c/em\u003e Assessment of Age-Related Differences in Functional Capacity Using the Virtual Reality Functional Capacity Assessment Tool (VRFCAT). \u003cem\u003eThe journal of prevention of Alzheimer's disease\u003c/em\u003e 2, 121\u0026ndash;127; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.14283/jpad.2015.61\u003c/span\u003e\u003cspan address=\"10.14283/jpad.2015.61\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2015).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eM\u0026uuml;hling, T. \u003cem\u003eet al.\u003c/em\u003e Virtual reality in medical emergencies training: benefits, perceived stress, and learning success. \u003cem\u003eMultimedia Systems\u003c/em\u003e 29, 2239\u0026ndash;2252; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00530-023-01102-0\u003c/span\u003e\u003cspan address=\"10.1007/s00530-023-01102-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWeiss, S. L. \u003cem\u003eet al\u003c/em\u003e. Surviving Sepsis Campaign International Guidelines for the Management of Septic Shock and Sepsis-Associated Organ Dysfunction in Children. \u003cem\u003ePediatric critical care medicine: a journal of the Society of Critical Care Medicine and the World Federation of Pediatric Intensive and Critical Care Societies\u003c/em\u003e 21, e52-e106; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/PCC.0000000000002198\u003c/span\u003e\u003cspan address=\"10.1097/PCC.0000000000002198\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eM\u0026uuml;hling, T., Schreiner, V., Appel, M., Leutritz, T. \u0026amp; K\u0026ouml;nig, S. Comparing Virtual Reality-Based and Traditional Physical Objective Structured Clinical Examination (OSCE) Stations for Clinical Competency Assessments: Randomized Controlled Trial. \u003cem\u003eJournal of medical Internet research\u003c/em\u003e 27, e55066; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/55066\u003c/span\u003e\u003cspan address=\"10.2196/55066\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHart, S. G. \u0026amp; Staveland, L. E. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In \u003cem\u003eAdvances in Psychology: Human Mental Workload\u003c/em\u003e, edited by P. A. Hancock \u0026amp; N. Meshkati (North-Holland1988), Vol. 52, pp. 139\u0026ndash;183.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFl\u0026auml;gel, K., Galler, B., Steinh\u0026auml;user, J. \u0026amp; G\u0026ouml;tz, K. Der \u0026bdquo;National Aeronautics and Space Administration-Task Load Index\u0026ldquo; (NASA-TLX) \u0026ndash; ein Instrument zur Erfassung der Arbeitsbelastung in der haus\u0026auml;rztlichen Sprechstunde: Bestimmung der psychometrischen Eigenschaften. \u003cem\u003eZeitschrift fur Evidenz, Fortbildung und Qualitat im Gesundheitswesen\u003c/em\u003e 147\u0026ndash;148, 90\u0026ndash;96; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.zefq.2019.10.003\u003c/span\u003e\u003cspan address=\"10.1016/j.zefq.2019.10.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKleckner, I. R. \u003cem\u003eet al.\u003c/em\u003e Simple, Transparent, and Flexible Automated Quality Assessment Procedures for Ambulatory Electrodermal Activity Data. \u003cem\u003eIEEE transactions on bio-medical engineering\u003c/em\u003e 65, 1460\u0026ndash;1467; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TBME.2017.2758643\u003c/span\u003e\u003cspan address=\"10.1109/TBME.2017.2758643\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLooff, P. de \u003cem\u003eet al.\u003c/em\u003e Wearables: An R Package With Accompanying Shiny Application for Signal Analysis of a Wearable Device Targeted at Clinicians and Researchers. \u003cem\u003eFrontiers in behavioral neuroscience\u003c/em\u003e 16, 856544; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fnbeh.2022.856544\u003c/span\u003e\u003cspan address=\"10.3389/fnbeh.2022.856544\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePosada-Quintero, H. F. \u0026amp; Chon, K. H. Innovations in Electrodermal Activity Data Collection and Signal Processing: A Systematic Review. \u003cem\u003eSensors (Basel, Switzerland)\u003c/em\u003e 20; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s20020479\u003c/span\u003e\u003cspan address=\"10.3390/s20020479\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBraun, V. \u0026amp; Clarke, V. Using thematic analysis in psychology. \u003cem\u003eQualitative Research in Psychology\u003c/em\u003e 3, 77\u0026ndash;101; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1191/1478088706qp063oa\u003c/span\u003e\u003cspan address=\"10.1191/1478088706qp063oa\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2006).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHanley, J. A. Simple and multiple linear regression: sample size considerations. \u003cem\u003eJournal of clinical epidemiology\u003c/em\u003e 79, 112\u0026ndash;119; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jclinepi.2016.05.014\u003c/span\u003e\u003cspan address=\"10.1016/j.jclinepi.2016.05.014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2016).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEysenbach, G. CONSORT-EHEALTH: improving and standardizing evaluation reports of Web-based and mobile health interventions. \u003cem\u003eJournal of medical Internet research\u003c/em\u003e 13, e126; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/jmir.1923\u003c/span\u003e\u003cspan address=\"10.2196/jmir.1923\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eValentine, A., Vrbik, P. \u0026amp; Thomas, R. A systematic review of paper-based versus computer-based testing in engineering and computing education, 364\u0026ndash;372; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/EDUCON52537.2022.9766469\u003c/span\u003e\u003cspan address=\"10.1109/EDUCON52537.2022.9766469\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKhoshsima, H., Hosseini, M. \u0026amp; Toroujeni, S. M. H. Cross-Mode Comparability of Computer-Based Testing (CBT) Versus Paper-Pencil Based Testing (PPT): An Investigation of Testing Administration Mode among Iranian Intermediate EFL Learners. \u003cem\u003eELT\u003c/em\u003e 10, 23; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5539/elt.v10n2p23\u003c/span\u003e\u003cspan address=\"10.5539/elt.v10n2p23\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAdedinsewo, D. \u003cem\u003eet al.\u003c/em\u003e Health Disparities, Clinical Trials, and the Digital Divide. \u003cem\u003eMayo Clinic proceedings\u003c/em\u003e 98, 1875\u0026ndash;1887; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.mayocp.2023.05.003\u003c/span\u003e\u003cspan address=\"10.1016/j.mayocp.2023.05.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShaw, J., Brewer, L. C. \u0026amp; Veinot, T. Recommendations for Health Equity and Virtual Care Arising From the COVID-19 Pandemic: Narrative Review. \u003cem\u003eJMIR formative research\u003c/em\u003e 5, e23233; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/23233\u003c/span\u003e\u003cspan address=\"10.2196/23233\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRiggins, F. \u0026amp; Dewan, S. The Digital Divide: Current and Future Research Directions. \u003cem\u003eJAIS\u003c/em\u003e 6, 298\u0026ndash;337; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.17705/1jais.00074\u003c/span\u003e\u003cspan address=\"10.17705/1jais.00074\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2005).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVayena, E., Blasimme, A. \u0026amp; Cohen, I. G. Machine learning in medicine: Addressing ethical challenges. \u003cem\u003ePLoS medicine\u003c/em\u003e 15, e1002689; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pmed.1002689\u003c/span\u003e\u003cspan address=\"10.1371/journal.pmed.1002689\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eObermeyer, Z., Powers, B., Vogeli, C. \u0026amp; Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. \u003cem\u003eScience (New York, N.Y.)\u003c/em\u003e 366, 447\u0026ndash;453; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1126/science.aax2342\u003c/span\u003e\u003cspan address=\"10.1126/science.aax2342\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBinstadt, E., Donner, S., Nelson, J., Flottemesch, T. \u0026amp; Hegarty, C. Simulator training improves fiber-optic intubation proficiency among emergency medicine residents. \u003cem\u003eAcademic emergency medicine: official journal of the Society for Academic Emergency Medicine\u003c/em\u003e 15, 1211\u0026ndash;1214; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/j.1553-2712.2008.00199.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1553-2712.2008.00199.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLohre, R. \u003cem\u003eet al.\u003c/em\u003e Effectiveness of Immersive Virtual Reality on Orthopedic Surgical Skills and Knowledge Acquisition Among Senior Surgical Residents: A Randomized Clinical Trial. \u003cem\u003eJAMA network open\u003c/em\u003e 3, e2031217; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamanetworkopen.2020.31217\u003c/span\u003e\u003cspan address=\"10.1001/jamanetworkopen.2020.31217\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eArthur, T. \u003cem\u003eet al.\u003c/em\u003e Examining the validity and fidelity of a virtual reality simulator for basic life support training. \u003cem\u003eBMC Digit Health\u003c/em\u003e 1; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s44247-023-00016-1\u003c/span\u003e\u003cspan address=\"10.1186/s44247-023-00016-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003evan Deursen, A. \u0026amp; van Dijk, J. Internet skills and the digital divide. \u003cem\u003eNew Media \u0026amp; Society\u003c/em\u003e 13, 893\u0026ndash;911; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1177/1461444810386774\u003c/span\u003e\u003cspan address=\"10.1177/1461444810386774\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWulf, G. \u0026amp; Shea, C. H. Principles derived from the study of simple skills do not generalize to complex skill learning. \u003cem\u003ePsychonomic bulletin \u0026amp; review\u003c/em\u003e 9, 185\u0026ndash;211; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3758/BF03196276\u003c/span\u003e\u003cspan address=\"10.3758/BF03196276\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. (2002).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTorkington, J., Smith, S. G., Rees, B. I. \u0026amp; Darzi, A. Skill transfer from virtual reality to a real laparoscopic task. \u003cem\u003eSurgical endoscopy\u003c/em\u003e 15, 1076\u0026ndash;1079; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s004640000233\u003c/span\u003e\u003cspan address=\"10.1007/s004640000233\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2001).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eClarke, D. B. \u003cem\u003eet al.\u003c/em\u003e Knowledge transfer and retention of simulation-based learning for neurosurgical instruments: a randomised trial of perioperative nurses. \u003cem\u003eBMJ simulation \u0026amp; technology enhanced learning\u003c/em\u003e 7, 146\u0026ndash;153; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmjstel-2019-000576\u003c/span\u003e\u003cspan address=\"10.1136/bmjstel-2019-000576\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePyke, W., Lunau, J. \u0026amp; Javadi, A.-H. Does difficulty moderate learning? A comparative analysis of the desirable difficulties framework and cognitive load theory. \u003cem\u003eQuarterly journal of experimental psychology (2006)\u003c/em\u003e, 17470218241308143; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1177/17470218241308143\u003c/span\u003e\u003cspan address=\"10.1177/17470218241308143\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMarsh, W. E., Hantel, T., Zetzsche, C. \u0026amp; Schill, K. Is the user trained? Assessing performance and cognitive resource demands in the Virtusphere, 15\u0026ndash;22; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/3DUI.2013.6550191\u003c/span\u003e\u003cspan address=\"10.1109/3DUI.2013.6550191\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2013).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eR\u0026ouml;lfing, J. D., N\u0026oslash;rskov, J. K., Paltved, C., Konge, L. \u0026amp; Andersen, S. A. W. Failure affects subjective estimates of cognitive load through a negative carry-over effect in virtual reality simulation of hip fracture surgery. \u003cem\u003eAdvances in simulation (London, England)\u003c/em\u003e 4, 26; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s41077-019-0114-9\u003c/span\u003e\u003cspan address=\"10.1186/s41077-019-0114-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChandler, P. \u0026amp; Sweller, J. Cognitive Load Theory and the Format of Instruction. \u003cem\u003eCognition and Instruction\u003c/em\u003e 8, 293\u0026ndash;332; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1207/s1532690xci0804_2\u003c/span\u003e\u003cspan address=\"10.1207/s1532690xci0804_2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (1991).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLeppink, J., Paas, F., van Gog, T., van der Vleuten, C. P. \u0026amp; van Merri\u0026euml;nboer, J. J. Effects of pairs of problems and examples on task performance and different types of cognitive load. \u003cem\u003eLearning and Instruction\u003c/em\u003e 30, 32\u0026ndash;42; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.learninstruc.2013.12.001\u003c/span\u003e\u003cspan address=\"10.1016/j.learninstruc.2013.12.001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYoung, J. Q., van Merrienboer, J., Durning, S. \u0026amp; Cate, O. ten. Cognitive Load Theory: implications for medical education: AMEE Guide No. 86. \u003cem\u003eMedical teacher\u003c/em\u003e 36, 371\u0026ndash;384; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3109/0142159X.2014.889290\u003c/span\u003e\u003cspan address=\"10.3109/0142159X.2014.889290\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNaismith, L. M. \u0026amp; Cavalcanti, R. B. Validity of Cognitive Load Measures in Simulation-Based Training: A Systematic Review. \u003cem\u003eAcademic medicine: journal of the Association of American Medical Colleges\u003c/em\u003e 90, S24-35; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/ACM.0000000000000893\u003c/span\u003e\u003cspan address=\"10.1097/ACM.0000000000000893\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2015).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAyres, P., Lee, J. Y., Paas, F. \u0026amp; van Merri\u0026euml;nboer, J. J. G. The Validity of Physiological Measures to Identify Differences in Intrinsic Cognitive Load. \u003cem\u003eFrontiers in psychology\u003c/em\u003e 12, 702538; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fpsyg.2021.702538\u003c/span\u003e\u003cspan address=\"10.3389/fpsyg.2021.702538\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHalbig, A. \u0026amp; Latoschik, M. E. A Systematic Review of Physiological Measurements, Factors, Methods, and Applications in Virtual Reality. \u003cem\u003eFront. Virtual Real.\u003c/em\u003e 2; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/frvir.2021.694567\u003c/span\u003e\u003cspan address=\"10.3389/frvir.2021.694567\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGlass, G. V. \u0026amp; Hopkins, K. D. \u003cem\u003eStatistical Methods in Education and Psychology\u003c/em\u003e. 3rd ed. (Allyn and Bacon, 1996).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eL\u0026ouml;vd\u0026eacute;n, M., Brehmer, Y., Li, S.-C. \u0026amp; Lindenberger, U. Training-induced compensation versus magnification of individual differences in memory performance. \u003cem\u003eFrontiers in human neuroscience\u003c/em\u003e 6, 141; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fnhum.2012.00141\u003c/span\u003e\u003cspan address=\"10.3389/fnhum.2012.00141\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBast, J. \u0026amp; Reitsma, P. Analyzing the development of individual differences in terms of Matthew effects in reading: results from a Dutch Longitudinal study. \u003cem\u003eDevelopmental psychology\u003c/em\u003e 34, 1373\u0026ndash;1399; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1037/0012-1649.34.6.1373\u003c/span\u003e\u003cspan address=\"10.1037/0012-1649.34.6.1373\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (1998).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMcVey, R. \u003cem\u003eet al.\u003c/em\u003e Baseline Laparoscopic Skill May Predict Baseline Robotic Skill and Early Robotic Surgery Learning Curve. \u003cem\u003eJournal of endourology\u003c/em\u003e 30, 588\u0026ndash;592; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1089/end.2015.0774\u003c/span\u003e\u003cspan address=\"10.1089/end.2015.0774\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2016).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu, L. \u003cem\u003eet al.\u003c/em\u003e Dose-response relationship between computerized cognitive training and cognitive improvement. \u003cem\u003eNPJ digital medicine\u003c/em\u003e 7, 214; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-024-01210-9\u003c/span\u003e\u003cspan address=\"10.1038/s41746-024-01210-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRadianti, J., Majchrzak, T. A., Fromm, J. \u0026amp; Wohlgenannt, I. A systematic review of immersive virtual reality applications for higher education: Design elements, lessons learned, and research agenda. \u003cem\u003eComputers \u0026amp; Education\u003c/em\u003e 147, 103778; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.compedu.2019.103778\u003c/span\u003e\u003cspan address=\"10.1016/j.compedu.2019.103778\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eM\u0026uuml;hling, T., Backhaus, J., Demmler, L. \u0026amp; K\u0026ouml;nig, S. How Personality and Affective Responses Are Associated with Skepticism Towards Virtual Reality in Medical Training-A Pre-Post Intervention Study. \u003cem\u003eCyberpsychology, behavior and social networking\u003c/em\u003e 28, 335\u0026ndash;341; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1089/cyber.2024.0567\u003c/span\u003e\u003cspan address=\"10.1089/cyber.2024.0567\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChamusca, I. L. \u003cem\u003eet al.\u003c/em\u003e Evaluating Design Guidelines for Intuitive, Therefore Sustainable, Virtual Reality Authoring Tools. \u003cem\u003eSustainability\u003c/em\u003e 16, 1744; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/su16051744\u003c/span\u003e\u003cspan address=\"10.3390/su16051744\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7660457/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7660457/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eVirtual reality (VR) is increasingly used for assessment in educational and clinical settings. However, users\u0026rsquo; immersive competence (IC) - the ability to navigate and operate VR systems - may introduce bias unrelated to clinical skills or patient functioning, a relationship that remains unexplored.\u003c/p\u003e\u003cp\u003eIn this randomized controlled trial, 88 medical students received either general IC training, general plus specific IC training, or no structured training before completing a VR-based assessment scenario. Multimodal outcome data were collected, including physiological stress markers (electrodermal activity), cognitive-load ratings (NASA-TLX), procedural efficiency, and self-reported usability barriers.\u003c/p\u003e\u003cp\u003eSpecific IC training improved performance compared with control (28.3%\u0026plusmn;10.3% vs. 21.2%\u0026plusmn;10.8%, p\u0026thinsp;=\u0026thinsp;.010, d\u0026thinsp;=\u0026thinsp;0.67), partly mediated by procedural efficiency (η\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.26) and increased cognitive load (η\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.15). Prior experience with 3D applications was unrelated to performance in the control group (ρ\u0026thinsp;=\u0026thinsp;0.165, p\u0026thinsp;=\u0026thinsp;.383) but significantly associated with higher performance in the specific training group (ρ\u0026thinsp;=\u0026thinsp;0.387, p\u0026thinsp;=\u0026thinsp;.034). Participants reported high enjoyment, although interface barriers distracted untrained users.\u003c/p\u003e\u003cp\u003eThese findings indicate that IC is a causal, modifiable factor in VR-based assessments and should be accounted for to ensure fair and valid evaluations.\u003c/p\u003e","manuscriptTitle":"Uncovering Immersive Competence as a Hidden Bias in VR-Based Clinical Assessment – A Randomized Controlled Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-17 02:14:39","doi":"10.21203/rs.3.rs-7660457/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-11-13T21:09:50+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-11-10T18:34:37+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-27T05:55:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"23822936305203754563908596647975485567","date":"2025-10-20T01:10:06+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"153578792362499634541079586860750923312","date":"2025-10-04T03:37:04+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-10-03T14:47:56+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-09-28T00:28:06+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-09-26T07:14:14+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Digital Medicine","date":"2025-09-19T16:48:18+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b8ef0c69-5239-46aa-98ac-59077c3dc550","owner":[],"postedDate":"October 17th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":56296169,"name":"Health sciences/Health care"},{"id":56296170,"name":"Health sciences/Medical research"},{"id":56296171,"name":"Biological sciences/Neuroscience"},{"id":56296172,"name":"Biological sciences/Psychology"},{"id":56296173,"name":"Social science/Psychology"}],"tags":[],"updatedAt":"2026-03-16T16:05:09+00:00","versionOfRecord":{"articleIdentity":"rs-7660457","link":"https://doi.org/10.1038/s41746-026-02482-z","journal":{"identity":"npj-digital-medicine","isVorOnly":false,"title":"npj Digital Medicine"},"publishedOn":"2026-03-09 15:59:40","publishedOnDateReadable":"March 9th, 2026"},"versionCreatedAt":"2025-10-17 02:14:39","video":"","vorDoi":"10.1038/s41746-026-02482-z","vorDoiUrl":"https://doi.org/10.1038/s41746-026-02482-z","workflowStages":[]},"version":"v1","identity":"rs-7660457","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7660457","identity":"rs-7660457","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.