Vocal Dynamics as Predictors of Competitive Success in E-Sports: Evidence from Professional Counter-Strike Matches

preprint OA: closed CC-BY-4.0

Abstract

Abstract Objective To identify acoustic and contextual predictors of competitive success in professional Counter-Strike: Global Offensive (CS:GO) players by analyzing voice-based biomarkers of emotional and cognitive dynamics. Methods Naturalistic voice recordings from official matches were processed to extract 68 temporal, spectral, and cepstral features using the PyAudioAnalysis library. Logistic regression models tested associations between these acoustic parameters and match outcomes (win/loss), first considering voice-only predictors and then combining them with contextual performance indicators (team ranking, opponent ranking, and ranking difference). Results Two vocal features—Chroma₁ and ΔMFCC₁₃—emerged as significant predictors of victory, indicating that greater tonal organization and spectral variability were associated with winning outcomes. The inclusion of contextual ranking variables improved model fit (AUC = 0.787), yet both acoustic predictors remained significant, demonstrating that vocal expression contributes unique information beyond historical performance. Conclusion The findings suggest that voice dynamics during competitive play reflect real-time affective and cognitive processes linked to arousal regulation, coordination, and engagement. By integrating behavioral acoustics with contextual performance data, this study advances a multimodal framework for understanding and predicting human performance in high-pressure environments such as e-sports.
Full text 112,451 characters · extracted from preprint-html · click to expand
Vocal Dynamics as Predictors of Competitive Success in E-Sports: Evidence from Professional Counter-Strike Matches | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Vocal Dynamics as Predictors of Competitive Success in E-Sports: Evidence from Professional Counter-Strike Matches Raphael Santos, Gabriel Kadri, Felipe Aguiar, Victor Otani, Ricardo Uchida, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9032928/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Objective To identify acoustic and contextual predictors of competitive success in professional Counter-Strike: Global Offensive (CS:GO) players by analyzing voice-based biomarkers of emotional and cognitive dynamics. Methods Naturalistic voice recordings from official matches were processed to extract 68 temporal, spectral, and cepstral features using the PyAudioAnalysis library. Logistic regression models tested associations between these acoustic parameters and match outcomes (win/loss), first considering voice-only predictors and then combining them with contextual performance indicators (team ranking, opponent ranking, and ranking difference). Results Two vocal features—Chroma₁ and ΔMFCC₁₃—emerged as significant predictors of victory, indicating that greater tonal organization and spectral variability were associated with winning outcomes. The inclusion of contextual ranking variables improved model fit (AUC = 0.787), yet both acoustic predictors remained significant, demonstrating that vocal expression contributes unique information beyond historical performance. Conclusion The findings suggest that voice dynamics during competitive play reflect real-time affective and cognitive processes linked to arousal regulation, coordination, and engagement. By integrating behavioral acoustics with contextual performance data, this study advances a multimodal framework for understanding and predicting human performance in high-pressure environments such as e-sports. Voice analysis E-sports Competitive performance Acoustic biomarkers Cognitive load Communication dynamics INTRODUCTION Competitive gaming (e-sports) has emerged as a complex psychophysiological environment in which cognitive control, emotional regulation, and team coordination are constantly challenged under high-pressure conditions. In games such as Counter-Strike: Global Offensive (CS:GO), players are required to process rapidly changing sensory information, maintain attentional focus, and make split-second decisions while engaging in continuous verbal communication with teammates. This dynamic scenario provides a unique opportunity to study how neurophysiological and behavioral processes—traditionally examined in laboratory contexts—manifest in naturalistic, high-stakes performance environments. Among the multiple channels of behavioral expression available in such settings, the human voice is particularly informative. Vocal output reflects moment-to-moment variations in arousal, cognitive load, and affective states, providing a continuous and non-invasive window into internal processes. Previous studies have shown that acoustic features such as fundamental frequency, intensity, spectral entropy, and formant structure can index autonomic activation, stress reactivity, and emotional valence across a range of contexts(Sondhi, 2015 ; Schewski 2025). In competitive or stressful conditions, modulations in these parameters are often observed, reflecting physiological mobilization and emotional tension(Alvear, 2013). Thus, voice analysis represents a promising avenue for capturing fine-grained affective and cognitive dynamics during real-world performance. While voice analytics has been increasingly explored in health domains—such as mental health monitoring(Briganti, 2025 ), fatigue detection(Gao, 2022), and cognitive workload assessment(Boyer, 2018)—its application in competitive gaming and performance prediction remains scarce. Most existing research on e-sports has focused on overt behavioral metrics (e.g., reaction time, accuracy, strategy patterns)(Saygı, 2023 ; Ersin, 2022) or physiological indices such as heart rate variability and electrodermal activity(Berga, 2023; Behnke, 2025). Given that communication plays a central role in coordination and decision-making during gameplay, understanding its acoustic underpinnings could provide valuable insights into collective performance and stress regulation under pressure. Moreover, the competitive context of professional e-sports allows for quantitative contextualization through objective measures such as global team ranking and opponent strength. These metrics make it possible to link real-world performance indicators with psychophysiological proxies, establishing a bridge between affective neuroscience, behavioral data science, and sports analytics. By combining vocal acoustics with contextual ranking variables, one can test whether subtle variations in vocal tone, energy, or spectral balance carry predictive value for match outcomes beyond historical performance alone. Despite these opportunities, no studies to date have systematically modeled the relationship between vocal features and competitive success in professional gaming. The current literature remains largely descriptive or focused on laboratory-based emotional speech paradigms (Berga, 2023; Behnke, 2025), limiting its ecological validity. Addressing this gap could not only advance the understanding of stress and communication dynamics in e-sports but also inform broader models of performance monitoring applicable to high-demand occupations such as aviation, military operations, and emergency response, where real-time vocal analysis may serve as an unobtrusive marker of cognitive and affective state. The present study was designed to explore these associations using a dataset of professional CS:GO matches, in which the same team’s in-game voice communications were recorded across multiple official and scrimmage matches. For each audio segment, 68 acoustic features encompassing temporal, spectral, and cepstral domains were extracted. These features were analyzed in relation to both contextual competitive variables—team and opponent rankings—and objective match outcomes. Our analytic strategy involved multiple regression models to test whether specific acoustic signatures were associated with competitive success, both independently and in combination with ranking indicators. We hypothesized that increased vocal energy and spectral complexity would reflect heightened arousal and engagement (Goudbeek, 2010 ), potentially predicting successful outcomes, whereas flatter spectral profiles might be linked to fatigue or cognitive overload during disadvantageous situations(Tran, 2020). Furthermore, we expected that integrating contextual ranking variables would enhance predictive accuracy, capturing interactions between emotional expression, physiological activation, and situational challenge within the naturalistic dynamics of professional competition. METHODS AND ANALYSIS Sampling methods, participants, and study design The present study is part of the Performance Prediction in E-Sports Players through Acoustic Analysis of Voice, Speech, and Facial Expression research initiative, conducted at the Department of Mental Health, Santa Casa de São Paulo School of Medical Sciences, Brazil. This project aims to identify behavioral biomarkers associated with performance and emotional states in competitive e-sports environments. For the current analysis, we focused exclusively on voice recordings collected from professional Counter-Strike: Global Offensive (CS:GO) players during official matches. The audio data was provided by a professional team that routinely records in-game communication for strategic review and performance evaluation. The dataset comprises recordings from five professional e-sports players belonging to the same team, which competes in ranked professional leagues. Inclusion criteria required that participants be at least 18 years old, have an active competitive ranking, and provide written consent for the use of their voice recordings for research purposes. Players reporting neurological, psychiatric, or speech/hearing impairments were excluded. The study followed a cross-sectional retrospective design, analyzing previously collected voice data from multiple official matches per team (83 matches), ensuring adequate data variability and statistical power for multivariate modeling. All data were de-identified prior to analysis and processed in accordance with institutional ethical approval (CAAE: 72933023.5.0000.5479). Demographic and Contextual Variables Basic contextual information was collected for each match, including the team’s international ranking, opponent ranking, and the difference between team rankings. These variables were extracted from publicly available e-sports databases and official competition records. Match outcome (0 = loss, 1 = win) was used as the dependent variable, while acoustic and contextual features served as independent predictors. Voice Data Acquisition and Preprocessing Voice data were extracted from original team communication recordings obtained during official Counter-Strike: Global Offensive (CS:GO) professional matches. Each audio file contained natural, spontaneous in-game speech used for coordination and strategic communication among team members. This dataset therefore represents a highly ecological measure of vocal expression and interpersonal synchrony during competitive performance. All raw files were acquired in .mp4 format and converted to 16-bit, 44.1 kHz mono .wav files using the PyDub library (v0.25.1). After conversion, only the first 60 seconds of each recording were retained for analysis, which could include the warm-up period, the pre-game phase, and/or the initial moments of the match. Non-speech segments and background noise were automatically removed through amplitude-based silence trimming, applying a −40 dB threshold and a minimum silence duration of 100 ms to ensure that only active speech periods were analyzed. Subsequently, acoustic features were extracted using the PyAudioAnalysis library (Giannakopoulos, 2015 ) in Python (v3.12). A total of 34 baseline acoustic descriptors were computed for each recording (see Table 1 ), and first-order derivatives (Δ) were calculated for every feature, resulting in a comprehensive set of 68 parameters per sample. These descriptors encompass time-domain, spectral-domain, cepstral-domain, and tonal representations of the voice, thereby providing a multidimensional characterization of vocal dynamics. For each match, acoustic features were averaged across frames, producing a single representative vector per observation. The resulting dataset included both acoustic features and contextual variables (match outcome, team ranking, opponent ranking, and ranking difference). All features were z-score normalized prior to regression analyses, ensuring comparability across participants and matches. Analyses were conducted to capture differences in vocal behavior associated with active performance. Table 1 Acoustic features extracted by the PyAudioAnalysis library. When deriving each of the characteristics to the first order, a total of 68 features are obtained. Adapted from Giannakopoulos T. ( 2015 ). Index Name Description 1 Zero Crossing Rate The rate of sign-changes of the signal during the duration of a particular frame. 2 Energy The sum of squares of the signal values, normalised by the length. 3 Entropy of Energy The entropy of sub-frames normalized energies. It can be interpreted as a measure of abrupt changes. 4 Spectral Centroid The centre of gravity of the spectrum. 5 Spectral Spread The second of gravity of the spectrum. 6 Spectral Entropy Entropy of the normalized spectral energies for a set of sub-frames 7 Spectral Flux The squared difference between the normalised magnitudes of the spectra of the two successive frames. 8 Spectral Rolloff The frequency below which 90% of the magnitude distribution of the spectrum is concentrated. 9–21 MFCCs Mel Frequency Cepstral Coefficients form a cepstral representation where the frequency bands are not linear but distributed according to the mel-scale. 22–33 Chroma Vector A 12-element representation of the spectral energy where the bins represent the 12 equal-tempered pitch classes of western-type music (semitone spacing). 34 Chroma Deviation The standard deviation of the 12 chroma coefficients. INSERT Table 1 HERE Statistical analysis The primary analytic strategy was based on classical logistic regression models, aiming to test associations between acoustic features and competitive outcomes (win vs. loss). Given the high dimensionality and potential multicollinearity of the acoustic dataset, analyses were conducted in two sequential stages: (1) voice-only models to isolate the predictive value of vocal biomarkers, and (2) combined models integrating both acoustic and contextual performance variables (team ranking, opponent ranking, and ranking difference). This two-step approach allowed us to contrast the intrinsic predictive capacity of behavioral voice parameters with the incremental explanatory power of historical team performance indicators. Regression analyses All regression analyses were performed using Python (v3.12) and Jamovi® 2.5. For the voice-only model, 68 acoustic features extracted from in-game communication (see Table 1 ) served as independent variables, and the match outcome (Result_game: 0 = loss, 1 = win) was used as the dependent variable. For the combined model, three contextual variables — team ranking, opponent ranking, and ranking difference — were included alongside the significant acoustic predictors identified in the first stage. Initially, univariate logistic regressions were conducted to identify potential acoustic predictors of match outcomes. Variables with p < 0.20 were retained for multivariate modeling. Subsequently, multivariate logistic regression models were constructed following the purposeful selection of variables approach (Bursac, 2008), which integrates theoretical relevance, confounding assessment (based on literature and changes > 10% in β coefficients), and statistical criteria. A stepwise backward elimination procedure was applied until only significant predictors (p < 0.05) and theoretically relevant variables remained in the final model. Regression diagnostics were systematically verified following Osborne and Waters (2002), including checks for linearity of the logit, absence of multicollinearity (VIF < 2.0), independence of residuals, and overall model adequacy. To evaluate model fit and discriminative performance, we computed the Hosmer–Lemeshow goodness-of-fit test, Nagelkerke’s R², and the Area Under the Receiver Operating Characteristic Curve (AUC). Model accuracy, sensitivity, and specificity were also calculated to quantify predictive capacity. All analyses were performed separately for pre-match and in-match voice segments to account for contextual variability in players’ affective and communicative states. Ethical Aspects This study was approved by the local ethics committee. All participants sign the informed consent form before starting the assessments according to the Declaration of HelsinkiThis study was approved by the local ethics committee (CAAE: 72933023.5.0000.5479) and conducted in accordance with the Declaration of Helsinki and national ethical standards (World Medical Association, 1964 ). The analyzed audio data were part of routine training and competitive recordings systematically collected by a professional e-sports team for performance monitoring. Each athlete had previously authorized the recording of their in-game communications as part of their contractual and professional activities. For this research, institutional consent was obtained from the team management, formally granting access to anonymized audio archives for scientific analysis. No personally identifiable information was collected or analyzed, and all data were de-identified prior to processing to ensure confidentiality and compliance with ethical and privacy regulations. RESULTS Data characteristics The dataset comprised 83 audio segments extracted from recordings of professional Counter-Strike: Global Offensive (CS:GO) players during official competitive matches. Each recording represented a continuous segment of team communication captured either in pre-match discussions or during gameplay. All recordings passed preprocessing steps, including silence removal and signal normalization, and none were excluded as outliers, since every file met the predefined acoustic quality standards. For each segment, 68 acoustic features were extracted via the PyAudioAnalysis library, covering temporal (e.g., zero-crossing rate, short-term energy), spectral (centroid, spread, entropy, flux, roll-off), and cepstral domains (13 MFCCs and 12 chroma coefficients plus chroma deviation), together with their first-order derivatives. These features were later matched to contextual information about each match. The final dataset also included competitive history variables describing team performance. The dependent variable describing game results indicated match outcome (0 = loss, 1 = win), with 37.3% wins and 62.7% losses, reflecting realistic outcome variability in high-level tournaments. The studied team’s ranking ranged from 149 to 258 (mean = 194.23, SD = 46.59), while opponents’ rankings ranged from 16 to 300 (mean = 165.24, SD = 88.42). The difference between opponent and team rankings spanned from − 240 to + 151 (mean = − 28.99, SD = 96.24), indicating that some matches involved stronger opponents (positive values) whereas others favored the studied team (negative values). The median difference (− 12) and interquartile range (− 97 to + 33) confirm a heterogeneous sample covering both balanced and unbalanced matchups. This broad dispersion across ranking differences provided an ideal structure to assess whether vocal dynamics could predict competitive success independently of historical performance metrics. The next sections detail two regression models: the first testing voice features alone, and the second integrating contextual ranking variables. Regression analyses Model 1: Voice-only predictors The first regression model aimed to evaluate whether acoustic features extracted from players’ voices could independently predict the outcome of competitive matches. Following the univariate screening and stepwise backward elimination procedure, two features remained as significant predictors of victory: Chroma₁ and ΔMFCC₁₃. The resulting model presented satisfactory adjustment indices (Deviance = 98.2; AIC = 104; McFadden’s R² = 0.105), suggesting that the inclusion of these vocal parameters explained a meaningful portion of the variance in match results while maintaining an adequate balance between model complexity and predictive performance. Collinearity diagnostics confirmed the statistical robustness of the final model, with VIF = 1.09 and Tolerance = 0.916 for both predictors, indicating the absence of multicollinearity and the independence of their effects. Both Chroma₁ and ΔMFCC₁₃ were positively associated with the likelihood of winning (Table 2 ). In other words, matches in which players’ voices presented greater tonal stability (Chroma₁) and dynamic spectral modulation (ΔMFCC₁₃) were more likely to result in victory. Predictive performance metrics supported the practical significance of the model, which achieved an overall accuracy of 0.663, with specificity = 0.846 and sensitivity = 0.355, resulting in an AUC of 0.694. Although sensitivity remained moderate, the model demonstrated high specificity, indicating a stronger capacity to correctly classify victories than defeats — a pattern consistent with the variability inherent to real-world competitive performance. The balance between adjustment indices and classification accuracy indicates that even when considered in isolation, vocal parameters can serve as meaningful behavioral biomarkers of performance in professional e-sports. The findings suggest that greater harmonic organization and vocal variability — potentially reflecting emotional engagement, cognitive coordination, or stress modulation — are associated with enhanced team outcomes. Table 2 Multivariate model for Voice-only predictors. Voice-only predictors Variables Coefficient Standard Error p R²McF 0.105 Chroma 1 367.78 144.12 0.011 Delta MFCC 13 2791.86 1256.46 0.026 INSERT Table 2 HERE Model 2: Voice and team performance predictors The second regression model integrated the two previously identified acoustic predictors (Chroma₁ and ΔMFCC₁₃) with contextual variables describing team performance history. The inclusion of these contextual indicators aimed to determine whether voice-derived features retained predictive power when considered alongside traditional performance metrics. The combined model exhibited improved adjustment indices compared to the voice-only model (Deviance = 89.5; AIC = 97.5; McFadden’s R² = 0.184), indicating a stronger explanatory capacity and better balance between fit and parsimony. All three predictors remained statistically significant, confirming the robustness of their contributions. Collinearity diagnostics confirmed the adequacy of the model, with VIF values ranging from 1.02 to 1.13 and Tolerance values between 0.89 and 0.98, suggesting no evidence of multicollinearity among variables. Chroma₁, ΔMFCC₁₃ and ranking difference were all positively associated with the likelihood of victory (Table 3 ). As expected, better relative ranking (i.e., when the team was ranked higher than its opponents) increased the probability of winning, but importantly, both acoustic features remained significant even when controlling for this contextual advantage. The predictive performance of this combined model was superior to the previous one, achieving an overall accuracy of 0.723, specificity = 0.846, sensitivity = 0.516, and AUC = 0.787. These results indicate a more balanced classification between victories and defeats, with substantial improvement in sensitivity (from 0.355 to 0.516) and area under the ROC curve (from 0.694 to 0.787). This enhancement demonstrates that integrating contextual variables complements the predictive value of acoustic markers without diminishing their unique contribution. Taken together, these findings suggest that both communicative expressivity and historical team performance contribute to competitive outcomes. However, the persistence of vocal predictors after statistical control for ranking differences reinforces the idea that voice-based markers capture real-time behavioral and emotional dynamics that transcend static performance indicators. Vocal features such as tonal organization and spectral variability likely reflect coordination, arousal regulation, and collective engagement processes that are integral to effective team performance in high-pressure contexts. Table 3 Multivariate model for Voice and team performance predictors. Voice and team performance predictors Variables Coefficient Standard Error p R²McF 0.108 Chroma 1 361.30 150.10 0.016 Delta MFCC 13 2960.01 1349.77 0.028 Delta ranking 0.01 0.01 0.006 INSERT Table 3 HERE DISCUSSION In this study, we investigated whether acoustic features extracted from team voice communications during professional Counter-Strike: Global Offensive (CS:GO) matches could predict competitive outcomes, both independently and in combination with contextual ranking variables. Using regression modeling, we found that specific voice-derived parameters—particularly those related to energy and spectral entropy—showed consistent associations with match results, while contextual factors such as team and opponent rankings remained strong but not exclusive predictors. Together, these findings support the feasibility of using naturalistic voice data as a behavioral and psychophysiological marker of competitive dynamics in e-sports, highlighting both its potential and current methodological challenges. Vocal dynamics and competitive outcomes Our analyses revealed that vocal energy and spectral features were among the most informative predictors of match results. Higher short-term energy and greater spectral complexity were associated with favorable outcomes, suggesting that increased vocal engagement may reflect heightened arousal, coordination, and motivational states during successful matches(Li, 2017). These results align with previous findings showing that vocal amplitude and spectral richness increase with physiological activation and emotional intensity in naturalistic contexts such as sports, teamwork, and high-stress decision-making(Anikin, 2020 ). Conversely, flatter spectral profiles and reduced vocal energy tended to occur during matches lost by the team, possibly reflecting fatigue, cognitive overload, or decreased collective engagement(Tran,2020). In high-performance settings, reduced prosodic variability has been linked to diminished cognitive control and attentional focus under stress(Bogdanov, 2021). From this perspective, the observed voice patterns may represent a proxy for the underlying psychophysiological state of the team—capturing transient shifts in alertness, motivation, and coordination that precede observable performance outcomes. These findings reinforce the view that voice is not merely a communicative channel but an embodied signal of affective and cognitive state (Kamiloğlu, 2019). In e-sports, where continuous verbal exchange mediates strategic alignment, vocal parameters can reveal the dynamics of team synchrony and stress regulation. Thus, our results extend evidence from traditional sports psychology and affective computing into the realm of competitive gaming, offering a scalable method for behavioral monitoring in real-world high-pressure contexts(Xie, 2025). The role of contextual ranking and historical performance Contextual variables—particularly team and opponent rankings—remained significant predictors of match outcome, as expected. These metrics capture the historical skill disparity and provide a reference for interpreting whether acoustic variations correspond to genuine emotional-cognitive modulation or to structural performance asymmetries(Van Mersbergen, 2020 ). Importantly, when combined with voice features, the models indicated that acoustic and contextual factors contributed complementary information, suggesting that vocal expression contains situational variance not explained by ranking alone(Leongómez, 2017). This integration highlights the dual nature of competitive performance: one component grounded in objective historical indicators of skill, and another reflecting the real-time affective and cognitive processes that shape moment-to-moment team functioning(Delice, 2019). By modeling both domains simultaneously, the study provides a novel empirical bridge between affective neuroscience, communication research, and performance analytics. Voice as a multimodal marker of affective and cognitive load From a neuropsychophysiological standpoint, vocal features such as energy, spectral entropy, and MFCC-based parameters have been linked to autonomic arousal, respiratory control, and emotional expressivity(Morales- Luque, 2025). Increased vocal intensity and broader spectral distribution correspond to sympathetic activation and heightened engagement, while reduced variability may signal cognitive strain or fatigue(Tran, 2020). The observed associations between these parameters and match outcomes align with this framework, suggesting that voice-based metrics may act as peripheral correlates of performance-related arousal regulation. In this sense, the current study adds to a growing literature supporting voice as a non-invasive marker for tracking mental states in naturalistic environments(Schewski, 2025). Unlike self-reports or physiological sensors, speech requires no additional instrumentation and provides ecologically valid data streams that can be continuously monitored. Such features make voice analytics an attractive candidate for both research and applied settings in performance psychology and digital health. Integrating regression and predictive modeling The analytic strategy of combining regression and contextual modeling was designed to assess both explanatory and predictive perspectives. Regression models allowed for the identification of individual acoustic predictors, while multivariate frameworks incorporating ranking variables captured interactions between affective expression and situational challenge. The convergence of these methods strengthens confidence in the results, indicating that while contextual factors remain essential, voice metrics provide unique incremental predictive value. Nonetheless, the modest effect sizes observed caution against overinterpretation. The relationships between vocal features and outcomes are likely multifactorial, influenced by contextual stressors, communication roles, and team composition. These complexities underscore the need for larger datasets and complementary machine learning analyses to detect nonlinear and higher-order interactions(Masri, 2025). In future work, integrating deep-learning acoustic representations or multimodal fusion with physiological data (e.g., heart rate or EEG) could improve model sensitivity and interpretability. Practical and theoretical implications From a practical standpoint, the identification of voice-based markers of competitive success offers promising applications in performance monitoring, training feedback, and team communication analysis. Automated voice analytics could provide real-time indicators of stress, focus, or cohesion, supporting adaptive coaching and psychological interventions(Diptimoni, 2025). From a theoretical perspective, these findings contribute to a complex systems view of performance, in which communication, physiology, and environment dynamically interact to shape outcomes(Pietromonaco,2017). Furthermore, by demonstrating that naturally occurring voice data can reflect psychophysiological adaptation in a real-world competitive setting, the present study aligns with emerging frameworks in affective and social neuroscience that emphasize ecological validity and multimodal assessment(Parsons, 2015 ). Limitations and future directions Several limitations should be acknowledged. First, the dataset, although ecologically rich, represents a relatively small number of matches and players, which constrains generalizability. Second, acoustic variability may be influenced by contextual factors not captured here, such as microphone distance, noise level, or linguistic content. Third, the study focused exclusively on global match outcomes, and future research could analyze intra-match temporal dynamics to identify critical moments of stress and coordination breakdown. Finally, while our preprocessing and feature extraction followed standardized protocols, more advanced representations—such as deep embeddings or temporal convolutional descriptors—may better capture the nonlinear structure of emotional speech in competitive contexts. Future studies should therefore employ larger, multicentric datasets and multimodal frameworks integrating voice, physiology, and behavioral performance metrics. The use of explainable machine learning methods and longitudinal data will be crucial to validate the robustness, interpretability, and ecological applicability of voice-based models of human performance. Conclusions In summary, this study provides new evidence that vocal acoustics can serve as a sensitive marker of emotional and cognitive dynamics in professional e-sports. Acoustic features related to vocal energy and spectral richness were associated with match outcomes, even after accounting for contextual ranking differences. These findings suggest that subtle modulations in voice reflect the interplay between arousal, engagement, and team coordination during competition. While ranking metrics remain powerful predictors of performance, the inclusion of vocal features enhances the explanatory framework, bridging affective neuroscience, communication dynamics, and data-driven performance analytics. Taken together, the results support a multimodal perspective on competitive behavior, where communication patterns, affective states, and contextual factors jointly shape outcomes. Replication in larger and longitudinal cohorts, integrating voice, physiology, and cognitive-behavioral data, will be essential to establish whether vocal dynamics can evolve into reliable biomarkers of performance and well-being in high-stakes environments. Declarations CONFLICT OF INTEREST VHO, FOA, TZSO, DACV, LMM and RRU disclose their roles as partners and researchers at Infinity Doctors, a Digital Healthcare Marketplace company. The remaining authors declare that this research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. CONSENT FOR PUBLICATION Not applicable. Competing Interests VHO, FOA, TZSO, DACV, LMM and RRU disclose their roles as partners and researchers at Infinity Doctors, a Digital Healthcare Marketplace company. The remaining authors declare that this research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. FUNDING This study was funded by Infinity Doctors Inc., a Digital Healthcare Marketplace company. The company provided salary, and the hardware used in the research. Author Contribution RIMS, GK, FOA and RRU contributed to the study concept and design. All authors participated in writing the manuscript, were involved in the analysis and interpretation of the results and approved the final version of the manuscript. Acknowledgement The authors would like to express their gratitude to Fernando Janson for his support with the English language review and to the professional e-sports team that granted access to their in-game communication data. Their collaboration made it possible to investigate the relationship between vocal dynamics and competitive performance in an authentic, real-world environment. Data Availability The datasets and code generated and analyzed during the current study are intellectual property of Infinity Doctors Inc. and cannot be made publicly available. However, anonymized acoustic-feature datasets and the machine learning model code may be shared with qualified researchers upon reasonable request to the corresponding author, subject to company approval and compliance with ethical and privacy regulations. Raw audio files cannot be shared publicly as they are inherently identifiable, but controlled access can be granted under the same conditions. STATEMENT OF ETHICS This study was submitted to and approved by the Research Ethics Committee of the Santa Casa de São Paulo School of Medical Sciences (Comitê de Ética em Pesquisa da Faculdade de Ciências Médicas da Santa Casa de São Paulo) (CAAE: 72933023.5.0000.5479). All participants were over 18 years old and provided written informed consent prior to participation. No minors were included in the study. To ensure privacy and confidentiality, all voice recordings were de-identified immediately after collection by removing any personal identifiers from the associated metadata and replacing participant names with randomly generated alphanumeric codes. Audio files were stored on secure, access-controlled institutional servers with encryption both in transit and at rest. Access to the raw audio data was restricted to the core research team and granted only for purposes directly related to the study, in accordance with the approved research protocol. Processed feature-extraction datasets contained no information that could be linked back to individual participants. Data sharing for reproducibility purposes will be performed only in anonymized form, in compliance with applicable regulations and ethical guidelines. All procedures adhered to the principles of the Declaration of Helsinki. References Alvear RM, Barón-López FJ, Alguacil MD, Dawid-Milner MS. Interactions between voice fundamental frequency and cardiovascular parameters: Preliminary results and physiological mechanisms. Logopedics Phoniatrics Vocology. 2013;38(2):52–8. ttps://doi.org/10.3109/14015439.2012.696140. Anikin A. A ligação entre a saliência auditiva e a intensidade da emoção. Cogn Emot. 2020;34:1246–59. ttps://doi.org/10.1080/02699931.2020.1736992. Behnke M, Krzyżaniak W, Nowak J, Kupiński S, Chwiłkowska P, Jęśko Białek S, Kłoskowski M, Maciejewski P, Szymański K, Lakens D, Petrova K, Jamieson JP, Gross JJ. The competitive esports physiological, affective, and video dataset. Sci Data. 2025;12(1):56. ttps://doi.org/10.1038/s41597-024-04364-z. Berga D, Pereda-Baños A, Nandi A, Febrer-Coll E, Reverte M, Russo L. (2023). Measuring arousal and stress physiology in esports: A League of Legends case study. TechRxiv. ttps://doi.org/10.36227/techrxiv.22140683 Bogdanov M, Nitschke J, LoParco S, Bartz J, Otto AR. Acute psychosocial stress increases cognitive-effort avoidance. Psychol Sci. 2021;32:1463–75. ttps://doi.org/10.1177/09567976211005465. Boyer S, Paubel PV, Ruiz R, El Yagoubi R, Daurat A. Human voice as a measure of mental load level. J Speech Lang Hear Res. 2018;61(11):2722–34. ttps://doi.org/10.1044/2018_JSLHR-S-18-0066. Briganti G, Lechien JR. Speech and voice quality as digital biomarkers in depression: A systematic review. J Voice. 2025. ttps://doi.org/10.1016/j.jvoice.2025.05.002. Bursac Z, Gauss CH, Williams DK, Hosmer DW. Purposeful selection of variables in logistic regression. Source Code Biol Med. 2008;3(1):17. ttps://doi.org/10.1186/1751-0473-3-17. Delice F, Rousseau V, Feitosa J. Advancing teams research: What, when, and how to measure team dynamics over time. Front Psychol. 2019;10:1324. ttps://doi.org/10.3389/fpsyg.2019.01324. Diptimoni N, Gypsy N, Uzzal S, Jyoti B. A machine learning-based approach for stress detection in sports students using vocal analysis. Int J Environ Sci. 2025. ttps://doi.org/10.64252/qj27qj19. Ersin A, Tezeren H, Asal B, Atabey A, Diri A, Gonen İ. The relationship between reaction time and gaming time in e-sports players. Kinesiology. 2022;54:36–42. ttps://doi.org/10.26582/k.54.1.4. Gao X, Ma K, Yang H, Wang K, Fu B, Zhu Y, She X, Cui B. A rapid, non-invasive method for fatigue detection based on voice information. Front Cell Dev Biology. 2022;10:994001. ttps://doi.org/10.3389/fcell.2022.994001. Giannakopoulos T. pyAudioAnalysis: An Open-Source Python Library for Audio Signal Analysis. PLoS ONE. 2015;10(12):e0144610. ttps://doi.org/10.1371/journal.pone.0144610. Goudbeek M, Scherer KR. Beyond arousal: Valence and potency/control cues in the vocal expression of emotion. J Acoust Soc Am. 2010;128(3):1322–36. ttps://doi.org/10.1121/1.3466853. Kamiloğlu R, Fischer AH, Sauter DA. Good vibrations: A review of vocal expressions of positive emotions. Psychon Bull Rev. 2019;27:237–65. ttps://doi.org/10.3758/s13423-019-01701-x. Leongómez J, Mileva V, Little AC, Roberts SC. Perceived differences in social status between speaker and listener affect the speaker's vocal characteristics. PLoS ONE. 2017;12:e0179407. ttps://doi.org/10.1371/journal.pone.0179407. Li A, Liao H, Tangirala S, Firth B. The content of the message matters: The differential effects of promotive and prohibitive team voice on team productivity and safety performance gains. J Appl Psychol. 2017;102(9):1259–70. ttps://doi.org/10.1037/apl0000215. Masri D, Yousef A, Turkistani L, Tadmori T, Barkat E, Kabbaj N. (2025). Classifying speech disorders using voice signals and machine learning. In Proceedings of the 22nd International Learning and Technology Conference (L&T) (pp. 349–353). IEEE. ttps://doi.org/10.1109/lt64002.2025.10940483 Morales-Luque C, Carrillo-Franco L, López-González MV, González-García M, Dawid-Milner MS. Mapping the neurophysiological link between voice and autonomic function: A scoping review. Biology. 2025;14:1382. ttps://doi.org/10.3390/biology14101382. Osbourne JW, Waters E. Four assumptions of multiple regression that researchers should always test. Volume 8. Practical Assessment, Research & Evaluation; 2002. 2. Parsons TD. Virtual reality for enhanced ecological validity and experimental control in the clinical, affective, and social neurosciences. Front Hum Neurosci. 2015;9:660. ttps://doi.org/10.3389/fnhum.2015.00660. Pietromonaco PR, Collins NL. Interpersonal mechanisms linking close relationships to health. Am Psychol. 2017;72(6):531–42. ttps://doi.org/10.1037/amp0000129. Saygı T, Odabaş İ. Comparing the accuracy and reaction times of esports players and sport sciences students. Eurasian Res Sport Sci. 2023;8(2):80–94. ttps://doi.org/10.29228/ERISS.32. Schewski L, Doss MM, Beldi G, Keller S. Measuring negative emotions and stress through acoustic correlates in speech: A systematic review. PLoS ONE. 2025;20(7):e0328833. ttps://doi.org/10.1371/journal.pone.0328833. Sondhi S, Khan M, Vijay R, Salhan AK, Chouhan S. Acoustic analysis of speech under stress. Int J Bioinform Res Appl. 2015;11(5):417–32. ttps://doi.org/10.1504/ijbra.2015.071942. Tran Y, Craig A, Craig R, Chai R, Nguyen H. The influence of mental fatigue on brain activity: Evidence from a systematic review with meta-analyses. Psychophysiology. 2020;e13554. ttps://doi.org/10.1111/psyp.13554. Van Mersbergen M, Payne A. Cognitive, emotional, and social influences on voice production elicited by three different Stroop tasks. Folia Phoniatr et Logopaedica. 2020;73:326–34. ttps://doi.org/10.1159/000508572. World Medical Association. (1964). Human experimentation: Code of ethics of W.M.A. BMJ, 2(5402), 177. ttps://doi.org/10.1136/bmj.2.5402.177 Xie N, Zhang X, Lu C. An exploratory framework for EEG-based monitoring of motivation and performance in athletic-like scenarios. Sci Rep. 2025;15:26156. ttps://doi.org/10.1038/s415. Additional Declarations Competing interest reported. VHO, FOA, TZSO, DACV, LMM and RRU disclose their roles as partners and researchers at Infinity Doctors, a Digital Healthcare Marketplace company. The remaining authors declare that this research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 06 Apr, 2026 Editor assigned by journal 31 Mar, 2026 Editor invited by journal 12 Mar, 2026 Submission checks completed at journal 12 Mar, 2026 First submitted to journal 11 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9032928","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":619850795,"identity":"a8fbaa9d-2c94-4229-8300-87ac36ebe9b6","order_by":0,"name":"Raphael Santos","email":"","orcid":"","institution":"Faculdade de Ciências Médicas da Santa Casa de São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Raphael","middleName":"","lastName":"Santos","suffix":""},{"id":619850803,"identity":"a62b7b3c-8c7f-47df-b584-39a30371fa38","order_by":1,"name":"Gabriel Kadri","email":"","orcid":"","institution":"Faculdade de Ciências Médicas da Santa Casa de São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Gabriel","middleName":"","lastName":"Kadri","suffix":""},{"id":619850804,"identity":"e8ef56cf-ae42-4d75-be97-809904ac244f","order_by":2,"name":"Felipe Aguiar","email":"","orcid":"","institution":"Faculdade de Ciências Médicas da Santa Casa de São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Felipe","middleName":"","lastName":"Aguiar","suffix":""},{"id":619850808,"identity":"380fa2f8-2c47-408b-92a7-0c7afb188e21","order_by":3,"name":"Victor Otani","email":"","orcid":"","institution":"Faculdade de Ciências Médicas da Santa Casa de São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Victor","middleName":"","lastName":"Otani","suffix":""},{"id":619850809,"identity":"642cd30d-1fbb-4c33-97c7-318d3044709d","order_by":4,"name":"Ricardo Uchida","email":"","orcid":"","institution":"Faculdade de Ciências Médicas da Santa Casa de São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Ricardo","middleName":"","lastName":"Uchida","suffix":""},{"id":619850815,"identity":"274bf99c-58b3-425c-baae-19d5fa55b12f","order_by":5,"name":"Lucas Marques","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA3UlEQVRIiWNgGAWjYBACPjiLGYg/ADEbOwEtbMhaGGeARJiJ1gLSxQOzDq8WieSHH3/mHM43b+c9/Nnm1zZ5PmYGxg8fc/BpSTOW5t122HLOYb406dy+24ZtzAzMkjO34dOSwyDNuO2wgQQzjxlzbs9tRqAWNmZe/FqYf/6EaDH+bNlz254YLWwSvBAtBtIMP24nEtbC88zMmndbOthhkr0Nt5PbmBmb8fqFnz358c2f26wNJPjPGH/48ee27fz25oMfPuLRggoY28BkA7HqQeAPKYpHwSgYBaNgpAAAVsxDa5JlqMQAAAAASUVORK5CYII=","orcid":"","institution":"Faculdade de Ciências Médicas da Santa Casa de São Paulo","correspondingAuthor":true,"prefix":"","firstName":"Lucas","middleName":"","lastName":"Marques","suffix":""}],"badges":[],"createdAt":"2026-03-04 17:53:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9032928/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9032928/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106993562,"identity":"5bfe01a9-43b0-486b-b245-683537eaaa39","added_by":"auto","created_at":"2026-04-15 14:37:41","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":855584,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9032928/v1/3cb1eefe-7d80-4606-8246-44666e391c46.pdf"}],"financialInterests":"Competing interest reported. VHO, FOA, TZSO, DACV, LMM and RRU disclose their roles as partners and researchers at Infinity Doctors, a Digital Healthcare Marketplace company. The remaining authors declare that this research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.","formattedTitle":"\u003cp\u003eVocal Dynamics as Predictors of Competitive Success in E-Sports: Evidence from Professional Counter-Strike Matches\u003c/p\u003e","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eCompetitive gaming (e-sports) has emerged as a complex psychophysiological environment in which cognitive control, emotional regulation, and team coordination are constantly challenged under high-pressure conditions. In games such as Counter-Strike: Global Offensive (CS:GO), players are required to process rapidly changing sensory information, maintain attentional focus, and make split-second decisions while engaging in continuous verbal communication with teammates. This dynamic scenario provides a unique opportunity to study how neurophysiological and behavioral processes\u0026mdash;traditionally examined in laboratory contexts\u0026mdash;manifest in naturalistic, high-stakes performance environments.\u003c/p\u003e \u003cp\u003eAmong the multiple channels of behavioral expression available in such settings, the human voice is particularly informative. Vocal output reflects moment-to-moment variations in arousal, cognitive load, and affective states, providing a continuous and non-invasive window into internal processes. Previous studies have shown that acoustic features such as fundamental frequency, intensity, spectral entropy, and formant structure can index autonomic activation, stress reactivity, and emotional valence across a range of contexts(Sondhi, 2015 ; Schewski 2025). In competitive or stressful conditions, modulations in these parameters are often observed, reflecting physiological mobilization and emotional tension(Alvear, 2013). Thus, voice analysis represents a promising avenue for capturing fine-grained affective and cognitive dynamics during real-world performance.\u003c/p\u003e \u003cp\u003eWhile voice analytics has been increasingly explored in health domains\u0026mdash;such as mental health monitoring(Briganti, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2025\u003c/span\u003e), fatigue detection(Gao, 2022), and cognitive workload assessment(Boyer, 2018)\u0026mdash;its application in competitive gaming and performance prediction remains scarce. Most existing research on e-sports has focused on overt behavioral metrics (e.g., reaction time, accuracy, strategy patterns)(Saygı, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Ersin, 2022) or physiological indices such as heart rate variability and electrodermal activity(Berga, 2023; Behnke, 2025). Given that communication plays a central role in coordination and decision-making during gameplay, understanding its acoustic underpinnings could provide valuable insights into collective performance and stress regulation under pressure.\u003c/p\u003e \u003cp\u003eMoreover, the competitive context of professional e-sports allows for quantitative contextualization through objective measures such as global team ranking and opponent strength. These metrics make it possible to link real-world performance indicators with psychophysiological proxies, establishing a bridge between affective neuroscience, behavioral data science, and sports analytics. By combining vocal acoustics with contextual ranking variables, one can test whether subtle variations in vocal tone, energy, or spectral balance carry predictive value for match outcomes beyond historical performance alone.\u003c/p\u003e \u003cp\u003eDespite these opportunities, no studies to date have systematically modeled the relationship between vocal features and competitive success in professional gaming. The current literature remains largely descriptive or focused on laboratory-based emotional speech paradigms (Berga, 2023; Behnke, 2025), limiting its ecological validity. Addressing this gap could not only advance the understanding of stress and communication dynamics in e-sports but also inform broader models of performance monitoring applicable to high-demand occupations such as aviation, military operations, and emergency response, where real-time vocal analysis may serve as an unobtrusive marker of cognitive and affective state.\u003c/p\u003e \u003cp\u003eThe present study was designed to explore these associations using a dataset of professional CS:GO matches, in which the same team\u0026rsquo;s in-game voice communications were recorded across multiple official and scrimmage matches. For each audio segment, 68 acoustic features encompassing temporal, spectral, and cepstral domains were extracted. These features were analyzed in relation to both contextual competitive variables\u0026mdash;team and opponent rankings\u0026mdash;and objective match outcomes. Our analytic strategy involved multiple regression models to test whether specific acoustic signatures were associated with competitive success, both independently and in combination with ranking indicators.\u003c/p\u003e \u003cp\u003eWe hypothesized that increased vocal energy and spectral complexity would reflect heightened arousal and engagement (Goudbeek, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2010\u003c/span\u003e), potentially predicting successful outcomes, whereas flatter spectral profiles might be linked to fatigue or cognitive overload during disadvantageous situations(Tran, 2020). Furthermore, we expected that integrating contextual ranking variables would enhance predictive accuracy, capturing interactions between emotional expression, physiological activation, and situational challenge within the naturalistic dynamics of professional competition.\u003c/p\u003e"},{"header":"METHODS AND ANALYSIS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003eSampling methods, participants, and study design\u003c/h2\u003e\n \u003cp\u003eThe present study is part of the Performance Prediction in E-Sports Players through Acoustic Analysis of Voice, Speech, and Facial Expression research initiative, conducted at the Department of Mental Health, Santa Casa de S\u0026atilde;o Paulo School of Medical Sciences, Brazil. This project aims to identify behavioral biomarkers associated with performance and emotional states in competitive e-sports environments.\u003c/p\u003e\n \u003cp\u003eFor the current analysis, we focused exclusively on voice recordings collected from professional Counter-Strike: Global Offensive (CS:GO) players during official matches. The audio data was provided by a professional team that routinely records in-game communication for strategic review and performance evaluation.\u003c/p\u003e\n \u003cp\u003eThe dataset comprises recordings from five professional e-sports players belonging to the same team, which competes in ranked professional leagues. Inclusion criteria required that participants be at least 18 years old, have an active competitive ranking, and provide written consent for the use of their voice recordings for research purposes. Players reporting neurological, psychiatric, or speech/hearing impairments were excluded.\u003c/p\u003e\n \u003cp\u003eThe study followed a cross-sectional retrospective design, analyzing previously collected voice data from multiple official matches per team (83 matches), ensuring adequate data variability and statistical power for multivariate modeling. All data were de-identified prior to analysis and processed in accordance with institutional ethical approval (CAAE: 72933023.5.0000.5479).\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eDemographic and Contextual Variables\u003c/h3\u003e\n\u003cp\u003eBasic contextual information was collected for each match, including the team\u0026rsquo;s international ranking, opponent ranking, and the difference between team rankings. These variables were extracted from publicly available e-sports databases and official competition records. Match outcome (0\u0026thinsp;=\u0026thinsp;loss, 1\u0026thinsp;=\u0026thinsp;win) was used as the dependent variable, while acoustic and contextual features served as independent predictors.\u003c/p\u003e\n\u003ch3\u003eVoice Data Acquisition and Preprocessing\u003c/h3\u003e\n\u003cp\u003eVoice data were extracted from original team communication recordings obtained during official Counter-Strike: Global Offensive (CS:GO) professional matches. Each audio file contained natural, spontaneous in-game speech used for coordination and strategic communication among team members. This dataset therefore represents a highly ecological measure of vocal expression and interpersonal synchrony during competitive performance.\u003c/p\u003e\n\u003cp\u003eAll raw files were acquired in .mp4 format and converted to 16-bit, 44.1 kHz mono .wav files using the PyDub library (v0.25.1). After conversion, only the first 60 seconds of each recording were retained for analysis, which could include the warm-up period, the pre-game phase, and/or the initial moments of the match.\u003c/p\u003e\n\u003cp\u003eNon-speech segments and background noise were automatically removed through amplitude-based silence trimming, applying a \u0026minus;40 dB threshold and a minimum silence duration of 100 ms to ensure that only active speech periods were analyzed.\u003c/p\u003e\n\u003cp\u003eSubsequently, acoustic features were extracted using the PyAudioAnalysis library (Giannakopoulos, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2015\u003c/span\u003e) in Python (v3.12). A total of 34 baseline acoustic descriptors were computed for each recording (see Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), and first-order derivatives (\u0026Delta;) were calculated for every feature, resulting in a comprehensive set of 68 parameters per sample. These descriptors encompass time-domain, spectral-domain, cepstral-domain, and tonal representations of the voice, thereby providing a multidimensional characterization of vocal dynamics.\u003c/p\u003e\n\u003cp\u003eFor each match, acoustic features were averaged across frames, producing a single representative vector per observation. The resulting dataset included both acoustic features and contextual variables (match outcome, team ranking, opponent ranking, and ranking difference). All features were z-score normalized prior to regression analyses, ensuring comparability across participants and matches. Analyses were conducted to capture differences in vocal behavior associated with active performance.\u003c/p\u003e\n\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eAcoustic features extracted by the PyAudioAnalysis library. When deriving each of the characteristics to the first order, a total of 68 features are obtained.\u003c/p\u003e\n \u003cdiv class=\"Credit\"\u003e\n \u003cp\u003eAdapted from Giannakopoulos T. (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2015\u003c/span\u003e).\u003c/p\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eIndex\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eName\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eDescription\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eZero Crossing Rate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe rate of sign-changes of the signal during the duration of a particular frame.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eEnergy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe sum of squares of the signal values, normalised by the length.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eEntropy of Energy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe entropy of sub-frames normalized energies. It can be interpreted as a measure of abrupt changes.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eSpectral Centroid\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe centre of gravity of the spectrum.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eSpectral Spread\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe second of gravity of the spectrum.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eSpectral Entropy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eEntropy of the normalized spectral energies for a set of sub-frames\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eSpectral Flux\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe squared difference between the normalised magnitudes of the spectra of the two successive frames.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eSpectral Rolloff\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe frequency below which 90% of the magnitude distribution of the spectrum is concentrated.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e9\u0026ndash;21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eMFCCs\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eMel Frequency Cepstral Coefficients form a cepstral representation where the frequency bands are not linear but distributed according to the mel-scale.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e22\u0026ndash;33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eChroma Vector\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eA 12-element representation of the spectral energy where the bins represent the 12 equal-tempered pitch classes of western-type music (semitone spacing).\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eChroma Deviation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThe standard deviation of the 12 chroma coefficients.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eINSERT\u003c/strong\u003e Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e \u003cstrong\u003eHERE\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n \u003ch2\u003eStatistical analysis\u003c/h2\u003e\n \u003cp\u003eThe primary analytic strategy was based on classical logistic regression models, aiming to test associations between acoustic features and competitive outcomes (win vs. loss). Given the high dimensionality and potential multicollinearity of the acoustic dataset, analyses were conducted in two sequential stages: (1) voice-only models to isolate the predictive value of vocal biomarkers, and (2) combined models integrating both acoustic and contextual performance variables (team ranking, opponent ranking, and ranking difference).\u003c/p\u003e\n \u003cp\u003eThis two-step approach allowed us to contrast the intrinsic predictive capacity of behavioral voice parameters with the incremental explanatory power of historical team performance indicators.\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eRegression analyses\u003c/h3\u003e\n\u003cp\u003eAll regression analyses were performed using Python (v3.12) and Jamovi\u0026reg; 2.5.\u003c/p\u003e\n\u003cp\u003eFor the voice-only model, 68 acoustic features extracted from in-game communication (see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) served as independent variables, and the match outcome (Result_game: 0\u0026thinsp;=\u0026thinsp;loss, 1\u0026thinsp;=\u0026thinsp;win) was used as the dependent variable. For the combined model, three contextual variables \u0026mdash; team ranking, opponent ranking, and ranking difference \u0026mdash; were included alongside the significant acoustic predictors identified in the first stage.\u003c/p\u003e\n\u003cp\u003eInitially, univariate logistic regressions were conducted to identify potential acoustic predictors of match outcomes. Variables with p\u0026thinsp;\u0026lt;\u0026thinsp;0.20 were retained for multivariate modeling. Subsequently, multivariate logistic regression models were constructed following the purposeful selection of variables approach (Bursac, 2008), which integrates theoretical relevance, confounding assessment (based on literature and changes\u0026thinsp;\u0026gt;\u0026thinsp;10% in \u0026beta; coefficients), and statistical criteria.\u003c/p\u003e\n\u003cp\u003eA stepwise backward elimination procedure was applied until only significant predictors (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05) and theoretically relevant variables remained in the final model. Regression diagnostics were systematically verified following Osborne and Waters (2002), including checks for linearity of the logit, absence of multicollinearity (VIF\u0026thinsp;\u0026lt;\u0026thinsp;2.0), independence of residuals, and overall model adequacy.\u003c/p\u003e\n\u003cp\u003eTo evaluate model fit and discriminative performance, we computed the Hosmer\u0026ndash;Lemeshow goodness-of-fit test, Nagelkerke\u0026rsquo;s R\u0026sup2;, and the Area Under the Receiver Operating Characteristic Curve (AUC). Model accuracy, sensitivity, and specificity were also calculated to quantify predictive capacity.\u003c/p\u003e\n\u003cp\u003eAll analyses were performed separately for pre-match and in-match voice segments to account for contextual variability in players\u0026rsquo; affective and communicative states.\u003c/p\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n \u003ch2\u003eEthical Aspects\u003c/h2\u003e\n \u003cp\u003eThis study was approved by the local ethics committee. All participants sign the informed consent form before starting the assessments according to the Declaration of HelsinkiThis study was approved by the local ethics committee (CAAE: 72933023.5.0000.5479) and conducted in accordance with the Declaration of Helsinki and national ethical standards (World Medical Association, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e1964\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eThe analyzed audio data were part of routine training and competitive recordings systematically collected by a professional e-sports team for performance monitoring. Each athlete had previously authorized the recording of their in-game communications as part of their contractual and professional activities.\u003c/p\u003e\n \u003cp\u003eFor this research, institutional consent was obtained from the team management, formally granting access to anonymized audio archives for scientific analysis. No personally identifiable information was collected or analyzed, and all data were de-identified prior to processing to ensure confidentiality and compliance with ethical and privacy regulations.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n \u003ch2\u003eData characteristics\u003c/h2\u003e\n \u003cp\u003eThe dataset comprised 83 audio segments extracted from recordings of professional Counter-Strike: Global Offensive (CS:GO) players during official competitive matches. Each recording represented a continuous segment of team communication captured either in pre-match discussions or during gameplay. All recordings passed preprocessing steps, including silence removal and signal normalization, and none were excluded as outliers, since every file met the predefined acoustic quality standards.\u003c/p\u003e\n \u003cp\u003eFor each segment, 68 acoustic features were extracted via the PyAudioAnalysis library, covering temporal (e.g., zero-crossing rate, short-term energy), spectral (centroid, spread, entropy, flux, roll-off), and cepstral domains (13 MFCCs and 12 chroma coefficients plus chroma deviation), together with their first-order derivatives. These features were later matched to contextual information about each match.\u003c/p\u003e\n \u003cp\u003eThe final dataset also included competitive history variables describing team performance. The dependent variable describing game results indicated match outcome (0\u0026thinsp;=\u0026thinsp;loss, 1\u0026thinsp;=\u0026thinsp;win), with 37.3% wins and 62.7% losses, reflecting realistic outcome variability in high-level tournaments. The studied team\u0026rsquo;s ranking ranged from 149 to 258 (mean\u0026thinsp;=\u0026thinsp;194.23, SD\u0026thinsp;=\u0026thinsp;46.59), while opponents\u0026rsquo; rankings ranged from 16 to 300 (mean\u0026thinsp;=\u0026thinsp;165.24, SD\u0026thinsp;=\u0026thinsp;88.42). The difference between opponent and team rankings spanned from \u0026minus;\u0026thinsp;240 to +\u0026thinsp;151 (mean\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;28.99, SD\u0026thinsp;=\u0026thinsp;96.24), indicating that some matches involved stronger opponents (positive values) whereas others favored the studied team (negative values). The median difference (\u0026minus;\u0026thinsp;12) and interquartile range (\u0026minus;\u0026thinsp;97 to +\u0026thinsp;33) confirm a heterogeneous sample covering both balanced and unbalanced matchups.\u003c/p\u003e\n \u003cp\u003eThis broad dispersion across ranking differences provided an ideal structure to assess whether vocal dynamics could predict competitive success independently of historical performance metrics. The next sections detail two regression models: the first testing voice features alone, and the second integrating contextual ranking variables.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003eRegression analyses\u003c/h2\u003e\n \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\n \u003ch2\u003eModel 1: Voice-only predictors\u003c/h2\u003e\n \u003cp\u003eThe first regression model aimed to evaluate whether acoustic features extracted from players\u0026rsquo; voices could independently predict the outcome of competitive matches. Following the univariate screening and stepwise backward elimination procedure, two features remained as significant predictors of victory: Chroma₁ and \u0026Delta;MFCC₁₃.\u003c/p\u003e\n \u003cp\u003eThe resulting model presented satisfactory adjustment indices (Deviance\u0026thinsp;=\u0026thinsp;98.2; AIC\u0026thinsp;=\u0026thinsp;104; McFadden\u0026rsquo;s R\u0026sup2; = 0.105), suggesting that the inclusion of these vocal parameters explained a meaningful portion of the variance in match results while maintaining an adequate balance between model complexity and predictive performance. Collinearity diagnostics confirmed the statistical robustness of the final model, with VIF\u0026thinsp;=\u0026thinsp;1.09 and Tolerance\u0026thinsp;=\u0026thinsp;0.916 for both predictors, indicating the absence of multicollinearity and the independence of their effects.\u003c/p\u003e\n \u003cp\u003eBoth Chroma₁ and \u0026Delta;MFCC₁₃ were positively associated with the likelihood of winning (Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). In other words, matches in which players\u0026rsquo; voices presented greater tonal stability (Chroma₁) and dynamic spectral modulation (\u0026Delta;MFCC₁₃) were more likely to result in victory. Predictive performance metrics supported the practical significance of the model, which achieved an overall accuracy of 0.663, with specificity\u0026thinsp;=\u0026thinsp;0.846 and sensitivity\u0026thinsp;=\u0026thinsp;0.355, resulting in an AUC of 0.694. Although sensitivity remained moderate, the model demonstrated high specificity, indicating a stronger capacity to correctly classify victories than defeats \u0026mdash; a pattern consistent with the variability inherent to real-world competitive performance.\u003c/p\u003e\n \u003cp\u003eThe balance between adjustment indices and classification accuracy indicates that even when considered in isolation, vocal parameters can serve as meaningful behavioral biomarkers of performance in professional e-sports. The findings suggest that greater harmonic organization and vocal variability \u0026mdash; potentially reflecting emotional engagement, cognitive coordination, or stress modulation \u0026mdash; are associated with enhanced team outcomes.\u003c/p\u003e\n \u003ctable float=\"Yes\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eMultivariate model for Voice-only predictors.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\n \u003cp\u003eVoice-only predictors\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cem\u003eVariables\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e\u003cem\u003eCoefficient\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e\u003cem\u003eStandard Error\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u003cem\u003ep\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e\u003cem\u003eR\u0026sup2;McF\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e0.105\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eChroma\u003csub\u003e1\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e367.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e144.12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e0.011\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eDelta MFCC\u003csub\u003e13\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e2791.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e1256.46\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e0.026\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eINSERT\u003c/strong\u003e Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e \u003cstrong\u003eHERE\u003c/strong\u003e\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n \u003ch2\u003eModel 2: Voice and team performance predictors\u003c/h2\u003e\n \u003cp\u003eThe second regression model integrated the two previously identified acoustic predictors (Chroma₁ and \u0026Delta;MFCC₁₃) with contextual variables describing team performance history. The inclusion of these contextual indicators aimed to determine whether voice-derived features retained predictive power when considered alongside traditional performance metrics.\u003c/p\u003e\n \u003cp\u003eThe combined model exhibited improved adjustment indices compared to the voice-only model (Deviance\u0026thinsp;=\u0026thinsp;89.5; AIC\u0026thinsp;=\u0026thinsp;97.5; McFadden\u0026rsquo;s R\u0026sup2; = 0.184), indicating a stronger explanatory capacity and better balance between fit and parsimony. All three predictors remained statistically significant, confirming the robustness of their contributions. Collinearity diagnostics confirmed the adequacy of the model, with VIF values ranging from 1.02 to 1.13 and Tolerance values between 0.89 and 0.98, suggesting no evidence of multicollinearity among variables.\u003c/p\u003e\n \u003cp\u003eChroma₁, \u0026Delta;MFCC₁₃ and ranking difference were all positively associated with the likelihood of victory (Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). As expected, better relative ranking (i.e., when the team was ranked higher than its opponents) increased the probability of winning, but importantly, both acoustic features remained significant even when controlling for this contextual advantage.\u003c/p\u003e\n \u003cp\u003eThe predictive performance of this combined model was superior to the previous one, achieving an overall accuracy of 0.723, specificity\u0026thinsp;=\u0026thinsp;0.846, sensitivity\u0026thinsp;=\u0026thinsp;0.516, and AUC\u0026thinsp;=\u0026thinsp;0.787. These results indicate a more balanced classification between victories and defeats, with substantial improvement in sensitivity (from 0.355 to 0.516) and area under the ROC curve (from 0.694 to 0.787). This enhancement demonstrates that integrating contextual variables complements the predictive value of acoustic markers without diminishing their unique contribution.\u003c/p\u003e\n \u003cp\u003eTaken together, these findings suggest that both communicative expressivity and historical team performance contribute to competitive outcomes. However, the persistence of vocal predictors after statistical control for ranking differences reinforces the idea that voice-based markers capture real-time behavioral and emotional dynamics that transcend static performance indicators. Vocal features such as tonal organization and spectral variability likely reflect coordination, arousal regulation, and collective engagement processes that are integral to effective team performance in high-pressure contexts.\u003c/p\u003e\n \u003ctable float=\"Yes\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eMultivariate model for Voice and team performance predictors.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e\n \u003cp\u003eVoice and team performance predictors\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cem\u003eVariables\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e\u003cem\u003eCoefficient\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e\u003cem\u003eStandard Error\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u003cem\u003ep\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e\u003cem\u003eR\u0026sup2;McF\u003c/em\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e0.108\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eChroma\u003csub\u003e1\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e361.30\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e150.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e0.016\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eDelta MFCC\u003csub\u003e13\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e2960.01\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e1349.77\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e0.028\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eDelta ranking\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.01\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e0.01\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e0.006\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eINSERT\u003c/strong\u003e Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e \u003cstrong\u003eHERE\u003c/strong\u003e\u003c/p\u003e\n\u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eIn this study, we investigated whether acoustic features extracted from team voice communications during professional Counter-Strike: Global Offensive (CS:GO) matches could predict competitive outcomes, both independently and in combination with contextual ranking variables. Using regression modeling, we found that specific voice-derived parameters\u0026mdash;particularly those related to energy and spectral entropy\u0026mdash;showed consistent associations with match results, while contextual factors such as team and opponent rankings remained strong but not exclusive predictors. Together, these findings support the feasibility of using naturalistic voice data as a behavioral and psychophysiological marker of competitive dynamics in e-sports, highlighting both its potential and current methodological challenges.\u003c/p\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eVocal dynamics and competitive outcomes\u003c/h2\u003e \u003cp\u003eOur analyses revealed that vocal energy and spectral features were among the most informative predictors of match results. Higher short-term energy and greater spectral complexity were associated with favorable outcomes, suggesting that increased vocal engagement may reflect heightened arousal, coordination, and motivational states during successful matches(Li, 2017). These results align with previous findings showing that vocal amplitude and spectral richness increase with physiological activation and emotional intensity in naturalistic contexts such as sports, teamwork, and high-stress decision-making(Anikin, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eConversely, flatter spectral profiles and reduced vocal energy tended to occur during matches lost by the team, possibly reflecting fatigue, cognitive overload, or decreased collective engagement(Tran,2020). In high-performance settings, reduced prosodic variability has been linked to diminished cognitive control and attentional focus under stress(Bogdanov, 2021). From this perspective, the observed voice patterns may represent a proxy for the underlying psychophysiological state of the team\u0026mdash;capturing transient shifts in alertness, motivation, and coordination that precede observable performance outcomes.\u003c/p\u003e \u003cp\u003eThese findings reinforce the view that voice is not merely a communicative channel but an embodied signal of affective and cognitive state (Kamiloğlu, 2019). In e-sports, where continuous verbal exchange mediates strategic alignment, vocal parameters can reveal the dynamics of team synchrony and stress regulation. Thus, our results extend evidence from traditional sports psychology and affective computing into the realm of competitive gaming, offering a scalable method for behavioral monitoring in real-world high-pressure contexts(Xie, 2025).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eThe role of contextual ranking and historical performance\u003c/h2\u003e \u003cp\u003eContextual variables\u0026mdash;particularly team and opponent rankings\u0026mdash;remained significant predictors of match outcome, as expected. These metrics capture the historical skill disparity and provide a reference for interpreting whether acoustic variations correspond to genuine emotional-cognitive modulation or to structural performance asymmetries(Van Mersbergen, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Importantly, when combined with voice features, the models indicated that acoustic and contextual factors contributed complementary information, suggesting that vocal expression contains situational variance not explained by ranking alone(Leong\u0026oacute;mez, 2017).\u003c/p\u003e \u003cp\u003eThis integration highlights the dual nature of competitive performance: one component grounded in objective historical indicators of skill, and another reflecting the real-time affective and cognitive processes that shape moment-to-moment team functioning(Delice, 2019). By modeling both domains simultaneously, the study provides a novel empirical bridge between affective neuroscience, communication research, and performance analytics.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eVoice as a multimodal marker of affective and cognitive load\u003c/h2\u003e \u003cp\u003eFrom a neuropsychophysiological standpoint, vocal features such as energy, spectral entropy, and MFCC-based parameters have been linked to autonomic arousal, respiratory control, and emotional expressivity(Morales- Luque, 2025). Increased vocal intensity and broader spectral distribution correspond to sympathetic activation and heightened engagement, while reduced variability may signal cognitive strain or fatigue(Tran, 2020). The observed associations between these parameters and match outcomes align with this framework, suggesting that voice-based metrics may act as peripheral correlates of performance-related arousal regulation.\u003c/p\u003e \u003cp\u003eIn this sense, the current study adds to a growing literature supporting voice as a non-invasive marker for tracking mental states in naturalistic environments(Schewski, 2025). Unlike self-reports or physiological sensors, speech requires no additional instrumentation and provides ecologically valid data streams that can be continuously monitored. Such features make voice analytics an attractive candidate for both research and applied settings in performance psychology and digital health.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eIntegrating regression and predictive modeling\u003c/h2\u003e \u003cp\u003eThe analytic strategy of combining regression and contextual modeling was designed to assess both explanatory and predictive perspectives. Regression models allowed for the identification of individual acoustic predictors, while multivariate frameworks incorporating ranking variables captured interactions between affective expression and situational challenge. The convergence of these methods strengthens confidence in the results, indicating that while contextual factors remain essential, voice metrics provide unique incremental predictive value.\u003c/p\u003e \u003cp\u003eNonetheless, the modest effect sizes observed caution against overinterpretation. The relationships between vocal features and outcomes are likely multifactorial, influenced by contextual stressors, communication roles, and team composition. These complexities underscore the need for larger datasets and complementary machine learning analyses to detect nonlinear and higher-order interactions(Masri, 2025). In future work, integrating deep-learning acoustic representations or multimodal fusion with physiological data (e.g., heart rate or EEG) could improve model sensitivity and interpretability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003ePractical and theoretical implications\u003c/h2\u003e \u003cp\u003eFrom a practical standpoint, the identification of voice-based markers of competitive success offers promising applications in performance monitoring, training feedback, and team communication analysis. Automated voice analytics could provide real-time indicators of stress, focus, or cohesion, supporting adaptive coaching and psychological interventions(Diptimoni, 2025). From a theoretical perspective, these findings contribute to a complex systems view of performance, in which communication, physiology, and environment dynamically interact to shape outcomes(Pietromonaco,2017).\u003c/p\u003e \u003cp\u003eFurthermore, by demonstrating that naturally occurring voice data can reflect psychophysiological adaptation in a real-world competitive setting, the present study aligns with emerging frameworks in affective and social neuroscience that emphasize ecological validity and multimodal assessment(Parsons, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2015\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eLimitations and future directions\u003c/h2\u003e \u003cp\u003eSeveral limitations should be acknowledged. First, the dataset, although ecologically rich, represents a relatively small number of matches and players, which constrains generalizability. Second, acoustic variability may be influenced by contextual factors not captured here, such as microphone distance, noise level, or linguistic content. Third, the study focused exclusively on global match outcomes, and future research could analyze intra-match temporal dynamics to identify critical moments of stress and coordination breakdown. Finally, while our preprocessing and feature extraction followed standardized protocols, more advanced representations\u0026mdash;such as deep embeddings or temporal convolutional descriptors\u0026mdash;may better capture the nonlinear structure of emotional speech in competitive contexts.\u003c/p\u003e \u003cp\u003eFuture studies should therefore employ larger, multicentric datasets and multimodal frameworks integrating voice, physiology, and behavioral performance metrics. The use of explainable machine learning methods and longitudinal data will be crucial to validate the robustness, interpretability, and ecological applicability of voice-based models of human performance.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn summary, this study provides new evidence that vocal acoustics can serve as a sensitive marker of emotional and cognitive dynamics in professional e-sports. Acoustic features related to vocal energy and spectral richness were associated with match outcomes, even after accounting for contextual ranking differences. These findings suggest that subtle modulations in voice reflect the interplay between arousal, engagement, and team coordination during competition. While ranking metrics remain powerful predictors of performance, the inclusion of vocal features enhances the explanatory framework, bridging affective neuroscience, communication dynamics, and data-driven performance analytics.\u003c/p\u003e\n\u003cp\u003eTaken together, the results support a multimodal perspective on competitive behavior, where communication patterns, affective states, and contextual factors jointly shape outcomes. Replication in larger and longitudinal cohorts, integrating voice, physiology, and cognitive-behavioral data, will be essential to establish whether vocal dynamics can evolve into reliable biomarkers of performance and well-being in high-stakes environments.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCONFLICT OF INTEREST\u003c/h2\u003e \u003cp\u003eVHO, FOA, TZSO, DACV, LMM and RRU disclose their roles as partners and researchers at Infinity Doctors, a Digital Healthcare Marketplace company. The remaining authors declare that this research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eCONSENT FOR PUBLICATION\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003cp\u003eVHO, FOA, TZSO, DACV, LMM and RRU disclose their roles as partners and researchers at Infinity Doctors, a Digital Healthcare Marketplace company. The remaining authors declare that this research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFUNDING\u003c/h2\u003e \u003cp\u003eThis study was funded by Infinity Doctors Inc., a Digital Healthcare Marketplace company. The company provided salary, and the hardware used in the research.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eRIMS, GK, FOA and RRU contributed to the study concept and design. All authors participated in writing the manuscript, were involved in the analysis and interpretation of the results and approved the final version of the manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe authors would like to express their gratitude to Fernando Janson for his support with the English language review and to the professional e-sports team that granted access to their in-game communication data. Their collaboration made it possible to investigate the relationship between vocal dynamics and competitive performance in an authentic, real-world environment.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets and code generated and analyzed during the current study are intellectual property of Infinity Doctors Inc. and cannot be made publicly available. However, anonymized acoustic-feature datasets and the machine learning model code may be shared with qualified researchers upon reasonable request to the corresponding author, subject to company approval and compliance with ethical and privacy regulations. Raw audio files cannot be shared publicly as they are inherently identifiable, but controlled access can be granted under the same conditions.\u003c/p\u003e\n \u003ch2\u003eSTATEMENT OF ETHICS\u003c/h2\u003e\n \u003cp\u003eThis study was submitted to and approved by the Research Ethics Committee of the Santa Casa de S\u0026atilde;o Paulo School of Medical Sciences (Comit\u0026ecirc; de \u0026Eacute;tica em Pesquisa da Faculdade de Ci\u0026ecirc;ncias M\u0026eacute;dicas da Santa Casa de S\u0026atilde;o Paulo) (CAAE: 72933023.5.0000.5479). All participants were over 18 years old and provided written informed consent prior to participation. No minors were included in the study. To ensure privacy and confidentiality, all voice recordings were de-identified immediately after collection by removing any personal identifiers from the associated metadata and replacing participant names with randomly generated alphanumeric codes. Audio files were stored on secure, access-controlled institutional servers with encryption both in transit and at rest. Access to the raw audio data was restricted to the core research team and granted only for purposes directly related to the study, in accordance with the approved research protocol. Processed feature-extraction datasets contained no information that could be linked back to individual participants. Data sharing for reproducibility purposes will be performed only in anonymized form, in compliance with applicable regulations and ethical guidelines. All procedures adhered to the principles of the Declaration of Helsinki.\u003c/p\u003e\n "},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAlvear RM, Bar\u0026oacute;n-L\u0026oacute;pez FJ, Alguacil MD, Dawid-Milner MS. Interactions between voice fundamental frequency and cardiovascular parameters: Preliminary results and physiological mechanisms. Logopedics Phoniatrics Vocology. 2013;38(2):52\u0026ndash;8. ttps://doi.org/10.3109/14015439.2012.696140.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnikin A. A liga\u0026ccedil;\u0026atilde;o entre a sali\u0026ecirc;ncia auditiva e a intensidade da emo\u0026ccedil;\u0026atilde;o. Cogn Emot. 2020;34:1246\u0026ndash;59. ttps://doi.org/10.1080/02699931.2020.1736992.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBehnke M, Krzyżaniak W, Nowak J, Kupiński S, Chwiłkowska P, Jęśko Białek S, Kłoskowski M, Maciejewski P, Szymański K, Lakens D, Petrova K, Jamieson JP, Gross JJ. The competitive esports physiological, affective, and video dataset. Sci Data. 2025;12(1):56. ttps://doi.org/10.1038/s41597-024-04364-z.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBerga D, Pereda-Ba\u0026ntilde;os A, Nandi A, Febrer-Coll E, Reverte M, Russo L. (2023). Measuring arousal and stress physiology in esports: A League of Legends case study. TechRxiv. ttps://doi.org/10.36227/techrxiv.22140683\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBogdanov M, Nitschke J, LoParco S, Bartz J, Otto AR. Acute psychosocial stress increases cognitive-effort avoidance. Psychol Sci. 2021;32:1463\u0026ndash;75. ttps://doi.org/10.1177/09567976211005465.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoyer S, Paubel PV, Ruiz R, El Yagoubi R, Daurat A. Human voice as a measure of mental load level. J Speech Lang Hear Res. 2018;61(11):2722\u0026ndash;34. ttps://doi.org/10.1044/2018_JSLHR-S-18-0066.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBriganti G, Lechien JR. Speech and voice quality as digital biomarkers in depression: A systematic review. J Voice. 2025. ttps://doi.org/10.1016/j.jvoice.2025.05.002.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBursac Z, Gauss CH, Williams DK, Hosmer DW. Purposeful selection of variables in logistic regression. Source Code Biol Med. 2008;3(1):17. ttps://doi.org/10.1186/1751-0473-3-17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDelice F, Rousseau V, Feitosa J. Advancing teams research: What, when, and how to measure team dynamics over time. Front Psychol. 2019;10:1324. ttps://doi.org/10.3389/fpsyg.2019.01324.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDiptimoni N, Gypsy N, Uzzal S, Jyoti B. A machine learning-based approach for stress detection in sports students using vocal analysis. Int J Environ Sci. 2025. ttps://doi.org/10.64252/qj27qj19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eErsin A, Tezeren H, Asal B, Atabey A, Diri A, Gonen İ. The relationship between reaction time and gaming time in e-sports players. Kinesiology. 2022;54:36\u0026ndash;42. ttps://doi.org/10.26582/k.54.1.4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao X, Ma K, Yang H, Wang K, Fu B, Zhu Y, She X, Cui B. A rapid, non-invasive method for fatigue detection based on voice information. Front Cell Dev Biology. 2022;10:994001. ttps://doi.org/10.3389/fcell.2022.994001.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiannakopoulos T. pyAudioAnalysis: An Open-Source Python Library for Audio Signal Analysis. PLoS ONE. 2015;10(12):e0144610. ttps://doi.org/10.1371/journal.pone.0144610.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoudbeek M, Scherer KR. Beyond arousal: Valence and potency/control cues in the vocal expression of emotion. J Acoust Soc Am. 2010;128(3):1322\u0026ndash;36. ttps://doi.org/10.1121/1.3466853.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKamiloğlu R, Fischer AH, Sauter DA. Good vibrations: A review of vocal expressions of positive emotions. Psychon Bull Rev. 2019;27:237\u0026ndash;65. ttps://doi.org/10.3758/s13423-019-01701-x.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeong\u0026oacute;mez J, Mileva V, Little AC, Roberts SC. Perceived differences in social status between speaker and listener affect the speaker's vocal characteristics. PLoS ONE. 2017;12:e0179407. ttps://doi.org/10.1371/journal.pone.0179407.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi A, Liao H, Tangirala S, Firth B. The content of the message matters: The differential effects of promotive and prohibitive team voice on team productivity and safety performance gains. J Appl Psychol. 2017;102(9):1259\u0026ndash;70. ttps://doi.org/10.1037/apl0000215.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMasri D, Yousef A, Turkistani L, Tadmori T, Barkat E, Kabbaj N. (2025). Classifying speech disorders using voice signals and machine learning. In Proceedings of the 22nd International Learning and Technology Conference (L\u0026amp;T) (pp. 349\u0026ndash;353). IEEE. ttps://doi.org/10.1109/lt64002.2025.10940483\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMorales-Luque C, Carrillo-Franco L, L\u0026oacute;pez-Gonz\u0026aacute;lez MV, Gonz\u0026aacute;lez-Garc\u0026iacute;a M, Dawid-Milner MS. Mapping the neurophysiological link between voice and autonomic function: A scoping review. Biology. 2025;14:1382. ttps://doi.org/10.3390/biology14101382.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOsbourne JW, Waters E. Four assumptions of multiple regression that researchers should always test. Volume 8. Practical Assessment, Research \u0026amp; Evaluation; 2002. 2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eParsons TD. Virtual reality for enhanced ecological validity and experimental control in the clinical, affective, and social neurosciences. Front Hum Neurosci. 2015;9:660. ttps://doi.org/10.3389/fnhum.2015.00660.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePietromonaco PR, Collins NL. Interpersonal mechanisms linking close relationships to health. Am Psychol. 2017;72(6):531\u0026ndash;42. ttps://doi.org/10.1037/amp0000129.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaygı T, Odabaş İ. Comparing the accuracy and reaction times of esports players and sport sciences students. Eurasian Res Sport Sci. 2023;8(2):80\u0026ndash;94. ttps://doi.org/10.29228/ERISS.32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchewski L, Doss MM, Beldi G, Keller S. Measuring negative emotions and stress through acoustic correlates in speech: A systematic review. PLoS ONE. 2025;20(7):e0328833. ttps://doi.org/10.1371/journal.pone.0328833.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSondhi S, Khan M, Vijay R, Salhan AK, Chouhan S. Acoustic analysis of speech under stress. Int J Bioinform Res Appl. 2015;11(5):417\u0026ndash;32. ttps://doi.org/10.1504/ijbra.2015.071942.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTran Y, Craig A, Craig R, Chai R, Nguyen H. The influence of mental fatigue on brain activity: Evidence from a systematic review with meta-analyses. Psychophysiology. 2020;e13554. ttps://doi.org/10.1111/psyp.13554.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Mersbergen M, Payne A. Cognitive, emotional, and social influences on voice production elicited by three different Stroop tasks. Folia Phoniatr et Logopaedica. 2020;73:326\u0026ndash;34. ttps://doi.org/10.1159/000508572.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Medical Association. (1964). Human experimentation: Code of ethics of W.M.A. BMJ, 2(5402), 177. ttps://doi.org/10.1136/bmj.2.5402.177\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie N, Zhang X, Lu C. An exploratory framework for EEG-based monitoring of motivation and performance in athletic-like scenarios. Sci Rep. 2025;15:26156. ttps://doi.org/10.1038/s415.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-psychology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"psyo","sideBox":"Learn more about [BMC Psychology](http://bmcpsychology.biomedcentral.com/)","snPcode":"","submissionUrl":"","title":"BMC Psychology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Voice analysis, E-sports, Competitive performance, Acoustic biomarkers, Cognitive load, Communication dynamics","lastPublishedDoi":"10.21203/rs.3.rs-9032928/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9032928/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjective\u003c/h2\u003e \u003cp\u003eTo identify acoustic and contextual predictors of competitive success in professional Counter-Strike: Global Offensive (CS:GO) players by analyzing voice-based biomarkers of emotional and cognitive dynamics.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eNaturalistic voice recordings from official matches were processed to extract 68 temporal, spectral, and cepstral features using the PyAudioAnalysis library. Logistic regression models tested associations between these acoustic parameters and match outcomes (win/loss), first considering voice-only predictors and then combining them with contextual performance indicators (team ranking, opponent ranking, and ranking difference).\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eTwo vocal features\u0026mdash;Chroma₁ and ΔMFCC₁₃\u0026mdash;emerged as significant predictors of victory, indicating that greater tonal organization and spectral variability were associated with winning outcomes. The inclusion of contextual ranking variables improved model fit (AUC\u0026thinsp;=\u0026thinsp;0.787), yet both acoustic predictors remained significant, demonstrating that vocal expression contributes unique information beyond historical performance.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThe findings suggest that voice dynamics during competitive play reflect real-time affective and cognitive processes linked to arousal regulation, coordination, and engagement. By integrating behavioral acoustics with contextual performance data, this study advances a multimodal framework for understanding and predicting human performance in high-pressure environments such as e-sports.\u003c/p\u003e","manuscriptTitle":"Vocal Dynamics as Predictors of Competitive Success in E-Sports: Evidence from Professional Counter-Strike Matches","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-10 14:21:53","doi":"10.21203/rs.3.rs-9032928/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-04-06T05:08:25+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-31T11:03:25+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-12T10:06:18+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-12T08:11:59+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Psychology","date":"2026-03-11T18:02:39+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-psychology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"psyo","sideBox":"Learn more about [BMC Psychology](http://bmcpsychology.biomedcentral.com/)","snPcode":"","submissionUrl":"","title":"BMC Psychology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"33d6b149-f916-4563-9699-37b7c6b56d5f","owner":[],"postedDate":"April 10th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-10T14:21:53+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-10 14:21:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9032928","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9032928","identity":"rs-9032928","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0