Full text
71,730 characters
· extracted from
preprint-html
· click to expand
Beat-to-beat aortic valve opening detection from impedance cardiography using machine learning | Authorea try { document.documentElement.classList.add('js'); } catch (e) { } var _gaq = _gaq || []; _gaq.push(['_setAccount', 'G-8VDV14Y67G']); _gaq.push(['_trackPageview']); (function() { var ga = document.createElement('script'); ga.type = 'text/javascript'; ga.async = true; ga.src = ('https:' == document.location.protocol ? 'https://ssl' : 'http://www') + '.google-analytics.com/ga.js'; var s = document.getElementsByTagName('script')[0]; s.parentNode.insertBefore(ga, s); })(); Skip to main content Preprints Collections Wiley Open Research IET Open Research Ecological Society of Japan All Collections About About Authorea FAQs Contact Us Quick Search anywhere Search for preprint articles, keywords, etc. Search Search ADVANCED SEARCH SCROLL This is a preprint and has not been peer reviewed. Data may be preliminary. 17 December 2025 V1 Latest version Share on Beat-to-beat aortic valve opening detection from impedance cardiography using machine learning Authors : Luca Abel 0000-0002-5044-7113 [email protected] , Sebastian Stühler , Tobias Steigleder , Christoph Ostgathe , Nicolas Rohleder 0000-0003-2602-517X , Bjoern M. Eskofier , and Robert Richer 0000-0003-0272-5403 Authors Info & Affiliations https://doi.org/10.22541/au.176595493.35093262/v1 178 views 119 downloads Contents Abstract 1. Introduction 2. Methods 2.2 Feature extraction 2.3 Machine Learning Model Evaluation 2.4 Statistical analysis 2.5 Availability of data and code 3. Results 3.2 Feature importances 4. Discussion 5. Conclusion Author Contributions Acknowledgments Funding Ethics Statement Conflicts of Interest Data and Code Availability Statement Information & Authors Metrics & Citations View Options References Figures Tables Media Share Abstract Accurate detection of the aortic valve opening, represented by the B-point in the impedance cardiogram (ICG), is a critical step in non-invasive hemodynamic monitoring. The B-point is essential for deriving key parameters like the pre-ejection period (PEP) and left ventricular ejection time (LVET). However, accurately identifying this fiducial point is challenging due to a high intra-individual difference in the waveform, interfering factors like body movement, and a historical lack of publicly available datasets for evaluation. Capitalizing on recently published datasets, we propose a novel machine learning-based approach for beat-to-beat B-point detection that leverages domain knowledge from 12 different established ICG-based fiducial point detection algorithms. Using these in addition to the RR-interval as the feature set, we trained different machine learning pipelines on over 11,000 cardiac cycles. The best-performing pipeline, based on a Random Forest Regressor, achieved a considerable improvement in B-point detection accuracy and robustness. Compared to the best-performing traditional algorithm, our approach reduced the mean absolute error (MAE) by 54.7 % and its standard deviation by 47.4 %, resulting in a final MAE of 8.13 ± 12.12 ms. Our novel B-point detection method will be made available for the research community by integrating it in the open-source framework PEPbench . Luca Abel*1,2, Sebastian Stühler*1, Tobias Steigleder3, Christoph Ostgathe3, Nicolas Rohleder4, Bjoern M. Eskofier1,2,5,6, Robert Richer1,2 1 Department Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, Germany 2 Munich Center for Machine Learning (MCML), Munich, Germany 3 Department for Palliative Medicine, University Hospital Erlangen, Erlangen, Germany 4 Chair of Health Psychology, Department of Psychology, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, Germany 5 Chair of AI-supported Therapy Decisions, Institute for Medical Information Processing, Biometry, and Epidemiology, LMU München, Munich, Germany 6 Translational Digital Health Group, Institute of AI for Health, Helmholtz Zentrum München – German Research Center for Environmental Health, Neuherberg, Germany * These authors contributed equally to this work. Corresponding author: Luca Abel, Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Nürnberger-Str. 74, 91052 Erlangen, Germany Author contact information (missing IDs will be added in later stages): L. Abel: [email protected] , 0000-0002-5044-7113 S. Stühler: [email protected] , n/a T. Steigleder: [email protected] , n/a C. Ostgathe: [email protected] , n/a N. Rohleder: [email protected] , n/a B. M. Eskofier: [email protected] , 0000-0002-0417-0336 R. Richer: [email protected] , 0000-0003-0272-5403 Abstract Accurate detection of the aortic valve opening, represented by the B-point in the impedance cardiogram (ICG), is a critical step in non-invasive hemodynamic monitoring. The B-point is essential for deriving key parameters like the pre-ejection period (PEP) and left ventricular ejection time (LVET). However, accurately identifying this fiducial point is challenging due to a high intra-individual difference in the waveform, interfering factors like body movement, and a historical lack of publicly available datasets for evaluation. Capitalizing on recently published datasets, we propose a novel machine learning-based approach for beat-to-beat B-point detection that leverages domain knowledge from 12 different established ICG-based fiducial point detection algorithms. Using these in addition to the RR-interval as the feature set, we trained different machine learning pipelines on over 11,000 cardiac cycles. The best-performing pipeline, based on a Random Forest Regressor, achieved a considerable improvement in B-point detection accuracy and robustness. Compared to the best-performing traditional algorithm, our approach reduced the mean absolute error (MAE) by 54.7 % and its standard deviation by 47.4 %, resulting in a final MAE of 8.13 ± 12.12 ms. Our novel B-point detection method will be made available for the research community by integrating it in the open-source framework PEPbench . 1. Introduction Accurate measurement of hemodynamic parameters, such as the pre-ejection period (PEP), left ventricular ejection time (LVET), and stroke volume (SV), relies on the reliable detection of the aortic valve opening. This physiological event is marked by the B-point within the first derivative of the impedance cardiogram (ICG), conventionally denoted as\(\text{dZ}/\text{dt}\) (Lozano et al., 2007; Karpiel et al., 2022). The parameters, such as PEP and LVET, serve as crucial non-invasive indices of myocardial contractility and autonomic nervous system regulation in both clinical monitoring and psychophysiological research (Sherwood et al., 1990; Alhakak et al., 2021). Given that manual B-point detection is time-intensive and susceptible to inter-rater variability, making it impractical for clinical application (Riese et al., 2003), considerable effort has been put into the development of automatic B-point detection algorithms in recent years. However, the automatic detection of the B-point poses a significant challenge due to several factors. The morphology of the B-point can exhibit substantial inter- and intra-individual variation (Sherwood et al., 1990; DeMarzo et al., 1996; Carvalho et al., 2011; Forouzanfar et al., 2018; Pale et al., 2021). These variations can manifest as subtle inflections in the \(\text{dZ}/\text{dt}\) signal, which are inherently difficult to identify (Lozano et al., 2007). Furthermore, the ICG waveform can be distorted by interfering factors, including respiration, body movement, and underlying cardiovascular dysfunction, thereby complicating reliable automated B-point detection (Lozano et al., 2007; Sherwood et al., 1990; Nagel et al., 1989). A common approach for automatic B-point detection is the use of ensemble averaging, a technique that reduces noise and artifacts to simplify the reliable identification of fiducial points. However, this method has notable drawbacks: it can lead to the loss of relevant waveform details and is not suitable for real-time applications (Cieslak et al., 2018; Kelsey & Guethlein, 1990; Riese et al., 2003; Benouar et al., 2020). While beat-to-beat B-point detection addresses these issues, it is considerably more challenging than B-point detection on ensemble-averged signals. Nevertheless, a wide variety of algorithms were proposed in recent years to address this beat-to-beat detection problem. The most prominent family of beat-to-beat extraction algorithms applies predefined rules on the \(\text{dZ}/\text{dt}\) signal or its first- (\(d^{2}Z/dt^{2}\)) and second- (\(d^{3}Z/dt^{3}\)) order derivatives to identify the point corresponding to a specific, or the most appropriate morphological pattern of the actual B-point (Arbol et al., 2017, Bagal et al., 2017, Carvalho et al., 2011, Debski et al., 1993, DeMarzo & Lang, 1996, Drost et al., 2022, Forouzanfar et al., 2018, Naidu et al., 2014, Ono et al., 2004, Pale et al., 2021, Sherwood et al., 1990, Stern et al., 1985). Other methods have employed weighted windows that transform the signal to emphasize specific sections of the waveform (Miljković & Šekara, 2022), adaptive search windows to constrain the search area (Karpiel et al., 2022), matched filters applied to signal templates (Nagel et al., 1989), and empirical mode decomposition (Trybek et al., 2023). Furthermore, a variety of algorithms based on time-frequency analysis (Wang et al., 1995) and wavelet transformations (Hu et al., 2014; Nagel et al., 1989; Naidu et al., 2014; Shuguang et al., 2005) have been utilized. Similarly, data-driven B-point detection algorithms, using linear and quadratic regression formulas, have also been proposed to estimate the B-point location (Lozano et al., 2007). Despite the variety of algorithms, so far, no algorithm has been established as the recommended one, unlike for R-peak detection in the ECG, where an algorithm originally proposed by Pan & Tompkins (1985), which was continuously improved over the years, has been established as best practice. The primary reason for this is the lack of reproducibility, as the algorithm implementations and the data used for evaluation are often not publicly disclosed. This has made it difficult to conduct a fair comparison among the proposed algorithms. Moreover, the long-standing absence of a publicly available database with annotated fiducial points may explain why machine learning (ML) has not been widely adopted in this field of research, in contrast to its prevalence in other domains (Richer et al., 2025). To the best of our knowledge, machine learning and deep learning algorithms for B-point detection have been explored in only three published studies so far. Cieslak et al. (2018) reported a mean absolute error (MAE) of 1.3 ± 1.5 ms, the model proposed by Sheikh et al. (2022) achieved a MAE of 3.5 ± 9.2 ms, and the deep learning model developed by Wang et al. (2025) results in a MAE of around 5 ms. While these results are promising, all approaches have major limitations. The models were either trained on ensemble-averaged data (Cieslak et al., 2018; Sheikh et al., 2022) or, the datasets used for training and evaluation were relatively small (Wang et al., 2025; 417 cardiac cycles) or, not published alongside (Sheikh et al., 2022; Wang et al., 2025). A major obstacle in developing and comparing B-point detection methods, the lack of publicly available, annotated data, has recently been addressed. Several new datasets containing annotated\(\text{dZ}/\text{dt}\) signals are now accessible, including those from Miljković & Šekara (2023) and Pale et al. (2021). Furthermore, we published two datasets featuring both annotated electrocardiogram (ECG) and \(\text{dZ}/\text{dt}\) signals in previous work: EmpkinS (Richer et al., 2025a) and Guardian (Richer et al., 2025b). Building on these publicly available datasets, we proposed the PEPbench framework for systematic benchmarking of B-point detection algorithms. Utilizing this, we systematically evaluated most of the existing B-point detection algorithms in the field (Richer et al., 2025). Our result showed that, on the EmpkinS dataset, the algorithm proposed by Drost et al. (2022), performed best with a MAE of 14.9 ± 14.7 ms. In contrast, the linear regression-based algorithm developed by Lozano et al. (2007) reported the lowest MAE on the Guardian dataset, 16.7 ± 14.8 ms. This comprehensive benchmarking was a major step forward for standardizing the comparison of B-point detection methods. Building on that, we set out to improve the B-point detection by combining the strengths of existing algorithms to achieve unprecedented accuracy for beat-to-beat B-point detection through the use of machine learning. In this paper, we introduce a ML-based approach for automated, beat-to-beat B-point detection, as outlined in overview Figure 1. We trained and evaluated our proposed model on the same two publicly available datasets ( EmpkinS and Guardian ) from our previous work (Richer et al., 2025), utilizing over 11,000 cardiac cycles for a large-scale, direct comparison against already established algorithms. We leverage the existing domain knowledge inherent in published B-point detection algorithms, each possessing unique strengths and weaknesses, by combining their estimations into a unified ML model. This approach is designed to greatly enhance the overall accuracy and robustness of automatic B-point detection. 2. Methods 2.1 Dataset The proposed machine learning models were trained and evaluated on data provided by the PEPbench Framework (Richer et al., 2025). This dataset comprises a total of 11,611 annotated cardiac cycles of synchronized ECG and \(\text{dZ}/\text{dt}\) data from the EmpkinS and Guardian studies. Two independent, trained annotators performed annotations. The EmpkinS dataset contains 5,000 cardiac cycles, from n = 15 healthy participants (60 % women) who participated between December 2022 and May 2023. Participants had a mean age of 23.1 ± 2.6 years and a body mass index (BMI) of 21.9 ± 2.4 \(\text{kg}/m^{2}\) (M ± SD). Modified versions of the Trier Social Stress Test (TSST), which is a widely used protocol for psychosocial stress induction in the laboratory (Kirschbaum et al., 1993), and a friendly stress-free version of the TSST (f-TSST) (Wiemers et al., 2013), were performed on two consecutive days. The standard (f-)TSST protocols involve a 5-minute preparation period, followed by two 5-minute tasks: a mock job interview and a mental arithmetic task. Both tasks were performed in front of a two-person panel providing social-evaluative feedback. Since the EmpkinS study was also designed to evaluate radar-based stress assessment, the (f-)TSST protocol was adapted by introducing speech pauses between the tasks, which resulted in the modified (f-)TSST protocols. The Guardian dataset is a subset of a previously published dataset (Schellenberger et al., 2020), consisting of 6,611 cardiac cycles from n = 24 participants (50 % women) recruited between February and July 2018 at the Palliative Care Unit of the University Hospital Erlangen. Participants had a mean age of 31.2 ± 11.0 years and a body mass index (BMI) of 23.8 ± 3.5 \(\text{kg}/m^{2}\) (M ± SD). Before inclusion, a physician screened participants for health status by measuring heart rate (HR), blood pressure, and heart sounds, ensuring all values were within clinically acceptable ranges. Following screening, eligible participants were positioned on a tilt table in a supine position where electrodes for ECG and ICG recordings were attached. Following the approach of Schellenberger et al. (2020), the five-phase study began with a 10-minute baseline rest. Next, the Valsalva maneuver (20 s forced exhalation against a closed glottis) was performed three times to induce sympathetic activation (Gorlin et al., 1957; Sharpey-Schafer, 1955). In the next phase, participants were asked to follow breathing instructions that include holding their breath after inhalation and exhalation, and breathing normally between these tasks. The final two phases constituted an orthostatic challenge: a 10-minute 70-degree head-up tilt (Kenny et al., 1986; Van Zanten et al., 2024) followed by a 10-minute return to the supine position. For both studies, written informed consent was obtained from all eligible participants upon arrival. The studies adhered to the Declaration of Helsinki and were approved by the local ethics committee of FAU (protocol #493_20 B and protocol #85_15B). The experimental results from Richer et al. (2025) show physiological states with an average HR of 101.5 ± 24.3 beats per minute (bpm) in the range of 47.2 bpm to 157.9 bpm, and an average PEP of 88.4 ± 25.0 ms for the EmpkinS dataset, and an average HR of 67.6 ± 13.4 bpm in the range of 39.1 bpm to 138.3 bpm, and an average PEP of 138.4 ± 27.1 ms for the Guardian dataset. Detailed information on data collection, study protocols, and signal preprocessing is given in the work of Richer et al. (2025). The two datasets ( EmpkinS and Guardian ) were assessed separately in previous work. To enable a comparison with our new approach, we re-evaluated the performance of the existing methods on the combined data. The algorithm proposed by Drost et al. (2022) achieved the lowest MAE of 17.9 ± 23.0 ms. The second and third lowest errors were achieved by the linear regression model of Lozano et al. (2007) and the method proposed by Debski et al. (1993), with MAEs of 19.6 ± 17.5 ms and 20.5 ± 21.4 ms, respectively. Therefore, we will use the method by Drost et al. (2022) as a reference for automatic beat-to-beat B-point detection to compare our results to. Detailed results of other existing methods are provided in the supplementary material (Table S1). Figure 1: Overview of the main contributions. In this work we present a machine learning-based beat-to-beat detection algorithm of the aortic valve opening from impedance cardiography. Our best performing machine learning model is publicly available in the PEPbench framework, introduced in our previous work. 2.2 Feature extraction To develop a robust machine learning model for B-point detection, we chose a feature engineering approach that leverages established domain knowledge. Therefore, we utilized the PEPbench framework, which implements 12 B-point detection algorithms proposed in the scientific literature. The estimated B-point locations relative to the start of the cardiac cycle were combined with the duration of the RR interval to form our beat-to-beat feature set. To segment the ECG and\(\text{dZ}/\text{dt}\) signal in cardiac cycles, we detected the R-peak using the NeuroKit2 Python package (Makowski et al., 2021). Subsequently, we determined the start sample of a cardiac cycle by subtracting 35 % of the preceding RR interval from the location of the current cardiac cycle’s R-peak. We localized the end of the cardiac cycle by adding 65 % of the preceding RR interval to the current cardiac cycle’s R-peak (Richer et al., 2025). A detailed description of these extracted features is provided in Table 1. Original Publication Abbreviation Description Arbol et al., 2016 Arb17IC Last isoelectric crossing of a cardiac cycle before the maximum of the \(\text{dZ}/\text{dt}\) signal in that cardiac cycle (C-Point) (last crossing of the mean of the \(\text{dZ}/\text{dt}\) signal of a cardiac cycle before the C-Point). Arbol et al., 2016 Arb17SD Peak of the \(d^{2}Z/dt^{2}\) signal in the time window 150ms to 100ms before the C-Point. Arbol et al., 2016 Arb17TD Peak of the \(d^{3}Z/dt^{3}\) signal in a time window starting 300 ms before the C-Point. Debski et al., 1993 Deb93 Local minimum of the \(d^{2}Z/dt^{2}\) signal with minimal distance to the C-Point. Drost et al., 2022 Dro22 Point of maximal distance between the \(\text{dZ}/\text{dt}\) signal and a straight line that connects the C-Point and the point of the \(\text{dZ}/\text{dt}\) signal 150 ms before the C-Point. Forouzanfar et al., 2018 For18 Local maximum or zero-crossing in the \(d^{3}Z/dt^{3}\) signal with minimal distance to the C-Point in the most prominent monotonically increasing segment between the A-Point and the C-Point. In case no local maximum or zero crossing can be found, the B-point corresponds to the first point of the segment. Lozano et al., 2007 Loz07LR Linear regression model RB = 0.55 RC + 4.45 that exploits the close relationship between the RB-Interval and the RC-Interval (R represents the location of the R-Peak in the ECG signal. B, and C represent the B-point and C-Point, respectively). Lozano et al., 2007 Loz07QR Quadratic regression model RB = 1.233 RC - 0.0032 \(RC^{2}\) - 31.59 that exploits the close relationship between the RB-Interval and the RC-Interval (R represents the location of the R-Peak in the ECG signal. B, and C represent the B-point and C-Point, respectively). Miljković & Šekara, 2022 Mil22 The segment of the \(\text{dZ}/\text{dt}\) signal prior to the C-point is weighted by a scaling window to enhance the characteristic morphologies of the B-point, which simplifies detection. Pale et al., 2021 Pal21 B-point is searched in a time window ranging from C-Point - 80 ms to the point where half the amplitude of the C-Point is exceeded for the last time before the C-Point. Then the B-point is defined as the closest local minimum or point where a slope of 0.11 is exceeded before the C-Point. If no B-point is found, the slope threshold is decreased, and if B-point detection is still not successful, the minimum of the signal is defined as the B-point. Sherwood et al., 1990 She90 Zero-crossing of \(\text{dZ}/\text{dt}\) signal with minimal distance to the C-Point. Stern et al., 1985 Ste85 Local minimum of \(\text{dZ}/\text{dt}\) signal with minimal distance to the C-Point. derived from Makowski et al., 2021 RR interval RR interval: Distance between the R-peak of the current and previous cardiac cycle. Table 1: Overview of features used for training of the machine learning pipelines. To ensure comparability between the EmpkinS and Guardian datasets, which were recorded at different sampling rates, we upsampled the Guardian dataset (500 Hz) to match the sampling rate of the EmpkinS dataset (1000 Hz). Next, we performed the previously described segmentation into cardiac cycles and applied the feature extraction algorithms. Lastly, we subtracted the start sample of the current cardiac cycles from the estimated B-point location, to obtain the B-point location relative to the start of the cardiac cycle. In our previous work we found out that the B-point detection algorithms within the PEPbench framework were prone to two main types of failure: either a B-point was not detected at all, or its location was physiologically implausible, for example if the detected B-point is located prior to the R-peak location (Richer et al., 2025). In both cases, the failed detection was marked by a missing value. This limitation of the B-point detection algorithms has implications for our feature set. If all implemented B-point detection algorithms jointly failed to identify a valid B-point for a given cardiac cycle, we excluded this cycle entirely from the training data. This resulted in the exclusion of 372 cardiac cycles (3.2 %) from the dataset, shrinking the amount of usable data to 11,239 cardiac cycles. For the remaining cycles, the resulting feature set contained sporadic missing values for individual algorithm outputs and, occasionally, for the RR interval duration. To comprehensively evaluate the trade-offs associated with handling missing data and to explore the performance gains offered by a wider array of machine learning models, we defined two experimental datasets: Experiment 1 (Missing value inclusion): In this experiment we preserve the raw integrity of the data and restrict the analysis to ML models (e.g., RandomForestRegressor ) that can inherently handle missing feature values, resulting in 11,239 cardiac cycles. Experiment 2 (Missing value imputation): In this experiment, we imputed missing values using the median of the available B-point algorithm outputs within the same cardiac cycle to maintain beat-to-beat independence. Cycles with failed RR interval calculation (101 cycles, 0.9 %) were excluded to avoid inter-cycle imputation, resulting in a final dataset of 11,138 cardiac cycles. This allows to use non-parametric models, like SupportVectorRegressor ( SVR ). 2.3 Machine Learning Model Evaluation We integrated preprocessing steps, including feature scaling and dimensionality reduction with different machine learning models ( KNeighborsRegressor , SupportVectorRegressor , DecisionTreeRegressor , and RandomForestRegressor ) in a single, comprehensive pipeline. Many machine learning models, particularly those based on distance metrics or regularization, like KNeighborsRegressor and SVR, are susceptible to differences in feature scaling. Therefore, we employed two different scaling strategies: StandardScaler (transformation to zero mean and unit variance) and MinMaxScaler (rescaling to [0,1] range). Furthermore, we performed feature selection on the subset used in Experiment 2 to reduce complexity and improve model efficiency. This results in a simplified model structure that enhances both interpretability and generalization performance. For this purpose, we used two feature selection techniques. The first, SelectKBest , selects the top-K features based on a statistical test (F statistic or mutual information). The second, SelectFromModel , fits a RandomForestRegressor , whose design allows to automatically select features that meet a given importance threshold (Breiman, 2001). We incorporated both methods directly into our pipeline to ensure a fair and comprehensive model evaluation. In contrast to the procedure in Experiment 2, we did not apply feature selection in Experiment 1, because tree-based regression models, such as the RandomForestRegressor , are known to perform feature selection implicitly during their training process (Breiman, 2001). To find the set of model parameters that minimize the mean absolute error between the ML model predictions and the ground truth labels, we performed hyperparameter optimization. Therefore, we used grid search for the KNeighborsRegressor and SupportVectorRegressor , while a randomized search with 4,000 iterations was used for the DecisionTreeRegressor and RandomForestRegressor . To find the best performing ML model and the corresponding hyperparameters, we employed a 5-fold nested cross-validation (CV) approach. In 5-fold cross-validation, the training data is split into 5 subsets of equal size. In each iteration, the models are trained on 4 subsets and evaluated on the remaining subset of the data that serves as a test set. This procedure is repeated 5 times, and the final model performance is estimated by averaging over the individual performance metrics per fold. The inner loop of CV is used to tune the hyperparameters of the pipeline components that are currently evaluated. The outer CV loop is used to evaluate the generalizability of the fine-tuned pipeline combinations. Thereby, the inner CV is exclusively performed on the training set of the outer CV. This design ensures that the final evaluation of the model with the optimized hyperparameters from the inner CV loop is conducted on a completely unseen test set. Therefore, cross-validation provides a robust estimate of a model’s generalization performance on unseen data while mitigating the risk that the model learns the noise and random fluctuations present in the training data rather than the true underlying relationship, also known as overfitting (Yarkoni and Westfall, 2017). To ensure the model learned generalized features rather than specific intra-individual characteristics, we enforced a strict participant-level separation: all data belonging to a single participant was kept mutually exclusive between the training and test sets. We evaluated the performance of the machine learning pipelines using the mean absolute error (MAE) computed against the manually annotated B-point locations. Furthermore, we compared the best-performing machine learning model obtained in Experiments 1 and 2 to the previously best-performing B-point detection algorithm by Drost et al. (2022). To gain detailed insights into the models performance on different subsets of the data, we conducted further analyses on the model that achieved the lowest overall MAE. Firstly, we investigated the performance of the model in the case that at least one of the expert designed features fails. Therefore, we divided the dataset into two subsets. The first subset was the complete feature case, consisting of the 10,305 cardiac cycles with no missing values. The second subset contained the remaining 934 (8.3 %) and 833 (7.5 %) cardiac cycles for Experiment 1 and Experiment 2, respectively, and was characterized by at least one B-point detection algorithm failing feature extraction. As we expect the B-point to be more challenging to detect in these cases, we evaluated potential differences between the subsets. For the same reason, we explored whether the performance of the models depended on the inter-rater variability of the manually labeled ground truth. The inter-rater variability was evaluated in our previous work and was categorized by the difference of the manually labeled B-point location between both raters (Richer et al., 2025). We categorized the rater agreement into three levels: High agreement is defined as a difference smaller than or equal to 4 ms, medium agreement corresponds to a difference between 4 ms and 10 ms, and low agreement corresponds to deviations larger than 10 ms. For the PEPbench datasets, this categorization resulted in 66.1 % of the cardiac cycles showing high agreement, while 17.2 % and 16.7 % showed medium and low agreement, respectively. Thirdly, we evaluated the performance of the models on subsets grouped by the duration of the HR, which is inversely related to the RR interval. Therefore, we divided the dataset into three groups. The HR is considered to be low if it is lower than 60 bpm (18.9 %), medium between 60 and 100 bpm (58.8 %), and highover 100 bpm (22.3 %). We chose thresholds of 60 bpm and 100 bpm, as they represent the borders of a normal HR (Olshansky et al., 2023). To better understand the contributions of individual inputs, we analyzed the feature importances of the best-performing model. The importance scores were calculated using the Mean Decrease in Impurity (MDI), which is calculated implicitly for the model in each cross-validation fold (Pedregosa et al., 2011). Subsequently, we averaged the MDI score across all folds to obtain a robust estimate of the feature importances. 2.4 Statistical analysis We statistically evaluated performance differences between our novel machine learning approach and the traditional best-performing algorithm, as well as between several data subsets. Given that the Shapiro-Wilk test indicated violations of the assumption of normal distribution across all performance metrics, we used non-parametric statistical methods throughout the analysis. To compare the performance of the competing algorithms (novel ML approaches vs. the previous best-performing algorithm), we applied the Friedman test. When the Friedman test showed statistically significant differences, we performed post-hoc pairwise comparisons using the non-parametric Wilcoxon signed-rank test to identify specific differences. The subset of the data, where all feature extraction algorithms successfully estimated a B-point location was compared with the subset where at least one feature extraction algorithm failed, using the Mann-Whitney U test. The statistical differences between inter-rater agreement levels and the grouped HRs (both involving three groups) were assessed using the Kruskal-Wallis H test. Significant results were followed up with post-hoc pairwise comparisons using the non-parametric Wilcoxon signed-rank test. The significance level was set to α = 0.05. We accounted for multiple comparisons using the Bonferroni correction. All effect sizes are reported as Hedges’ g . Statistical significance is denoted as: * p < 0.05, ** p < 0.01, *** p < 0.001. 2.5 Availability of data and code The code used for training and evaluation of the proposed models is publicly available in the GitHub repository of the PEPbench framework (https://github.com/empkins/pepbench) under the MIT license. This framework also provides pipelines for automatic fiducial point detection based on algorithms from related work, along with comprehensive documentation. For further information, we refer to the introductory paper on the framework (Richer et al., 2025). The datasets used in this study are publicly available on the Open Science Framework (OSF) ( EmpkinS Dataset : https://doi.org/10.17605/OSF.IO/SH3XN, Guardian Dataset : https://doi.org/10.17605/OSF.IO/GYH75). All analyses were performed in Python (v3.12.6), using the package BioPsykit (v0.13.1; Richer et al., 2021), based on pingouin (v0.5.5; Vallat, 2018), and scikit-learn (v1.5.2; Pedregosa et al., 2011). 3. Results 3.1 General Model Performance The model evaluation revealed that the best performance was achieved by a pipeline consisting of a MinMaxScaler , no explicit feature selection, and a RandomForestRegressor . This pipeline achieved a MAE of 8.15 ± 1.14 ms, which was obtained in Experiment 1 (missing values were included). The three best-performing machine learning pipelines per Experiment with their respective mean absolute error (MAE), mean error (ME), and mean absolute relative error (MARE) are listed in Table 2. Detailed results of all pipeline combinations are provided in the supplementary material (Table S2). MAE (ms) ME (ms) MARE (%) Dataset Machine Learning Pipeline CV-folds Sample-wise Sample-wise Sample-wise Experiment 1 (Missing values included) MinMaxScaler - None - RandomForestRegressor 8.15 ± 1.14 8.13 ± 12.12 -0.03 ± 14.59 2.46 ± 3.92 StandardScaler - None - RandomForestRegressor 8.20 ± 1.14 8.18 ± 12.19 -0.02 ± 14.68 2.47 ± 3.95 MinMaxScaler - None - DecisionTreeRegressor 10.08 ± 0.89 10.06 ± 13.72 0.06 ± 17.02 3.02 ± 4.37 Experiment 2 (Median imputation) StandardScaler - SelectFromModel - RandomForestRegressor 8.31 ± 1.06 8.30 ± 12.30 -0.03 ±14.84 2.51 ± 3.92 MinMaxScaler - KBest - RandomForestRegressor 8.35 ± 1.31 8.33 ± 12.30 0.00 ± 14.85 2.53 ± 3.93 MinMaxScaler - SelectFromModel - RandomForestRegressor 8.42 ± 1.15 8.40 ± 12.36 -0.06 ± 14.95 2.55 ± 3.97 Table 2: B-point detection performance of the three best machine learning pipelines evaluated on the datasets of Experiment 1 and Experiment 2. Mean absolute error (MAE), mean error (ME), and mean absolute relative error (MARE) were evaluated across cross-validation folds and sample-wise. All values are depicted as Mean ± Standard Deviation. The machine learning pipelines developed in this work consistently and significantly outperformed all traditional B-point detection algorithms, previously evaluated on the same datasets (Richer et al., 2025). Specifically, RandomForestRegressor -based pipelines demonstrated the lowest errors. In Experiment 1, the pipeline (MinMaxScaler - RandomForestRegressor), which is the overall best-performing pipeline, significantly outperformed the best traditional method by Drost et al. (2022) ( p < 0.001, g = 0.53), by reducing the MAE by 54.7 % from 17.9 ms to 8.1 ms. The standard deviation dropped by 47.4 % from 23.0 ms to 12.1 ms. In Experiment 2, the pipeline (StandardScaler - SelectFromModel - RandomForestRegressor) achieved the lowest MAE of 8.3 ± 12.3 ms. This pipeline is also significantly better than the traditional method proposed by Drost et al. (2022) ( p < 0.001, g = 0.52; Figure 2). Figure 2: Comparison of the absolute error (ms) of the best-performing classic B-point detection algorithm (Drost et al., 2022) and the proposed best pipeline for Experiment 1, and Experiment 2. Mean values are denoted by the white cross. Comparing the best-performing machine learning pipelines of each experiment (missing value inclusion vs. missing value imputation) to each other, a minimal, but significant ( p < 0.001) difference in MAE of 0.17 ms is obtained. However, the effect size ( g = -0.01) is small, indicating high similarity between the outcomes of the two experiments. Because Experiment 1 included more cardiac cycles and obtained a slightly lower mean absolute error than Experiment 2, the following detailed analyses refer to the best-performing model ( MinMaxScaler - RandomForestRegressor ) obtained in Experiment 1. The residual plot of this best-performing machine learning pipeline against the manually labeled reference is illustrated in Figure 4. The error on the subset with a full feature set is significantly lower than the error obtained from the set, where at least one feature detection algorithm failed (7.76 ± 11.42 vs. 12.20 ± 17.60 ms MAE; p < 0.001, g = -0.37; Figure 3a). Similarly, the inter-rater agreement had an effect on model performance with the error decreasing for higher agreement ( low : 15.17 ± 13.76 ms MAE, medium : 9.00 ± 12.76 ms MAE, high : 5.86 ± 9.93 ms MAE) (Figure 3b). The observed changes between all groups are statistically significant ( high vs. low : p < 0.001, g = -0.86; high vs. medium : p < 0.001, g = -0.30; medium vs. low : p < 0.001, g = -0.46). The error of the machine learning model across different HRs decreases slightly from low HRs to high HRs (Figure 3c). Statistical analysis revealed significant differences between the model’s performance on the high HR and low HR subsets ( p = 0.002, g = -0.17) as well as between the high HR and medium HR subsets ( p = 0.008, g = -0.11). No statistically significant changes could be found between the subsets of medium HR and high HR durations ( p = 0.608, g = 0.05). Figure 3: Evaluation of the best performing machine learning pipeline (MinMaxScaler - RandomForestRegressor) on three subsets of the data. 3a: Comparison of the subsets where either all features were detected or at least one feature extraction algorithm failed. 3b: Evaluation on the subsets of low, medium, and high inter-rater agreement. 3c: Evaluation for different heart rate ranges. Mean values are denoted by the white cross. Using the annotations of the second annotator as reference values for computing the performance of our proposed model we obtained a MAE of 10.02 \(\pm\) 14.54 ms, which is an increased error of 1.67 ms compared to the MAE of 8.35 \(\pm\) 14.78 ms when evaluated on the labels of the first annotator. Figure 4: Residual plot of the best performing machine learning pipeline (MinMaxScaler - RandomForestRegressor) Figure 5: Regression plot of the best performing machine learning pipeline (MinMaxScaler - RandomForestRegressor) over the heart rate. 3.2 Feature importances Averaging the feature importances of the best-performing model across the cross-validation folds revealed that the model is primarily influenced by the B-point detection algorithm proposed by Drost et al. (2022), which had a feature importance score of 74.9 %. The linear regression model by Lozano et al. (2007) was also a critical contributor to the model’s decision-making, with a feature importance of 15.9 %, followed by input from the algorithm by Debski et al. (1993) and the RR interval, with feature importances of 4.4 % and 1.2 %, respectively. The combined feature importance of these four most influential features reached 96.4 %. A comprehensive overview of all feature importance scores is provided in Table 3. Short description Abbreviation Importance Score (%) Point of maximal distance to a straight line between the C-point and the point 150 ms before (Drost et al., 2022) Dro22 74.9 ± 1.8 Linear regression model RB = 0.55 RC + 4.45 (Lozano et al., 2007) Loz07LR 15.9 ± 1.5 Local minimum of second derivative (Debski et al., 1993) Deb93 4.4 ± 0.7 RR-Interval (derived from Makowski et al., 2021) RR interval 1.2 ± 0.3 Local maximum or zero crossing in \(d^{3}Z/dt^{3}\) signal (Forouzanfar et al., 2018) For18 0.8 ± 0.1 Quadratic regression model RB = 1.233 RC - 0.0032 \(RC^{2}\) - 31.59 (Lozano et al., 2007) Loz07QR 0.7 ± 0.1 Peak of the \(d^{2}Z/dt^{2}\) signal in the time window 150ms to 100ms before the C-Point. (Arbol et al., 2016) Arb16SD 0.6 ± 0.1 Local minimum of \(\text{dZ}/\text{dt}\) signal with minimal distance to the C-Point (Stern et al., 1985) Ste85 0.6 ± 0.0 Weighted window approach (Miljković & Šekara, 2022) Mil22 0.3 ± 0.1 Peak of the \(d^{3}Z/dt^{3}\) signal in a time window (Arbol et al., 2016) Arb16TD 0.3 ± 0.3 Zero-crossing of \(\text{dZ}/\text{dt}\) signal with minimal distance to the C-Point. (Sherwood et al., 1990) She90 0.2 ± 0.1 Last isoelectric crossing of a cardiac cycle before the C-Point (Arbol et al., 2016) Arb16IC 0.1 ± 0.0 Rule-based detection in \(\text{dZ}/\text{dt}\) signal (Pale et al., 2021) Pal21 0.1 ± 0.0 Table 3: Feature importance scores (Mean ± Standard Deviation) of the best performing machine learning model (MinMaxScaler - RandomForestRegressor) averaged over the 5 cross-validation folds. Refer to Table 1 for a detailed description of the B-point detection methods. 4. Discussion In this paper, we introduced and evaluated a novel beat-to-beat B-point detection method using machine learning. Our results show that we can reliably detect the B-point from ECG and \(\text{dZ}/\text{dt}\) signals with a significantly reduced absolute error compared to the best-performing traditional algorithm by Drost et al. (2022), from 17.9 ± 23.0 ms to 8.13 ± 12.12 ms. While the machine learning models provided in related work achieved lower MAEs than our proposed method (Cieslak et al., 2018; Sheikh et al., 2022), our approach has some key advantages. Most importantly, we performed B-point detection on a beat-to-beat basis. Therefore, we do not rely on ensemble averaging, which substantially simplifies automatic detection by smoothing the \(\text{dZ}/\text{dt}\) signal. Ensemble averaging, and even moving ensemble averaging, may lack sufficient temporal resolution to reflect rapid changes in hemodynamic parameters such as PEP or LVET (Carvalho et al., 2011). To our knowledge, only Wang et al. (2025) performed machine learning-based B-point detection on a beat-to-beat basis and achieved an MAE of less than 5 ms with their deep learning model. However, the dataset they used for training and evaluation contained only 417 cardiac cycles. Furthermore, the dataset was labeled by a single annotator, and no detailed information on the study population or physiological ranges of the dataset was provided. In contrast, we trained our model on a dataset with more than 11,000 cardiac cycles, labeled by two trained annotators. This allowed us to evaluate our machine learning pipeline on subsets of low, medium, and high inter-rater variability. We also calculated the performance against the annotations of the second rater, which were never seen during training. The achieved MAE of 10.02 ± 14.54 ms indicates that our model can produce high-quality results on unseen labels. To gain a deeper understanding of the best model’s performance, further analyses were carried out on different subsets of the data, where we expect to see differences in B-point detection difficulty. The model exhibited a significantly higher absolute error on the subset where at least one feature extraction algorithm was unable to detect a B-point. A potential explanation for this difference is that the failure of a feature detection algorithm serves as an indicator of an inherently challenging B-point morphology. The inability of an existing algorithm to find a B-point suggests that one or more of its typical defining characteristics might be absent or distorted in that specific cardiac cycle. This leads to a decrease in model performance, explaining the observed increase in absolute error within this subset of data. Further supporting this hypothesis, inter-rater agreement serves as an additional measure of B-point detection difficulty. As depicted in Figure 3b, the model’s error significantly decreases with increasing inter-rater agreement. This finding indicates that when raters exhibit low agreement on the reference B-point position, the task is inherently more challenging, not only for the human annotators but also for our proposed method. Nevertheless, the performance we obtained for the low agreement subset is on a comparable level to the performance of traditional models for the high agreement subset (15.17 ± 13.76 ms and 16.11 ± 20.67 ms, respectively). Therefore, our approach is able to effectively reduce the error for the difficult-to-detect B-points. Our results revealed a trend that the absolute error of the B-point detection decreases with higher HRs, as shown in Figure 5. This is further backed by the significant differences in absolute error between the low HR subset and both the medium and high HR subsets. Two potential explanations for this are: Specific B-point morphologies may be more distinct at higher HRs, and the compressed timescale of shorter RR intervals (high HRs) may inherently reduce the potential for a large absolute error. In this scenario, even if the model’s prediction is incorrect, the predicted B-point location is geometrically closer to the true reference B-point, thereby reducing the magnitude of the error. In conclusion, even in cases where reliable detection proves more challenging, the proposed machine learning model is still consistently outperforming the algorithm proposed by Drost et al. (2022), which showed the best performance in the comparative analysis in our previous work (Richer, 2025). While our analyses revealed that the algorithm proposed by Drost et al. (2022) is estimating the actual B-point location too early, which can be seen by the negative MEs of -10.1\(\pm\) 18.4 ms ( EmpkinS ) and -14.8 \(\pm\) 18.6 ms ( Guardian ), the proposed machine learning pipeline achieves a ME of -0.03 ± 14.59 and therefore reduces this bias to almost zero. Furthermore, the model achieves limits of agreement (\(\pm\ 1.96\ \times\ \text{SD}\)) of -28.64 ms and 28.57 ms, thereby reducing them compared to the best traditional algorithm ( EmpkinS dataset: -46.0 ms and 29.9 ms; Richer et al., 2025). Therefore, the developed machine learning pipeline is able to reduce the range from the upper limit of agreement to the lower limit of agreement substantially. The error of our proposed model (8.13 ± 12.12 ms) is close to the inter-rater error (8.35 ± 14.78 ms). Despite the inherent challenges in beat-to-beat B-point detection, this suggests that the performance of our automatic extraction method is on a comparable level to the current standard of manual annotation, while eliminating the bias introduced by individual human raters. The analysis of feature importances revealed a strong correlation between feature importance and the accuracy of its underlying B-point detection algorithm. While the model is primarily influenced by the algorithms with the lowest error in the benchmarking, specifically those proposed by Drost et al. (2022), Lozano et al. (2007), and Debski et al. (1993), algorithms with high errors have less influence on the model’s decision-making. Fusing the algorithms in a single model, even though some of them have minor influence, potentially reduces the large error that some of the algorithms obtain with certain waveforms. While there are differences in model performance across different HR ranges, the importance of the RR interval is low. A potential explanation for this is that the RR interval may be a critical indicator only for a small subset of extreme cases. Consequently, while it might provide critical information for these few instances, its overall importance is low when averaged across the entire dataset, where most cardiac cycles fall within a medium RR interval range. The proposed method, while demonstrating significant improvement, reveals two principal areas for necessary future investigation to ensure utility in psychophysiological research and external validity across diverse experimental conditions. The most immediate limitation concerns the generalizability of the model, which must be evaluated beyond the current datasets. Given that the datasets primarily comprise data from young and healthy participants, the demonstrated accuracy is confined to a cohort with minimal hemodynamic complexity. The model’s current validation must be expanded to include groups where physiology is inherently altered. This requires validation on populations from studies with pharmacological interventions (e.g. Clark et al., 2015), and clinical populations, such as major depressive disorder (Bair et al., 2021), where the B-point is used for PEP computation. A secondary, yet equally important limitation stems from the insufficient representation of physiological extremes within the sampled data. The small amount of cardiac cycles corresponding to states of very low and high HRs compromises the model’s reliability. Since psychophysiological experiments often involve stimuli designed to induce rapid and extreme changes in HR (e.g., startling sounds or intense cognitive tasks), the model’s robustness across the full range of HR variation remains to be established. Therefore, future research should include more data specifically focusing on these conditions to prove the method’s robustness. To select a model for practical application, we compared the best-performing machine learning model from each experiment. The model showing the overall best performance in terms of MAE ( MinMaxScaler - RandomForestRegressor ) is trained on a slightly larger dataset (Experiment 1), as it includes missing values of the RR interval feature. Consequently, this model has a more robust feature detection pipeline and is proposed to be used in future work. To enable this, we retrained this pipeline on the two entire datasets and implemented this model as part of a new algorithm BPointExtractionStuehler2025 11Will be made available upon potential acceptance of this manuscript in the PEPbench framework. 5. Conclusion In this paper, we introduced a novel machine learning approach for accurate aortic valve opening detection on a beat-to-beat level using the B-point. By building our feature set on existing B-point detection algorithms, we leveraged existing domain knowledge in impedance cardiography analysis to achieve a significant boost in performance over related work. Despite promising advancements in data-driven B-point detection methods (Cieslak et al., 2018; Sheikh et al., 2022; Wang et al., 2025), limitations remain, such as the reliance on unpublished datasets and the use of ensemble averaging. In contrast, our proposed algorithm was trained and evaluated on the publicly available PEPbench beat-to-beat datasets and utilizes features that can be easily extracted using the pipelines within that framework (Richer et al., 2025). This approach provides a straightforward method with results on human rater level, establishing a new benchmark for further developments in this area. Therefore, the accurate and reliable extraction of key time-domain measures, like the Pre-Ejection Period (PEP) and Left Ventricular Ejection Time (LVET), is possible. In future work, we plan to evaluate how the performance of the proposed machine learning model generalizes to unseen datasets. Furthermore, deep learning methods for automatic B-point detection could be explored. We hypothesize that these techniques further improve automated B-point detection, consistent with the notable advancements that have been achieved in other research domains with the application of deep learning. Author Contributions Luca Abel: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing; Sebastian Stuehler: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing; Tobias Steigleder: Data curation, Resources, Supervision, Writing – review & editing; Christoph Ostgathe: Conceptualization, Funding acquisition, Project administration, Resources, Supervision, Writing – review & editing; Nicolas Rohleder: Conceptualization, Funding acquisition, Project administration, Resources, Supervision, Writing – review & editing; Bjoern M Eskofier: Conceptualization, Funding acquisition, Project administration, Resources, Supervision, Writing – review & editing; Robert Richer: Conceptualization, Data curation, Investigation, Methodology, Software, Supervision, Validation, Writing – original draft, Writing – review & editing Acknowledgments We would like to thank all participants who took part in our study. During the preparation of this work, AI technologies were used to assist in the writing process. Specifically, Grammarly (Grammarly Inc., San Francisco, CA, USA) was used in order to check for grammar and style consistency. ChatGPT (GPT-4o) (OpenAI, San Francisco, CA, USA) and Gemini 3 (Google, Mountain View, CA, USA) were used to rephrase selected passages and improve clarity and readability. These tools were used solely for linguistic and editorial support, not for conceptual content generation. All scientific content, ideas, and interpretations presented in this manuscript are the original work of the authors. The final version was carefully reviewed and revised by the authors, and the use of AI tools is disclosed here in accordance with COPE guidelines. The authors take full responsibility for the integrity and accuracy of the published content. Funding This work was partly funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – SFB 1483 – Project-ID 442419336, EmpkinS. Ethics Statement Both studies in this manuscript were approved by the local ethics committee of FAU (protocol #493_20 B and #264_17 B) and were conducted in accordance with the Declaration of Helsinki. Conflicts of Interest The authors declare no conflicts of interest. Data and Code Availability Statement The two datasets used in this experiment are available on the Open Science Framework (OSF) platform (EmpkinS Dataset: https://doi.org/10.17605/OSF.IO/SH3XN, Guardian Dataset: https://doi.org/10.17605/OSF.IO/GYH75). The source code is available on GitHub (https://github.com/empkins/pepbench). References Árbol, J. R., Perakakis, P., Garrido, A., Mata, J. L., Fernández‐Santaella, M. C., & Vila, J. (2017). Mathematical detection of aortic valve opening (B point) in impedance cardiography: A comparison of three popular algorithms. Psychophysiology , 54 (3), 350–357. https://doi.org/10.1111/psyp.12799 Alhakak, A. S., Teerlink, J. R., Lindenfeld, J., Böhm, M., Rosano, G. M. C., & Biering‐Sørensen, T. (2021). The significance of left ventricular ejection time in heart failure with reduced ejection fraction. European Journal of Heart Failure, 23(4), 541–551. https://doi.org/10.1002/ejhf.2125 Bagal, U. R., Pandey, P. C., Naidu, S. M. M., & Hardas, S. P. (2017). Detection of opening and closing of the aortic valve using impedance cardiography and its validation by echocardiography. Biomedical Physics & Engineering Express , 4 (1), 015012. https://doi.org/10.1088/2057-1976/aa8bf5 Bair, A., Marksteiner, J., Falch, R., Ettinger, U., Reyes Del Paso, G. A., & Duschek, S. (2021). Features of autonomic cardiovascular control during cognition in major depressive disorder. Psychophysiology , 58 (1), e13628. https://doi.org/10.1111/psyp.13628 Benouar, S., Hafid, A., Kedir-Talha, M., & Seoane, F. (2021). First Steps Toward Automated Classification of Impedance Cardiography dZ/dt Complex Subtypes. In T. Jarm, A. Cvetkoska, S. Mahnič-Kalamiza, & D. Miklavcic (Eds.), 8th European Medical and Biological Engineering Conference (Vol. 80, pp. 563–573). Springer International Publishing. https://doi.org/10.1007/978-3-030-64610-3_64 Breiman, L. (2001). Random Forests. Machine Learning , 45 (1), 5–32. https://doi.org/10.1023/A:1010933404324 Carvalho, P., Paiva, R. P., Henriques, J., Antunes, M., & Quintal, I. (2011). Robust Characteristic Points for ICG - Definition and Comparative Analysis. Proceedings of the International Conference on Bio-Inspired Systems and Signal Processing , 161–168. https://doi.org/10.5220/0003134901610168 Cieslak, M., Ryan, W. S., Babenko, V., Erro, H., Rathbun, Z. M., Meiring, W., Kelsey, R. M., Blascovich, J., & Grafton, S. T. (2018). Quantifying rapid changes in cardiovascular state with a moving ensemble average. Psychophysiology , 55 (4), e13018. https://doi.org/10.1111/psyp.13018 Clark, C. M., Frye, C. G., Wardle, M. C., Norman, G. J., & De Wit, H. (2015). Acute effects of MDMA on autonomic cardiac activity and their relation to subjective prosocial and stimulant effects. Psychophysiology , 52 (3), 429–435. https://doi.org/10.1111/psyp.12327 Debski, T. T., Zhang, Y., Jennings, J. R., & Kamarck, T. W. (1993). Stability of cardiac impedance measures: Aortic opening (B-point) detection and scoring. Biological Psychology , 36 (1–2), 63–74. https://doi.org/10.1016/0301-0511(93)90081-I DeMarzo, A. P., & Lang, R. M. (1996). A new algorithm for improved detection of aortic valve opening by impedance cardiography. Computers in Cardiology 1996 , 373–376. https://doi.org/10.1109/CIC.1996.542551 Drost, L., Finke, J. B., Port, J., & Schächinger, H. (2022). Comparison of TWA and PEP as indices of α2- and ß-adrenergic activation. Psychopharmacology , 239 (7), 2277–2288. https://doi.org/10.1007/s00213-022-06114-8 Forouzanfar, M., Baker, F. C., Colrain, I. M., Goldstone, A., & De Zambotti, M. (2019). Automatic analysis of pre‐ejection period during sleep using impedance cardiogram. Psychophysiology , 56 (7), e13355. https://doi.org/10.1111/psyp.13355 Forouzanfar, M., Baker, F. C., De Zambotti, M., McCall, C., Giovangrandi, L., & Kovacs, G. T. A. (2018). Toward a better noninvasive assessment of preejection period: A novel automatic algorithm for B‐point detection and correction on thoracic impedance cardiogram. Psychophysiology , 55 (8), e13072. https://doi.org/10.1111/psyp.13072 Gorlin, R., Knowles, J. H., & Storey, C. F. (1957). The valsalva maneuver as a test of cardiac function. The American Journal of Medicine , 22 (2), 197–212. https://doi.org/10.1016/0002-9343(57)90004-9 Hu, X., Chen, X., Ren, R., Zhou, B., Qian, Y., Null, H. L., & Xia, S. (2014). Adaptive Filtering and Characteristics Extraction for Impedance Cardiography. Journal of Fiber Bioengineering and Informatics , 7 (1), 81–90. https://doi.org/10.3993/jfbi03201407 Karpiel, I., Richter-Laskowska, M., Feige, D., Gacek, A., & Sobotnicki, A. (2022). An Effective Method of Detecting Characteristic Points of Impedance Cardiogram Verified in the Clinical Pilot Study. Sensors , 22 (24), 9872. https://doi.org/10.3390/s22249872 Kelsey, R. M., & Guethlein, W. (1990). An Evaluation of the Ensemble Averaged Impedance Cardiogram. Psychophysiology , 27 (1), 24–33. https://doi.org/10.1111/j.1469-8986.1990.tb02173.x Kenny, R. A., Bayliss, J., Ingram, A., & Sutton, R. (1986). Head-Up Tilt: A Useful Test For Invstigating Unexplained Syncope. The Lancet , 327 (8494), 1352–1355. https://doi.org/10.1016/S0140-6736(86)91665-X Kirschbaum, C., Kudielka, B. M., Gaab, J., Schommer, N. C., & Hellhammer, D. H. (1999). Impact of gender, menstrual cycle phase, and oral contraceptives on the activity of the hypothalamus-pituitary-adrenal axis. Psychosomatic Medicine , 61 (2), 154–162. https://doi.org/10.1097/00006842-199903000-00006 Lozano, D. L., Norman, G., Knox, D., Wood, B. L., Miller, B. D., Emery, C. F., & Berntson, G. G. (2007). Where to B in dZ/dt. Psychophysiology , 44 (1), 113–119. https://doi.org/10.1111/j.1469-8986.2006.00468.x Makowski, D., Pham, T., Lau, Z. J., Brammer, J. C., Lespinasse, F., Pham, H., Schölzel, C., & Chen, S. H. A. (2021). NeuroKit2: A Python toolbox for neurophysiological signal processing. Behavior Research Methods, 53(4), 1689–1696. https://doi.org/10.3758/s13428-020-01516-y Miljković, N., & Šekara, T. B. (2022). A New Weighted Time Window-based Method to Detect B-point in Impedance Cardiogram (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2207.04490 Miljković, N., & Šekara B, T. (2023). Software for Detection of B-point in Imepedance Cardiogram with Dataset Consisting of 20 Healthy Participants (Version 3) [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.6813716 Nagel, J. H., Shyu, L. Y., Reddy, S. P., Hurwitz, B. E., McCabe, P. M., & Schneiderman, N. (1989). New signal processing techniques for improved precision of noninvasive impedance cardiography. Annals of Biomedical Engineering , 17 (5), 517–534. https://doi.org/10.1007/BF02368071 Naidu, S. M. M., Bagal, U. R., Pandey, P. C., Hardas, S., & Khambete, N. D. (2014). Detection of characteristic points of impedance cardiogram and validation using Doppler echocardiography. 2014 Annual IEEE India Conference (INDICON) , 1–6. https://doi.org/10.1109/INDICON.2014.7030596 Olshansky, B., Ricci, F., & Fedorowski, A. (2023). Importance of resting heart rate. Trends in Cardiovascular Medicine , 33 (8), 502–515. https://doi.org/10.1016/j.tcm.2022.05.006 Ono, T., Miyamura, M., Yasuda, Y., Ito, T., Saito, T., Ishiguro, T., Yoshizawa, M., & Yambe, T. (2004). Beat-to-Beat Evaluation of Systolic Time Intervals during Bicycle Exercise Using Impedance Cardiography. The Tohoku Journal of Experimental Medicine , 203 (1), 17–29. https://doi.org/10.1620/tjem.203.17 Pale, U., Muller, N., Arza, A., & Atienza, D. (2021). ReBeatICG: Real-time Low-Complexity Beat-to-beat Impedance Cardiogram Delineation Algorithm. 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) , 5618–5624. https://doi.org/10.1109/EMBC46164.2021.9630170 Pan, J., & Tompkins, W. J. (1985). A Real-Time QRS Detection Algorithm. IEEE Transactions on Biomedical Engineering , BME-32 (3), 230–236. https://doi.org/10.1109/TBME.1985.325532 Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research , 12 (Oct), 2825–2830. Richer, R., Jorkowitz, J., Stühler, S., Abel, L., Kurz, M., Oesten, M., Griesshammer, S. G., Albrecht, N. C., Küderle, A., Ostgathe, C., Kölpin, A., Steigleder, T., Rohleder, N., & Eskofier, B. M. (2025). PEPbench—Open, Reproducible, and Systematic Benchmarking of Automated Pre‐Ejection Period Extraction Algorithms. Psychophysiology , 62 (11), e70176. https://doi.org/10.1111/psyp.70176 Richer, R., Jorkowitz, J., Stühler, S., Abel, L., Kurz, M., Oesten, M., … Eskofier, B. (2025a, July 30). PEP Benchmarking – EmpkinS Dataset. https://doi.org/10.17605/OSF.IO/SH3XN Richer, R., Jorkowitz, J., Stühler, S., Abel, L., Oesten, M., Kurz, M., … Eskofier, B. (2025b, July 30). PEP Benchmarking – Guardian Dataset. https://doi.org/10.17605/OSF.IO/GYH75 Richer, R., Küderle, A., Ullrich, M., Rohleder, N., & Eskofier, B. (2021). BioPsyKit: A Python package for the analysis of biopsychological data. Journal of Open Source Software , 6 (66), 3702. https://doi.org/10.21105/joss.03702 Riese, H., Groot, P. F. C., Van Den Berg, M., Kupper, N. H. M., Magnee, E. H. B., Rohaan, E. J., Vrijkotte, T. G. M., Willemsen, G., & De Geus, E. J. C. (2003). Large-scale ensemble averaging of ambulatory impedance cardiograms. Behavior Research Methods, Instruments, & Computers , 35 (3), 467–477. https://doi.org/10.3758/BF03195525 Schellenberger, S., Shi, K., Steigleder, T., Malessa, A., Michler, F., Hameyer, L., Neumann, N., Lurz, F., Weigel, R., Ostgathe, C., & Koelpin, A. (2020). A dataset of clinically recorded radar vital signs with synchronised reference sensor signals. Scientific Data , 7 (1), 291. https://doi.org/10.1038/s41597-020-00629-5 Sharpey-Schafer, E. P. (1955). Effects of Valsalva’s Maneuver on the Normal and Failing Circulation. BMJ , 1 (4915), 693–695. https://doi.org/10.1136/bmj.1.4915.693 Sheikh, S. A., Gurel, N. Z., Gupta, S., Chukwu, I. V., Levantsevych, O., Alkhalaf, M., Soudan, M., Abdulbaki, R., Haffar, A., Vaccarino, V., Inan, O. T., Shah, A. J., Clifford, G. D., & Bahrami Rad, A. (2022). Data‐driven approach for automatic detection of aortic valve opening: B point detection from impedance cardiogram. Psychophysiology , 59 (12), e14128. https://doi.org/10.1111/psyp.14128 Sherwood, A., Allen, M. T., Fahrenberg, J., Kelsey, R. M., Lovallo, W. R., & Van Doornen, L. J. P. (1990). Methodological Guidelines for Impedance Cardiography. Psychophysiology , 27 (1), 1–23. https://doi.org/10.1111/j.1469-8986.1990.tb02171.x Smyth, J. M., Ockenfels, M. C., Gorin, A. A., Catley, D., Porter, L. S., Kirschbaum, C., Hellhammer, D. H., & Stone, A. A. (1997). Individual differences in the diurnal cycle of cortisol. Psychoneuroendocrinology , 22 (2), 89–105. https://doi.org/10.1016/S0306-4530(96)00039-X Stern, H., Wolf, G., & Belz, G. (1985). Comparative measurements of left ventricular ejection time by mechano-, echo- and electrical impedance cardiography. Arzneimittel-Forschung , 35 (10), 1582—1586. Trybek, P., Sobotnicka, E., Wawrzkiewicz-Jałowiecka, A., Machura, Ł., Feige, D., Sobotnicki, A., & Richter-Laskowska, M. (2023). A New Method of Identifying Characteristic Points in the Impedance Cardiography Signal Based on Empirical Mode Decomposition. Sensors , 23 (2), 675. https://doi.org/10.3390/s23020675 Vallat, R. (2018). Pingouin: Statistics in Python. Journal of Open Source Software , 3 (31), 1026. https://doi.org/10.21105/joss.01026 Van Zanten, S., Sutton, R., Hamrefors, V., Fedorowski, A., & De Lange, F. J. (2024). Tilt table testing, methodology and practical insights for the clinic. Clinical Physiology and Functional Imaging , 44 (2), 119–130. https://doi.org/10.1111/cpf.12859 Wang, D., Richter, M., Zhu, X., Pedrycz, W., Gacek, A., Sobotnicki, A., & Li, Z. (2025). Detecting Characteristic Points for the Analysis of Bioimpedance Signal Through a Synergy of Fuzzy Rule-Based Models and Granular Neural Networks. IEEE Transactions on Fuzzy Systems , 33 (8), 2460–2468. https://doi.org/10.1109/TFUZZ.2024.3454335 Wiemers, U. S., Schoofs, D., & Wolf, O. T. (2013). A friendly version of the Trier Social Stress Test does not activate the HPA axis in healthy men and women. Stress , 16 (2), 254–260. https://doi.org/10.3109/10253890.2012.714427 Xiang Wang, Sun, H. H., & Van De Water, J. M. (1995). An advanced signal processing technique for impedance cardiography. IEEE Transactions on Biomedical Engineering , 42 (2), 224–230. https://doi.org/10.1109/10.341836 Yarkoni, T., & Westfall, J. (2017). Choosing Prediction Over Explanation in Psychology: Lessons From Machine Learning. Perspectives on Psychological Science , 12 (6), 1100–1122. https://doi.org/10.1177/1745691617693393 Zhao Shuguang, Fang Yanhong, Zhao Hailong, & Tang Min. (2005a). Detection of Impedance Cardiography’s Characteristic Points Based on Wavelet Transform. 2005 IEEE Engineering in Medicine and Biology 27th Annual Conference , 2730–2732. https://doi.org/10.1109/IEMBS.2005.1617035 Information & Authors Information Version history V1 Version 1 17 December 2025 Copyright This work is licensed under a Non Exclusive No Reuse License. Authors Affiliations Luca Abel 0000-0002-5044-7113 [email protected] Friedrich-Alexander-Universitat Erlangen-Nurnberg View all articles by this author Sebastian Stühler Friedrich-Alexander-Universitat Erlangen-Nurnberg View all articles by this author Tobias Steigleder Universitatsklinikum Erlangen Palliativmedizinische Abteilung View all articles by this author Christoph Ostgathe Universitatsklinikum Erlangen Palliativmedizinische Abteilung View all articles by this author Nicolas Rohleder 0000-0003-2602-517X Friedrich-Alexander-Universitat Erlangen-Nurnberg View all articles by this author Bjoern M. Eskofier Friedrich-Alexander-Universitat Erlangen-Nurnberg View all articles by this author Robert Richer 0000-0003-0272-5403 Friedrich-Alexander-Universitat Erlangen-Nurnberg View all articles by this author Metrics & Citations Metrics Article Usage 178 views 119 downloads .FvxKWukQNSOunydq8rnd { width: 100px; } Citations Download citation Luca Abel, Sebastian Stühler, Tobias Steigleder, et al. Beat-to-beat aortic valve opening detection from impedance cardiography using machine learning. Authorea . 17 December 2025. DOI: https://doi.org/10.22541/au.176595493.35093262/v1 If you have the appropriate software installed, you can download article citation data to the citation manager of your choice. Simply select your manager software from the list below and click Download. For more information or tips please see 'Downloading to a citation manager' in the Help menu . Format Please select one from the list RIS (ProCite, Reference Manager) EndNote BibTex Medlars RefWorks Direct import Tips for downloading citations document.getElementById('citMgrHelpLink').addEventListener('click', function() { popupHelp(this.href); return false; }); $(".js__slcInclude").on("change", function(e){ if ($(this).val() == 'refworks') $('#direct').prop("checked", false); $('#direct').prop("disabled", ($(this).val() == 'refworks')); }); View Options View options PDF View PDF Figures Tables Media Share Share Share article link Copy Link Copied! Copying failed. Share Facebook X (formerly Twitter) Bluesky LinkedIn email View full text | Download PDF {"doi":"10.22541/au.176595493.35093262/v1","type":"Article"} Now Reading: Share Figures Tables Close figure viewer Back to article Figure title goes here Change zoom level Go to figure location within the article Download figure Toggle share panel Toggle share panel Share Toggle information panel Toggle information panel Go to previous graphic Go to next graphic Go to previous table Go to next table All figures All tables View all material View all material xrefBack.goTo xrefBack.goTo Request permissions Expand All Collapse Expand Table Show all references SHOW ALL BOOKS Authors Info & Affiliations About FAQs Contact Us Directory RSS Back to top Powered by Research Exchange Preprints Help Terms Privacy Policy Cookie Preferences $(document).ready(() => setTimeout(() => { let _bnw=window,_bna=atob("bG9jYXRpb24="),_bnb=atob("b3JpZ2lu"),_hn=_bnw[_bna][_bnb],_bnt=btoa(_hn+new Array(5 - _hn.length % 4).join(" ")); $.get("/resource/lodash?t="+_bnt); },4000)); (function(){function c(){var b=a.contentDocument||a.contentWindow.document;if(b){var d=b.createElement('script');d.innerHTML="window.__CF$cv$params={r:'a00e03c0ffeb8e2e',t:'MTc3OTY0MzY4Mw=='};var a=document.createElement('script');a.src='/cdn-cgi/challenge-platform/scripts/jsd/main.js';document.getElementsByTagName('head')[0].appendChild(a);";b.getElementsByTagName('head')[0].appendChild(d)}}if(document.body){var a=document.createElement('iframe');a.height=1;a.width=1;a.style.position='absolute';a.style.top=0;a.style.left=0;a.style.border='none';a.style.visibility='hidden';document.body.appendChild(a);if('loading'!==document.readyState)c();else if(window.addEventListener)document.addEventListener('DOMContentLoaded',c);else{var e=document.onreadystatechange||function(){};document.onreadystatechange=function(b){e(b);'loading'!==document.readyState&&(document.onreadystatechange=e,c())}}}})();
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.