Performance-weighted ensemble learning for detecting patients with FTD and ALS from short reading tasks

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract We developed a novel machine learning model named performance-weighted ensemble learning (PWEL) to detect frontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS) using speech from word and sentence reading tasks. Overall, 197 participants (30 with FTD, 74 with ALS, and 93 healthy controls) were enrolled. The brief oral reading tasks consisted of 16 materials adapted from the Western Aphasia Battery and took 1 min to complete. For each task, 405 speech features (i.e., 384 acoustic, 17 linguistic, and 4 temporal) were extracted. Five-fold cross-validation highlighted acoustic features—especially MFCCs—as key discriminators of FTD and ALS from healthy controls. PWEL, which included under-bagging and adaptive task selection, achieved an area under the curve of 0.840, a sensitivity of 0.828, and an overall accuracy of 0.792. The sensitivity for FTD subtypes was 0.700-0.933, and that for ALS reached 0.812. Our proposed model surpassed both a simple ensemble learning model and task-wise classification in overall sensitivity, accuracy, and macro-F1. Its robust generalizability was demonstrated by consistent performance across stratified cross-validation folds and three recording sites. In conclusion, PWEL is practical, scalable, and generalizable across multiple sites, making it a promising tool for detecting neurodegenerative disorders.
Full text 185,078 characters · extracted from preprint-html · click to expand
Performance-weighted ensemble learning for detecting patients with FTD and ALS from short reading tasks | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Performance-weighted ensemble learning for detecting patients with FTD and ALS from short reading tasks Yuki Ito, Reiko Ohdake, Shohei Kato, Michihito Masuda, Maki Suzuki, and 12 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7917087/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract We developed a novel machine learning model named performance-weighted ensemble learning (PWEL) to detect frontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS) using speech from word and sentence reading tasks. Overall, 197 participants (30 with FTD, 74 with ALS, and 93 healthy controls) were enrolled. The brief oral reading tasks consisted of 16 materials adapted from the Western Aphasia Battery and took 1 min to complete. For each task, 405 speech features (i.e., 384 acoustic, 17 linguistic, and 4 temporal) were extracted. Five-fold cross-validation highlighted acoustic features—especially MFCCs—as key discriminators of FTD and ALS from healthy controls. PWEL, which included under-bagging and adaptive task selection, achieved an area under the curve of 0.840, a sensitivity of 0.828, and an overall accuracy of 0.792. The sensitivity for FTD subtypes was 0.700-0.933, and that for ALS reached 0.812. Our proposed model surpassed both a simple ensemble learning model and task-wise classification in overall sensitivity, accuracy, and macro-F1. Its robust generalizability was demonstrated by consistent performance across stratified cross-validation folds and three recording sites. In conclusion, PWEL is practical, scalable, and generalizable across multiple sites, making it a promising tool for detecting neurodegenerative disorders. Biological sciences/Computational biology and bioinformatics Health sciences/Neurology Biological sciences/Neuroscience Machine learning Frontotemporal dementia Amyotrophic lateral sclerosis Diagnostic screening tool Reading task Speech features Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Frontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS) belong to the same spectrum of neurodegenerative diseases and present with overlapping clinical manifestations and neuropathological features 1 . FTD encompasses a heterogeneous group of clinical syndromes characterized by progressive behavioral changes, executive dysfunction, or language impairments. The three main FTD subtypes are behavioral variant frontotemporal dementia (bvFTD) and two primary progressive aphasia (PPA) subtypes, namely, progressive non-fluent aphasia (PNFA) and semantic dementia (SD). Notably, 10–15% of individuals with FTD fulfill the diagnostic criteria for ALS 1 . Furthermore, approximately half of patients with ALS exhibit varying degrees of cognitive and behavioral impairments, similar to those seen in FTD 2 . There is currently no effective treatment to halt or delay disease progression in FTD. Its prevalence in patients with early-onset dementia (age < 65 years) is comparable to that of Alzheimer’s disease, affecting approximately 15 per 100,000 individuals 3 . Meanwhile, ALS predominantly occurs in adults aged 45–75 years, with annual prevalence rates of 4–8 per 100,000 people 4 . ALS results in substantial social and personal burdens to both patients and their caregivers. Despite improvements in their clinical understanding, both FTD and ALS can be difficult to accurately diagnose, and thus delayed or incorrect diagnoses are common. In one study, 60% of patients initially diagnosed with bvFTD in community settings did not meet the diagnostic criteria when evaluated by specialists 5 . This discrepancy underscores the complexity of FTD presentations and highlights the need for specialist involvement in the diagnostic workup. In ALS, certain clinical features, such as concomitant FTD, bulbar onset, rapid functional decline, significant weight loss, older age at onset, and reduced forced vital capacity, are associated with shorter survival 4 . Moreover, ALS is a clinically heterogeneous disease in which frontotemporal lobe involvement may be undetected, unless specifically assessed. Collectively, these challenges highlight the urgent need for a low-cost, non-invasive, quantitative, and time-efficient screening tool to identify patients with FTD and with ALS and reduce diagnostic delays. Accumulating evidence suggests that speech and language changes arise early in FTD, and thus, they may be noninvasive parameters for screening 6 . However, although prior machine learning research in this area focused primarily on linguistic features derived from spontaneous speech, these approaches often require lengthy and sometimes burdensome data collection protocols, as well as time-consuming manual transcription by trained annotators 7 , 8 . Clinically, patients with FTD or ALS may present with psychiatric or behavioral symptoms (e.g., apathy and disinhibition), speech stereotypies, and reduced verbal output, all of which can complicate lengthy interviews or spontaneous speech tasks. Similarly, specific variants such as patients with SD may feature pronounced anomia, repetitive speech, and behavioral symptoms that hamper extensive data collection 9 . Although existing methods are promising, they may be impractical for rapid screening or large-scale implementation. To overcome these challenges, recent investigations have shifted towards analyzing abnormal acoustic properties in the speech of patients with FTD and ALS 10 , 11 , leading to preliminary machine learning models that capture paralinguistic changes associated with emotional and cognitive deficits 12 . Building on these developments, we propose a machine learning approach focused on acoustic features extracted from a brief reading task to discriminate patients with FTD and with ALS from healthy controls (HCs). Our method requires only approximately 1 min of data collection: participants are asked to read single words or short sentences displayed on a screen without further prompts or examiner intervention. In addition to linguistic and temporal measurements, we emphasize the acoustic and paralinguistic features of the recorded speech. Furthermore, to facilitate scalability and reduce the examiner burden, our pipeline is designed for automated pre-processing and analysis, potentially allowing at-home recordings and adaptation to multiple languages. In this paper, we present our machine-learning framework for distinguishing patients with FTD and with ALS from healthy individuals based on acoustic, linguistic, and temporal features derived from a simple reading task. Results Diagnostic accuracy of performance-weighted ensemble learning Our proposed model, performance-weighted ensemble learning (PWEL), showed superior performance compared to a simple ensemble learning model and task-wise classification, with an area under the curve (AUC) of 0.840 (95% confidence interval [CI]: 0.786–0.891); sensitivity, 0.828ฏ (95% CI: 0.750–0.894); specificity, 0.751 (95% CI: 0.667–0.828); accuracy, 0.792 (95% CI: 0.731–0.843); and macro-F1, 0.788 (95% CI: 0.730–0.842) (Table 1 , Fig. 1 ). Figure 1 shows the receiver operating characteristic (ROC) curves for 5-fold cross validation (CV). The AUC values for each fold ranged from 0.799 to 0.879. Figure 2 shows the PWEL confusion matrix combining the results of 5-fold CV. The optimal values of \(\:{k}^{*}\) for each fold were [8, 1, 3, 3, and 12]. The optimal classification thresholds \(\:\theta\) determined for each fold were [0.500, 0.500, 0.333, 0.500, and 0.292]. For weak learners, support vector machine (SVM) showed higher accuracy and macro-F1 values than did LightGBM and random forest. Therefore, we adopted SVM as the inner/outer weak learner in PWEL. Table 1 Diagnostic accuracy in PWEL (FTD and ALS vs. HC). Sensitivity for each disease Spec Acc macro-F1 FTD and ALS total PNFA bvFTD SD ALS Ensemble Learning PWEL 0.828 0.800 0.700 0.933 0.812 0.751 0.792 0.788 Majority voting 0.759 1.000 0.700 0.867 0.715 0.784 0.772 0.770 Task-wise classification SVM 0.648 0.806 0.550 0.775 0.615 0.718 0.681 0.679 RandomForest 0.681 0.808 0.575 0.841 0.631 0.659 0.671 0.667 LightGBM 0.704 0.881 0.550 0.846 0.674 0.654 0.680 0.676 The task-wise classification indicates the mean classification performance across the 16 tasks. Highest records are reported in bold type. PWEL, Performance-Weighted Ensemble Learning; SVM, Support Vector Machine; Spec, Specificity; Acc, Accuracy. Ablation study Table 2 presents the comparison results for the four models, with and without under-bagging and adaptive task selection. Under-bagging or adaptive task selection improved specificity in the four models, and the combination of under-bagging and adaptive task selection applied to the PWEL model improved accuracy and macro-F1. Table 2 Comparison results for the four models combined with and without under-bagging and ATS. under-bagging ATS Sensitivity for each disease Spec Acc macro-F1 FTD an ALS total PNFA bvFTD SD ALS ✓ ✓ 0.828 0.800 0.700 0.933 0.812 0.751 0.792 0.788 ✓ 0.759 1.000 0.700 0.867 0.715 0.784 0.772 0.770 ✓ 0.817 1.000 0.600 0.867 0.810 0.719 0.771 0.769 0.865 1.000 0.700 0.867 0.864 0.677 0.776 0.772 Highest records are reported in bold type. ATS, Adaptive Task Selection; Spec, Specificity; Acc, Accuracy. Accuracy comparison among the three facilities A total of 140, 28, and 29 participants were recruited from one facility each in Nagoya, Toyoake, and Osaka, respectively (Table S1 ). There was no significant difference in the accuracy rate of PWEL among these facilities (P = 0.088). Speech features of reading tasks selected through 5-fold CV Among the tasks selected through adaptive task selection, three tasks (task 3, 8, and 15) were commonly selected across 4 of 5 folds. An examination of the top 10 features for each of the three tasks revealed that acoustic features were the most prevalent (96.7%), followed by temporal features (3.3%). No linguistic features were included in the top 10 features. MFCC—representing the vocal tract transfer function—was the most common acoustic feature, accounting for 66.7% (e.g., mfcc_sma_de[1]_maxPos, mfcc_sma[1]_min, mfcc_sma[4]_minPos). It was followed by Pcm zcr as a proxy for pitch at 13.3% (e.g., zcr_sma_de_kurtosis, zcr_sma_de_maxPos), RMS energy as an indicator of volume at 10.0% (e.g., RMSenergy_sma_linregc1, RMSenergy_sma_min), and voice probability as a measure of voicing at 6.7% (e.g., voiceProb_sma_de_skewness). Discrimination performance by task and disease subtype The discrimination performance by task and disease subtype is presented in Table S2. In patients with PNFA, the sensitivity to multiple word and sentence tasks reached 1.000. In patients with SD, the sensitivity to sentence tasks reached 1.000. Meanwhile, sensitivity was lower than 0.900 in all tasks in patients with bvFTD and ALS. The sensitivity for each task varied in patients with ALS, with the highest sensitivity being for the longest-sentence task (task No.15) (0.699), followed by that for three-letter words (No.3) (0.690). Therefore, task order (i.e., fatigue associated with performing tasks) did not affect diagnostic accuracy. In all study participants, accuracy and macro-F1 were the highest for the longest-sentence task (0.751 and 0.747, respectively) and lowest for the non-word task (0.594 and 0.591, respectively). Clinical characteristics of misclassified participants There was no significant difference in demographic variables, cognitive scores, or language functions between the HC error group (HCs classified to have FTD and ALS) and the HC correct group (Table S3). Compared to the FTD and ALS correct groups, the FTD and ALS error groups (classified as HCs) were younger (P = 0.042), had higher cognitive function (MMSE: P = 0.047), and had milder ALS Functional Rating Scale-Revised (ALSFRS-R) bulbar symptoms (P = 0.018). The 18 patients with misclassification in the disease group (false negatives in Fig. 2 ) included 14 patients with ALS, 2 patients with bvFTD, 1 patient with SD, and 1 patient with PNFA. The ALS group had the highest misclassification rate among the disease groups, with 18.9% of patients classified as having HCs. Discussion In this study, PWEL, which included under-bagging and adaptive task selection, achieved good discrimination ability for identifying patients with FTD and ALS. It showed very high sensitivity for FTD subtypes and acceptable sensitivity for ALS. Further, PWEL demonstrated good generalization performance with high accuracy across 5-fold CV and consistent accuracy across recording sites. The strong performance of PWEL can be attributed to two key components of its design: (i) the correction of class imbalance through under-bagging and (ii) the use of ensemble learning with adaptive task weighting. First, under-bagging addresses the inherent class imbalance in clinical datasets, particularly the smaller number of patients with FTD than that of patients with ALS and HCs. Models trained on imbalanced data tend to favor the majority class, which can degrade the sensitivity for underrepresented groups. By implementing under-bagging, combining undersampling with bagging, PWEL ensures that each training subset maintains class balance, enabling better detection of rare subtypes such as FTD. Second, PWEL employs ensemble learning by integrating the predictions from multiple classifiers trained on different reading tasks. This approach leverages the strengths of individual classifiers and outperforms single learners in various domains 13 . Through adaptive task selection, the model dynamically identifies and weighs tasks that contribute the most to the classification accuracy. By combining task-level predictions with learned weights, PWEL enhances the overall robustness and sensitivity of the findings. As shown by the results of our ablation study, the combination of under-bagging and adaptive task selection yielded the highest accuracy and macro-F1 scores. Collectively, these design elements enable PWEL to achieve superior discriminative performance in comparison to conventional machine learning approaches, particularly in the context of clinically heterogeneous and imbalanced datasets. Among the 405 speech features used in this study, 384 were acoustic features derived from the INTERSPEECH 2009 Emotion Challenge (IS09) feature set. Numerous studies have adopted IS09 subsets for analyzing speech in affective and neurological contexts 14 . However, to our best knowledge, this is the first study to apply a machine learning model that primarily leverages acoustic features to detect FTD and ALS in a relatively large cohort that involves a large number of patients with ALS. This emphasis on acoustic features may be especially relevant in ALS, where bulbar symptoms contribute to altered speech. In our analysis, patients with ALS who were misclassified as being HCs exhibited significantly milder bulbar symptoms, as measured using the ALSFRS-R subscores. This suggests that the model may detect bulbar motor involvement. However, PWEL had higher sensitivity for SD (0.933) as an FTD subtype than for ALS (0.812). Previous studies have suggested that the fundamental frequency (f0) range is associated with bulbar motor dysfunction in ALS, whereas increased pause duration and speech disruption are associated with cognitive deficits 10 . Moreover, patients with progressive apraxia of speech may exhibit articulatory distortions similar to those with motor neuron disease 15 . These findings support the premise that PWEL can capture both dysarthria and FTD-like speech anomalies in ALS, thereby contributing to its diagnostic utility. PWEL also demonstrated comparable performance to models that relied on linguistic features while offering notable practical advantages. Prior machine learning approaches for FTD and PPA have employed natural language processing on free speech (e.g., narrative descriptions or daily life conversations), reporting AUCs ranging from 0.700 to 0.900 and accuracies ranging from 80% to 90% 7,8,16–18 . Although rich in diagnostic information, such approaches often require labor-intensive pre-processing (e.g., correction of phonological errors, agrammatism, and fillers) that is typically performed manually. Consequently, the results may be influenced by the accuracy and consistency of the transcription pipelines. In contrast, PWEL automatically extracts acoustic features from a short, structured reading task, thus minimizing examiner bias and transcription variability. This approach not only reduces the workload, but also enhances scalability and objectivity in real-world screening scenarios. Although the IS09 was originally designed for emotion classification, its use here is justified by overlaps in speech alterations observed in both affective and neurological disorders. Particularly, prosodic and spectral features are effective in detecting dysarthria and motor speech anomalies. Our choice to prioritize acoustic features over linguistic features reflects a deliberate tradeoff between reducing the need for manual transcription and increasing model robustness in patients with low verbal output. Nevertheless, future studies are needed to systematically evaluate and optimize the feature sets for neurological disease classification. We employed a short reading task in PWEL that lasted approximately 1 min, using words and sentences from the WAB. We hypothesized that this approach would enhance task feasibility for patients with FTD and ALS while also improving the discriminative power of the machine learning model based on acoustic features. A short assessment duration is particularly important for patients with ALS who often experience fatigue due to motor dysfunction. The use of reading aloud tasks has several clinical advantages for FTD subtypes. Patients with PNFA typically exhibit effortful and halted speech due to articulatory difficulties; however, they often have mild and phonemic reading problems 19 . Meanwhile, patients with SD generally retain fluent speech, and, unless affected by surface dyslexia, their ability to read aloud is comparable to that of healthy individuals 20 . In patients with bvFTD, compulsive reading, a behavior associated with environmental dependency syndrome, may paradoxically increase engagement and feasibility, even in those with disinhibition or inertia. Moreover, as the tasks are presented visually, issues related to working memory or hearing impairments do not interfere with data collection. The enhanced classification performance observed with PWEL may be attributable to the emphasis on acoustic properties. By constraining the verbal content to standardized WAB items, we reduced linguistic variability and amplified interindividual differences in acoustic patterns. This may have enabled a more precise capture of paralinguistic features, such as emotional prosody, vocal effort, and articulation. Language independence is a notable advantage for international application, particularly in settings where linguistic resources and expert transcribers are limited. Moreover, reading tasks adopted from the WAB encompass a range of linguistic complexities that can indirectly reflect cognitive linguistic abilities. These include variations in word frequency, syllable count, sentence length, and grammatical structure 21 . Importantly, the use of a structured reading task minimizes the burden on both participants and examiners and facilitates standardized, reproducible data collection across settings. Furthermore, it is noteworthy that the PWEL reading task was developed based on the WAB. The WAB has been translated into at least 33 languages since its original development by Kertesz et al. 22 and is now widely used worldwide as a standardized tool for aphasia assessment. It has been proven effective in evaluating aphasia not only due to cerebrovascular disorders, but also due to neurodegenerative diseases 22 . Given that the WAB has been linguistically and psychometrically validated across multiple languages, it may be possible to develop equivalent reading tasks with levels of difficulty and linguistic complexity matched to our study. However, cross-linguistic applicability is a future direction to be empirically evaluated. In this study, we discuss the generalization performance of PWEL from the following two perspectives: (i) the generalization performance of AUC scores across each fold in 5-fold CV and (ii) the percentage of correct responses by each institution. First, the mean AUC across the 5 folds was 0.840, with fold-specific values of 0.836, 0.832, 0.799, 0.856, and 0.879. Although 1 fold showed an AUC slightly below 0.800, the majority of folds achieved values close to or above this threshold, and the overall performance remained within an acceptable range for diagnostic applications 23 , 24 . Furthermore, the 95% CI width of 0.105 is relatively narrow, likely representing a precise estimate with limited sampling variation. Such an interval suggests that the PWEL model’s performance metric is more reliable for diagnosis in the sense that it has much less variability and edge cases. Second, the generalizability of our approach was supported by consistent classification performance across the three participating institutions. Data were collected by three examiners in various clinical settings, including outpatient departments, inpatient wards, and affiliated facilities. As such, differences in background noise across recording locations are inevitable and can affect model training and performance. To mitigate this, we applied the spectral subtraction method 25 as a preprocessing step to reduce the background noise in all audio recordings. This technique helped standardize the acoustic quality of the input data and may have contributed to the consistency in model performance across sites. Overall, the findings support the generalizability and robustness of the proposed method. The development of robust, low-cost, and scalable tools, such as PWEL, has important implications for the early screening of ALS and FTD. Such tools can be deployed not only in clinical environments, but also in home-based or telemedicine settings, facilitating timely referral to specialists and reducing diagnostic delays. This study had limitations. First, the sample size, particularly for the FTD subtypes, was relatively small, limiting the statistical power and increasing the risk of inflated performance metrics in the subgroup analyses. To mitigate the risk of overfitting, we employed 5-fold CV and ablation studies as 5-fold CV is considered suitable for our moderate-sized datasets (n = 197), offering a practical balance between bias and variance 26 . Importantly, AUC values were above 0.800 in 4 of the 5 folds, suggesting that the model maintained robustness even under limited data conditions. Second, although the acoustic features were derived from a validated and widely used feature set (IS09), the assumption that they are optimal for disease classification remains to be tested empirically. Future research should explore whether alternative or task-specific features can improve discriminative performance. Third, missing linguistic features were imputed as zero in patients in whom automatic speech recognition (ASR) failed. However, ASR failure occurred in only 2/197 patients (1.0%), and the affected patients were evenly distributed across subtypes (1 patient with bvFTD and 1 patient with SD). To preserve input dimensionality, we imputed zero values, which arguably reflected severe speech impairment. Speech impairment is a potentially informative signal rather than random noise. Given the negligible frequency and distribution, this handling was unlikely to introduce systematic bias. However, more refined imputation and error-handling approaches must be developed. Finally, the high sensitivity observed in some FTD subtypes of each reading task (e.g., 1.000) may reflect the small number of test participants and should thus be interpreted cautiously. Although a separate external validation set was not used, speech data were collected from three distinct institutions across various clinical settings (inpatient, outpatient, and affiliated clinics). The consistency in the classification accuracy across these centers supports preliminary generalizability. However, the classification accuracy of PWEL could have been further improved by tailoring the noise-reduction strategies to the specific environmental characteristics of each facility. External validation using geographically and linguistically distinct datasets is planned in future studies. In conclusion, PWEL can be useful in distinguishing patients with FTD and ALS from healthy individuals through simply reading words and sentences for 1 min. Furthermore, noise reduction and the use of WAB tasks increase feasibility across different locations and languages, and automated machine learning with acoustic features improves time efficiency for the identification of patients with FTD and ALS. Collectively, our findings provide evidence that PWEL is a practical, scalable, and a promising screening tool adaptable to multiple sites that can be used for the clinical detection of neurodegenerative disorders. Methods Study design and participants This prospective study was approved by the Ethics Committees at Fujita Health University (approval number: HM19-216), Nagoya University (approval number: 2018 − 0243), The University of Osaka (approval number: 19183), and Nagoya Institute of Technology (approval number: 2019-011) and conformed to the Ethical Guidelines for Medical and Health Research Involving Human Subjects, as endorsed by the Japanese government. All patients provided written informed consent after the procedure was explained. Patients diagnosed with FTD and ALS at the Department of Neurology at Fujita Health University and Nagoya University in accordance with the established diagnostic criteria for FTD 27 – 29 and ALS 30 were enrolled from November 2018 to February 2023. Patients diagnosed with FTD at the Department of Psychiatry, Osaka University were also recruited from December 2019 to February 2023. The exclusion criteria were as follows: (i) a history of other neurological diseases, stroke, or traumatic brain injury; (ii) unintelligible speech, defined as an ALSFRS-R speech subscore ≤ 2; (iii) Mini-Mental State Examination (MMSE) score < 26 31 and Japanese version of Addenbrooke’s Cognitive Examination-Revised (ACE-R) score < 89 32 in HCs; and (iv) severe hearing and visual impairments that could interfere with task performance. A total of 104 patients (30 patients with FTD and 74 patients with ALS) were included. Among the 30 patients with FTD, 15, 7, and 8 patients had SD, PNFA, and bvFTD, respectively. We also included 93 age-matched HCs from the three institutions. All participants (n = 197) were native Japanese speakers. The participants were matched by age but not by sex or educational attainment (Table 3 ). Consequently, age and sex were matched within each fold of the training and test datasets (Table S4), and year of education was considered in the analysis of the discrimination outcomes. Table 3 Demographics, clinical characteristics, and cognitive function of the participants. FTD and ALS HC p - value Total, n = 104 PNFA, n = 7 bvFTD, n = 8 SD, n = 15 ALS, n = 74 n = 93 age 66.4 (10.2) 73.1 (6.7) 62.5 (9.0) 69.6 (7.3) 65.6 (10.8) 67.9 (8.7) 0.425 gender, M / F 54 / 50 3/4 3/5 6/9 42/ 32 35 / 58 0.044 education 12.2 (2.4) 12.0 (1.7) 14.5 (3.0) 12.5 (2.2) 11.9 (2.3) 13.4 (2.1) < 0.001 MMSE 1 24.3 (5.5) 23.0 (4.0) 21.4 (5.5) 16.5 (7.0) 26.3 (3.2) 28.8 (1.1) < 0.0001 ACE-R 1 76.1 (22.6) 70.9 (13.6) 61.9 (18.2) 35.4 (14.8) 86.5 (12.4) 95.7 (3.2) < 0.0001 WAB, AQ 1 86.5 (14.5) 77.2 (10.4) 79.5 (11.1) 60.2 (15.1) 93.6 (4.5) 97.1 (2.4) < 0.0001 Spontaneous Speech (/20) 17.0 (2.6) 13.9 (1.6) 15.1 (3.6) 13.5 (2.7) 18.2 (1.2) 19.3 (0.8) < 0.0001 Auditory comprehension (/10) 1 8.9 (1.5) 8.7 (0.9) 7.7 (1.4) 6.6 (2.3) 9.6 (0.5) 9.8 (0.3) < 0.0001 Repetition (/10) 9.3 (1.4) 7.8 (2.6) 9.2 (1.3) 7.3 (1.9) 9.8 (0.4) 9.9 (0.2) < 0.0001 Naming (/10) 1 8.1 (2.6) 8.3 (1.0) 7.7 (1.1) 2.6 (1.9) 9.3 (0.8) 9.6 (0.4) < 0.0001 disease duration 1 2.1 (2.1) 2.2 (1.2) 2.5 (1.9) 6.0 (2.8) 1.3 (1.1) ALSFRS, total 1 40.0 (3.8) bulbar 1–3 (/12) 11.0 (1.5) motor 4–9 (/24) 17.5 (3.4) respiratory 10–12 (/12) 11.5 (1.0) Data are mean (SD). FTD, frontotemporal dementia; ALS, amyotrophic lateral sclerosis; HC, healthy control; bvFTD, behavioral variant of frontotemporal dementia; PNFA, progressive non fluent aphasia; SD, semantic dementia; ALSFRS-R, ALS functional rating scale-revised; MMSE, Mini-Mental State Examination; ACE-R, Addenbrooke's Cognitive Examination-Revised; WAB, Western Aphasia Battery; AQ, Aphasia Quotient. the Mann-Whitney U test and the chi-square tests for demographic comparisons between the two independent study groups. 1 The number of participants was different (for disease duration: SD, n = 13; for ALSFRS: ALS, n = 73; for MMSE, ACE-R: ALS, n = 73; for WAB AQ, Auditory comprehension: ALS, n = 71; for WAB Naming: ALS, n = 73). Clinical data Language function and global cognition were evaluated through a series of neuropsychological assessments (Table 3 ). Language ability was assessed using the WAB 21 , 33 , while global cognitive function was measured using the MMSE and ACE-R. The AQ, a composite score based on spontaneous speech, auditory comprehension, repetition, and naming, was derived from the WAB. Disease duration was assessed in patients with both FTD and ALS. Functional disability at the time of testing was assessed using the ALSFRS-R in patients with ALS. To further characterize ALS symptoms, the ALSFRS-R was divided into three subscores according to previous studies 34 , 35 : bulbar subscore (maximum score = 12, items 1–3), motor subscore (maximum score = 24, items 4–9), and respiratory subscore (maximum = 12, items 10–12). In relation to the outputs of our machine learning model, the associations between classification performance and clinical variables, such as motor function and cognitive scores, were also examined. Reading tasks We applied the words and sentences of the repetition task from the WAB to our reading task. The WAB repetition task included various levels of language ability, such as frequent words with an increasing number of syllables, compound words, numbers, number-word combinations, common sentences, rare sentences, and sentences of increasing length, as well as grammatical complexity. Meanwhile, the WAB reading tasks were designed to assess reading comprehension, and the participants were asked to read aloud and act on the text. Therefore, we hypothesized that the content of the WAB repetition task would be suitable for assessing language ability in reading texts. Finally, one compound non-word was added to examine the influence of the semantic properties of language, and the reading task in this study consisted of 10 words and 6 sentences. Given that the Japanese language has three characters (i.e., Hiragana, Katakana, and Kanji), we added the Kana characters above Kanji characters to account for the differences between them. The participants were asked to read 16 tasks that were presented as a single line on a screen while their speech was digitally recorded. Voice recordings were collected using an AT9921 with a unidirectional microphone (Audio-Technica Corporation, Japan) and a PCM-A10 (SONY, Japan) placed approximately 30–40 cm from the mouth. Voice samples were recorded in a linear PCM format (.wav) at a sampling rate of 44.1 kHz, with a 16-bit sample size. Proposed method: PWEL The system identified the optimal combination of reading tasks using PWEL. Figure 3 illustrates an overview of the proposed system, from preprocessing to output. Preprocessing: noise reduction Denoising was applied to the audio data as a preprocessing step. The recordings were collected at multiple medical institutions. As background noise varied depending on the recording environment and could adversely affect model training, we applied the spectral subtraction method 36 . In this approach, the observed signal is modeled as $$\:x\left(t\right)=s\left(t\right)+n\left(t\right)$$ , where \(\:x\left(t\right)\) as the recorded audio at an arbitrary time \(\:t\) ; \(\:s\left(t\right)\) , the underlying clean speech signal; and \(\:n\left(t\right)\) is background noise \(\:.\) After applying the discrete Fourier transform (DFT), the relationship becomes $$\:X\left(\omega\:\right)=S\left(\omega\:\right)+N\left(\omega\:\right),$$ where \(\:X\left(\omega\:\right)\) , \(\:S\left(\omega\:\right)\) , and \(\:N\:\left(\omega\:\right)\) denote the spectra of the observed signal, clean speech, and noise, respectively. The noise power spectrum \(\:N\:\left(\omega\:\right)\) was estimated from manually identified non-speech intervals that contained only background noise and no vocal activity. The denoised spectrum was then obtained by subtracting this estimate from \(\:N\:\left(\omega\:\right)\) , and the time-domain signal \(\:\widehat{s}\left(t\right)\:\) was reconstructed using the inverse DFT. Importantly, non-speech intervals were not removed from the recordings; they were used solely to estimate the noise statistics, thereby preserving pause-related temporal features. Feature extraction Definition of speech features Speech features were extracted from voice recordings of 16 reading tasks. A total of 405 features, including 384 acoustic, 17 linguistic, and 4 temporal features, were extracted from the speech samples. Each feature type is detailed below. Acoustic features The acoustic features used in this study were selected from the IS09 37 set, a feature set widely used in affective computing and neurodegenerative disease studies 14 . The IS09 feature set includes prosodic, spectral, and voice quality characteristics that reflect emotional states and motor impairments in speech. Although originally designed for emotion classification, it has also been applied in multiple neurological domains owing to its robustness and broad coverage of speech signal properties. The low-level descriptor features and statistical measures are shown in Tables 4 , respectively. All 384 features were extracted by calculating 12 types of statistical measures for each of the 16 types of basic features related to vocal tract characteristics, volume, pitch, and voice probability. Feature extraction was performed using openSMILE software version 2.3.0 (audEERING GmbH, Gilching, Germany) 38 . Table 4 LLD and the 12 statistical measures calculated from LLD. Acoustic Property LLD Description Volume RMSenergy Root-mean-square signal frame energy Vocal Tract Transfer Function MFCC 1–12 Mel-Frequency cepstral coefficients 1–12 Pitch Pcm zcr Zero-crossing rate of time signal (frame-based) Voicing Probability Voice Prob The voicing probability computed from the ACF Pitch F0 The fundamental frequency computed from the Cepstrum *_de The first derivative of each of the above features functionals Description max The maximum value of the contour min The minimum value of the contour range max-min maxPos The absolute position of the maximum value (in frames) minPos The absolute position of the minimum value (in frames) amean The arithmetic mean of the contour linregc1 The slope (m) of a linear approximation of the contour linregc2 The offset (t) of a linear approximation of the contour linregerrO The quadratic error computed as the difference of the linear approximation and the actual contour stddev The standard deviation of the values in the contour skewness The skewness (3rd order moment) kurtosis The kurtosis (4th order moment) LLD, Low-Level Descriptors Linguistic features Linguistic features were extracted from the participants’ speech by first transcribing audio into text using the Azure Speech-to-Text system, a cloud-based ASR service provided by Microsoft Azure AI. In rare instances where ASR failed owing to incapable vocal intensity, all linguistic features were set to zero to maintain consistent input dimensionality across participants. Only 2 (1 patient with bvFTD and 1 patient with SD) of the 197 patients (1.0%) were affected by this imputation. Notably, no ASR failures occurred among participants with ALS. The transcribed text was further processed using MeCab 39 . Briefly, MeCab is an open-source morphological analysis engine tool for Japanese morphological analysis. Based on the morphological output, 17 linguistic features were extracted. Each feature is described in Table S5. The total number of words was denoted by \(\:N\) , and the number of unique words was denoted by \(\:V\left(N\right)\) . The type-token ratio (TTR) and Simpson’s D value 40 were also calculated to assess lexical diversity, as follows: $$\:TTR=\:\frac{V\left(N\right)}{N}.$$ $$\:D=\:\sum\:_{m\:=\:1}^{N}V\left(m,N\right)\frac{m\left(m-1\right)}{N\left(N-1\right)}$$ , where \(\:V\:(m,\:N)\) is the number of unique words that appear m times in a text of length \(\:N\) . Simpson’s D value is high when the same word is repeatedly used, suggesting low lexical diversity. Although the content of the reading task was predetermined, responses varied and included distortion or substitution of sounds due to effortful speech, mispronunciations of words, and repetitive speech. Therefore, we adopted an index to evaluate the lexical richness. Temporal features The temporal features were measured manually. Table S6 lists the temporal features and their corresponding descriptions. The reaction time (i.e., the time between the monitor display the duration from speech onset to speech offset for and the start of the response) and speech time (i.e., the duration from speech onset to speech offset for each of the 16 tasks) were extracted. In addition, the speed and linguistic features \(\:N\) and \(\:V\left(N\right)\) were combined to define the two types of features related to speech speed. Feature selection Considering that a total of 405 features were extracted, which was greater than the number of training data points, overlearning could occur. Therefore, feature selection was performed using a combination of filter and wrapper methods before training the model. The filter method selected characteristics using statistics such as variance and correlation coefficients 41 . The wrapper method used a subset of features to evaluate the performance of the model and select optimal features 42 . Although the filter method could select features in a short time, it could not consider the relationships between the features. In contrast, the wrapper method considered the relationships between features. Therefore, this study proposed a feature selection method that adapted the wrapper method to the features selected by the filter method. First, the filter method was applied to remove features that exhibited no variance across samples and those that showed high pairwise correlations (correlation coefficient > 0.80). Second, in the wrapper method, a forward stepwise procedure guided by the Akaike Information Criterion was employed to sequentially add features and identify those contributing to the classification. The number of selected features was limited to less than the number of training data points (maximum of 157 or 158) to prevent model overtraining. PWEL model PWEL is an ensemble-learning method that optimizes task combinations and weighting based on the classification performance of each disease subtype. It considers the continuity among subtypes. In clinical settings, the difficulty of data collection varies according to subtype. For rare subtypes, data scarcity makes it difficult to build machine learning models. Therefore, it is preferable to treat the subtypes as a single group and compare the classification performance between patients and HCs. PWEL evaluates the classification performance of each subtype. It assigns higher weights to weak learners with better performances. This approach effectively utilizes information from rare subtypes. The number of subtypes \(\:d\) and speech task types \(\:t\) can be adjusted. Weights are computed based on \(\:d\) and \(\:t\) . PWEL is composed of two methods: (i) under-bagging, which combines undersampling and bagging, and (ii) adaptive task selection, which consists of an inner weak learner, selecting tasks, and weighted ensemble learning (where the number of selected tasks is denoted by \(\:T\) ). As a weak learner, a binary classifier (e.g., SVM, decision tree, and gradient boosting) was considered. The model with the highest discrimination accuracy was adopted for the adaptive task selection. Addressing unbalanced data by under-bagging There were imbalances in the clinical data between the disease and non-disease groups. Directly inputting such imbalanced data into a machine learning model makes it difficult for the model to adequately learn from rare cases and correctly identify them. A recent systematic review on imbalanced medical datasets 43 shows that undersampling is frequently adopted in medical research because it avoids the risk of generating unrealistic synthetic data inherent to oversampling. Nevertheless, as undersampling inevitably discards part of the majority class, it is recommended to combine it with ensemble methods (e.g., bagging) to mitigate information loss and improve robustness. Based on this rationale, we applied an under-bagging strategy to balance the participant groups: the ALS group was undersampled to match the number of FTD samples, and the HC group was undersampled to match the combined total of the FTD and ALS groups. To assess the robustness of our analysis with respect to sample size, we employed both 5-fold stratified CV. In 5-fold CV, each training fold involved 157 participants (59 patients with ALS, 6 patients with PNFA, 6 patients with bvFTD, 12 patients with SD, and 74 HCs), whereas the corresponding test folds involved 40 participants (15 patients with ALS, 1 patient with PNFA, 2 patients with bvFTD, 3 patients with SD, and 19 HCs). Stratified 5-fold CV was conducted using mutually exclusive sample partitions across folds to mitigate data leakage and enhance the generalizability of the model. The slight variation in fold size reflects the inherent consequences of stratified partitioning with limited class frequencies and does not introduce systematic bias into model evaluation. Adaptive task selection Adaptive task selection dynamically calculated the ensemble weights to integrate the outputs of multiple models. The details of this approach were as follows. 1. Adaptive task selection selects the optimal combination of speech tasks and their weights through CV within the training data. It applies leave-one-out cross-validation (LOOCV), in which one sample is left out, and the remaining data are used for training. This process is referred to as inner CV. In the current study, adaptive task selection was used for disease groups, including the FTD subtypes of PNFA, bvFTD, and SD and ALS. Figure 4 shows the detailed computational process. a. Within the training data, the detection performance (macro-F1) is calculated for each disease subtype \(\:{d}_{j}\:\in\:\:D\) in 16 tasks \(\:({t}_{1},\dots\:,\:{t}_{16})\) . These values form the performance matrix \(\:\mathbf{F}\:\in\:\:{\mathbf{R}}^{16\times\:4}\) . Next, for each disease subtype \(\:{d}_{j}\) , the tasks are sorted in descending order of macro-F1 scores, creating the matrix \(\:{\mathbf{S}}_{j\:}=\:\text{a}\text{r}\text{g}\text{s}\text{o}\text{r}\text{t}\left({\mathbf{F}}_{:,j}\right)\) . b. Based on \(\:\varvec{S}\) , the weight of each task at threshold \(\:k\) is computed. The weight \(\:{w}_{l,k}\:(l=\:1,\dots\:,16)\) is defined as the number of times task \(\:{t}_{l}\) appears within the top \(\:k\) rankings across all disease subtypes. $$\:{w}_{l,k}=\sum\:_{{d}_{j}\in\:\:D}\mathbf{I}(\text{r}\text{a}\text{n}\text{k}\left({t}_{l},\:{d}_{j}\right)\leqq\:k),$$ where \(\:\mathbf{I}\) is an indicator function, and rank \(\:{(t}_{l},\:{d}_{j})\) represents the rank of task \(\:{t}_{l}\) for disease subtype \(\:{d}_{j}\) . Using these weights, the prediction value \(\:{P}_{k}\left(x\right)\) for a sample \(\:x\) is computed as follows: $$\:{P}_{k}\left(x\right)=\:\frac{{\sum\:}_{\:l=\:1}^{16}{w}_{l,\:k}{f}_{l}\left(x\right)}{{\sum\:}_{l\:=\:1}^{16}{w}_{l,\:k}}$$ , where \(\:{f}_{l}\left(x\right)\) represents the prediction value of \(\:x\) for task \(\:l\) . c. The optimal \(\:{k}^{*}\) is determined by performing LOOCV on the training data. The macro-F1 score is calculated for each \(\:k\) , and the value of \(\:k\) that maximizes the performance is selected. $$\:{k}^{*}={argmax}_{k\in\:\left\{1,\dots\:,16\right\}}\text{m}\text{a}\text{c}\text{r}\text{o}‐\text{F}1\left({P}_{k}\right).$$ 2. Weighted ensemble learning combines the predicted values of each task using the weights computed during adaptive task selection. The final prediction for a test sample \(\:{x}_{test}\) is given by the following equation, using the optimal \(\:{k}^{*}\) : $$\:\widehat{y}=\:\left\{\begin{array}{c}1\:if\:{P}_{{k}^{*}}\left({x}_{test}\right)\geqq\:\theta\:\\\:0\:otherwise\:\:\:\:\:\:\:\:\:\:\end{array}\right.$$ . For example, in Fig. 4 , if \(\:{k}^{*}=\:2\) , the task weights are \(\:{w}_{\text{5,2}}\:=\:3,\:{w}_{\text{1,2}}=\:2,\:{w}_{\text{4,2}}=\:2,\) and \(\:{w}_{\text{3,2}}=\:1\) , while the weights of other weak learners are set to 0. \(\:\theta\:\) represented the classification threshold, which was determined from the ROC curve constructed using the LOOCV-based predictions on the training data as the threshold that maximized the Youden index (TPR - FPR). This experiment evaluated the classification performances of weak learner and ensemble learning. In the weak learner method, SVM, random forest, and LightGBM were applied to identify which model performed the best. Table S7ฏ lists the parameters, search ranges, and methods for each model. In this experiment, the number of bagging cycles was set to nine. For ensemble learning, PWEL was compared with the majority voting method to confirm its effectiveness. Moreover, to examine the effectiveness of the two methods employed in PWEL, we compared their accuracy with and without (i) under-bagging and (ii) adaptive task selection. Statistical analysis Continuous variables were compared using the Mann–Whitney U test, Kruskal-Wallis test, chi-square test, or Fisher’s exact test, as appropriate. ROC analyses were performed to verify the diagnostic accuracy (AUC, sensitivity, specificity, accuracy, and macro-F1) of PWEL for FTD and ALS, and the optimal cut-off values for classification were determined based on the Youden index 44 , 45 . As FTD was divided into three subtypes, sensitivity was evaluated separately for each subtype. All statistical analyses were performed using JMP Pro version 16.0 (SAS Institute Inc., Cary, NC, USA) and Python (version 3.8.12) and scikit-learn (version 1.0.2). Statistical significance was set at P < 0.05 except for specific instances mentioned above. Declarations Competing interests The authors declare no competing interests. Funding Declaration This work was supported in part by the Ministry of Education, Culture, Sports, Science and Technology-Japan, Grant-in-Aid for Scientific Research under grant #JP24H00741, and part by the commissioned research by National Institute of Information and Communications Technology (NICT), JAPAN. Author Contribution YI, RO, SK, MM, and HW contributed to the conception and design of the study. RO, MM, MS, HW, TS, AO, KH, SS, YM, MH, MK, MI, and GS contributed to the acquisition of the data. YI, RO, KI, SK, and HW contributed the analysis of the data. YI, RO, SK, and HW contributed to the interpretation of the data. YI, RO, SK, and HW contributed to the draft of the article. YI, RO, SK, and HW revised the manuscript critically for important intellectual content. All authors read and approved the final version of the manuscript. Data Availability The datasets generated during and/or analysed during the current study are available from the corresponding author on reasonable request. References Burrell, J. R., Kiernan, M. C., Vucic, S. & Hodges, J. R. Motor neuron dysfunction in frontotemporal dementia. Brain 134 , 2582–2594 (2011). Iazzolino, B. et al. Validation of the revised classification of cognitive and behavioural impairment in ALS. J. Neurol. Neurosurg. Psychiatry . 90 , 734–739 (2019). Ratnavalli, E., Brayne, C., Dawson, K. & Hodges, J. R. The prevalence of frontotemporal dementia. Neurology 58 , 1615–1621 (2002). Masrori, P. & Van Damme, P. Amyotrophic lateral sclerosis: a clinical review. Eur. J. Neurol. 27 , 1918–1929 (2020). Shinagawa, S., Catindig, J. A., Block, N. R., Miller, B. L. & Rankin, K. P. When a little knowledge can be dangerous: False-positive diagnosis of behavioral variant frontotemporal dementia among community clinicians. Dement. Geriatr. Cogn. Disord . 41 , 99–108 (2016). Coppieters, R. et al. A systematic review of the quantitative markers of speech and language of the frontotemporal degeneration spectrum and their potential for cross-linguistic implementation. Neurosci. Biobehav Rev. 167 , 105909 (2024). Cho, S. et al. Automatic classification of AD pathology in FTD phenotypes using natural speech. Alzheimer’s Dement. 20 , 3416–3428 (2024). Zimmerer, V. C. et al. Automated profiling of spontaneous speech in primary progressive aphasia and behavioral-variant frontotemporal dementia: an approach based on usage-frequency. Cortex 133 , 103–119 (2020). Seeley, W. W. et al. The natural history of temporal variant frontotemporal dementia. Neurology 64 , 1384–1390 (2005). Nevler, N. et al. Automated analysis of natural speech in amyotrophic lateral sclerosis spectrum disorders. Neurology 95 , E1629–E1639 (2020). Vogel, A. P. et al. Motor speech signature of behavioral variant frontotemporal dementia: Refining the phenotype. Neurology 89 , 837–844 (2017). Gong, Y. et al. Exploring emotion and emotional variability as digitalbiomarkers in frontotemporal dementia speech. IEEE Access. 12 , 71419–71432 (2024). Sagi, O. & Rokach, L. Ensemble learning: a survey. Wiley Interdiscip Rev. Data Min. Knowl. Discov . 8 , 1–18 (2018). Ke, X., Mak, M. W. & Meng, H. M. Automatic selection of spoken language biomarkers for dementia detection. Neural Netw. 169 , 191–204 (2024). Duffy, J. R., Peach, R. K. & Strand, E. A. Progressive apraxia of speech as a sign of motor neuron disease. Am. J. Speech-Language Pathol. 16 , 198–208 (2007). da Cunha, P. L. et al. Automated free speech analysis reveals distinct markers of Alzheimer’s and frontotemporal dementia. PLoS One . 19 , 1–19 (2024). Garrard, P., Rentoumi, V., Gesierich, B., Miller, B. & Gorno-Tempini, M. L. Machine learning approaches to diagnosis and laterality effects in semantic dementia discourse. Cortex 55 , 122–129 (2014). Fraser, K. C. et al. Automated classification of primary progressive aphasia subtypes from narrative speech transcripts. Cortex 55 , 43–60 (2014). Christidi, F. et al. Gray matter and white matter changes in non-demented amyotrophic lateral sclerosis patients with or without cognitive impairment: a combined voxel-based morphometry and tract-based spatial statistics whole-brain analysis. Brain Imaging Behav. 12 , 547–563 (2018). Brambati, S. M., Ogar, J., Neuhaus, J. & Miller, B. L. Gorno-Tempini, M. L. Reading disorders in primary progressive aphasia: A behavioral and neuroimaging study. Neuropsychologia 47 , 1893–1900 (2009). Kertesz, A. The Western Aphasia Battery (Grune & Stratton, 1982). Kertesz, A. The Western Aphasia Battery: a systematic review of research and clinical applications. Aphasiology 36 , 21–50 (2022). Hanley, J. A. & McNeil, B. J. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143 , 29–36 (1982). Akobeng, A. K. Understanding diagnostic tests 3: Receiver operating characteristic curves. Acta Paediatr. 96 , 644–647 (2007). Jeancolas, L. et al. C,o,mparison of telephone recordings and professional microphone recordings for early detection of Parkinson’s disease, using mel-frequency cepstral coefficients with Gaussian mixture models to cite this version: HAL Id : hal-02474486 (2020). Comparison of Tele. Rodriguez, J. D., Perez, A. & Lozano, J. A. Sensitivity analysis of k-fold cross validation in prediction error estimation. IEEE Trans. Pattern Anal. Mach. Intell. 32 , 569–575 (2010). Neary, D. et al. Frontotemporal lobar degeneration: a consensus on clinical diagnostic criteria. Neurology 51 , 1546–1554 (1998). Gorno-Tempini, M. L. et al. Classification of primary progressive aphasia and its variants. Neurology 76 , 1006–1014 (2011). Rascovsky, K. et al. Sensitivity of revised diagnostic criteria for the behavioural variant of frontotemporal dementia. Brain 134 , 2456–2477 (2011). Brooks, B. R., Miller, R. G., Swash, M. & Munsat, T. L. El Escorial revisited: revised criteria for the diagnosis of amyotrophic lateral sclerosis. Amyotroph. Lateral Scler. 1 , 293–299 (2000). Folstein, M. F., Folstein, S. E. & McHugh, P. R. Mini-mental state. A practical method for grading the cognitive state of patients for the clinician. J. Psychiatr Res. 12 , 189–198 (1975). Mioshi, E., Dawson, K., Mitchell, J., Arnold, R. & Hodges, J. R. The Addenbrooke’s Cognitive Examination Revised (ACE-R): a brief cognitive test battery for dementia screening. Int. J. Geriatr. Psychiatry . 21 , 1078–1085 (2006). Kertesz, A. Aphasia and associated disorders: taxonomy, localization, and recovery (Grune & Stratton, 1979). Franchignoni, F., Mora, G., Giordano, A., Volanti, P. & Chiò, A. Evidence of multidimensionality in the ALSFRS-R Scale: a critical appraisal on its measurement properties using Rasch analysis. J. Neurol. Neurosurg. Psychiatry . 84 , 1340–1345 (2013). Rooney, J., Burke, T., Vajda, A., Heverin, M. & Hardiman, O. What does the ALSFRS-R really measure? A longitudinal and survival analysis of functional dimension subscores in amyotrophic lateral sclerosis. J. Neurol. Neurosurg. Psychiatry . 88 , 381–385 (2017). Boll, S. F. Suppression of acoustic noise in speech using spectral subtraction. IEEE Trans. Acoust. 27 , 113–120 (1979). Schuller, B., Steidl, S. & Batliner, A. The INTERSPEECH 2009 emotion challenge. Proc. Annu. Conf. Int. Speech Commun. Assoc. INTERSPEECH 312–315 (2009). 10.21437/interspeech.2009-103 Schuller, B. et al. A Survey on perceived speaker traits: personality, likability, pathology, and the first challenge. Comput. Speech Lang. 29 , 100–131 (2015). Kudo, T. Applying conditional random fields to Japanese morphological analysis. Proc. 2004 Conf. Empir. Methods Nat. Lang. Process. 230–237 (2004). SIMPSON, E. H. Measurement of Diversity. Nature 163 , 688–688 (1949). Guyon, I. & Elisseeff, A. An introduction to variable and feature selection. J. Mach. Learn. Res. 3 , 1157–1182 (2003). Blum, A. L. & Langley, P. Selection of relevant features and examples in machine learning. Artif. Intell. 97 , 245–271 (1997). Salmi, M., Atif, D., Oliva, D., Abraham, A. & Ventura, S. Handling imbalanced medical datasets: review of a decade of research. Artif. Intell. Rev. 57 , 273 (2024). Youden, W. J. Index for rating diagnostic tests. Cancer 3 , 32–35 (1950). Fluss, R., Faraggi, D. & Reiser, B. Estimation of the Youden Index and its associated cutoff point. Biometrical J. 47 , 458–472 (2005). Additional Declarations No competing interests reported. Supplementary Files supplementarymaterials20251014.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 18 Feb, 2026 Editor invited by journal 31 Oct, 2025 Editor assigned by journal 29 Oct, 2025 Submission checks completed at journal 29 Oct, 2025 First submitted to journal 21 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7917087","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":594628956,"identity":"be6347e1-f49f-4384-949e-855cc2b716fd","order_by":0,"name":"Yuki Ito","email":"","orcid":"","institution":"Nagoya Institute of Technology Graduate School of Engineering","correspondingAuthor":false,"prefix":"","firstName":"Yuki","middleName":"","lastName":"Ito","suffix":""},{"id":594628957,"identity":"17380a1e-f421-443d-93df-1f20fba25ab3","order_by":1,"name":"Reiko Ohdake","email":"","orcid":"","institution":"Fujita Health University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Reiko","middleName":"","lastName":"Ohdake","suffix":""},{"id":594628958,"identity":"0992fa11-941e-4b66-b47e-78ba07e8ec17","order_by":2,"name":"Shohei Kato","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCElEQVRIie2QsWrDMBCGTwjUxdSrjEF9hSsaQt5GwnsJZCm0UBeDuvgBDHmJPEKCIVkKWQMJ1CbQOdAlUA+9tjQQqAzZStE3SJy4j/9OAIHAH0R+XzNAYI8NnT9PokfhR6VAQDxLAUHNR8VP8rSaN9Bt1eCicHejUafiSc72B7i88SlplHFk7lUPy7nbVIhabmc8KUGMfYqCTEiW13a6tm4TIdqpNJDSlDb3KfGOVuhIeWndmJQHUvh7n5JKSgHxmcIcJ8WgNKI3Jal2Gq2rNT7bIo1QX1c04bBE/y5yZdtm39UKl8v2LerUVVxl9fpwu/D+2BfmtGQ0Ei7ML4393J+vBAKBwH/lA9eKUtl0QdpmAAAAAElFTkSuQmCC","orcid":"","institution":"Nagoya Institute of Technology Graduate School of Engineering","correspondingAuthor":true,"prefix":"","firstName":"Shohei","middleName":"","lastName":"Kato","suffix":""},{"id":594628959,"identity":"0df48b96-1f7c-4288-af2b-fc33b3b94b92","order_by":3,"name":"Michihito Masuda","email":"","orcid":"","institution":"Nagoya University Graduate School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Michihito","middleName":"","lastName":"Masuda","suffix":""},{"id":594628960,"identity":"6bdbe7e4-45c9-4aa0-baa4-99a4454138d5","order_by":4,"name":"Maki Suzuki","email":"","orcid":"","institution":"The University of Osaka Graduate School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Maki","middleName":"","lastName":"Suzuki","suffix":""},{"id":594628961,"identity":"1d5555f6-0f1a-4d07-9249-0d3f7b9cfb82","order_by":5,"name":"Hiroyuki Watanabe","email":"","orcid":"","institution":"The University of Osaka Graduate School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Hiroyuki","middleName":"","lastName":"Watanabe","suffix":""},{"id":594628962,"identity":"4b84415b-c987-405f-8216-ceab5cbabe82","order_by":6,"name":"Takuto Sakuma","email":"","orcid":"","institution":"Nagoya Institute of Technology Graduate School of Engineering","correspondingAuthor":false,"prefix":"","firstName":"Takuto","middleName":"","lastName":"Sakuma","suffix":""},{"id":594628963,"identity":"b1eb7696-0a4b-4589-af7a-404d8c96743c","order_by":7,"name":"Kazuhiro Ishibashi","email":"","orcid":"","institution":"Nagoya Institute of Technology Faculty of Engineering","correspondingAuthor":false,"prefix":"","firstName":"Kazuhiro","middleName":"","lastName":"Ishibashi","suffix":""},{"id":594628964,"identity":"2bcf1a0a-fb87-4a41-9754-aa876882b8d0","order_by":8,"name":"Aya Ogura","email":"","orcid":"","institution":"Nagoya University","correspondingAuthor":false,"prefix":"","firstName":"Aya","middleName":"","lastName":"Ogura","suffix":""},{"id":594628965,"identity":"e3688cd7-2a75-406e-b52a-a4e93e2226e7","order_by":9,"name":"Kazuhiro Hara","email":"","orcid":"","institution":"Nagoya University Graduate School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Kazuhiro","middleName":"","lastName":"Hara","suffix":""},{"id":594628966,"identity":"2bc6c049-322a-4b23-8461-71de0b70cbb5","order_by":10,"name":"Sayuri Shima","email":"","orcid":"","institution":"Fujita Health University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Sayuri","middleName":"","lastName":"Shima","suffix":""},{"id":594628967,"identity":"4a815425-296c-4ca8-9a75-e6f4525fbb73","order_by":11,"name":"Yasuaki Mizutani","email":"","orcid":"","institution":"Fujita Health University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Yasuaki","middleName":"","lastName":"Mizutani","suffix":""},{"id":594628968,"identity":"ed07ccc3-43d3-466e-a22a-326d4d35ac97","order_by":12,"name":"Mamoru Hashimoto","email":"","orcid":"","institution":"Kindai University Faculty of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Mamoru","middleName":"","lastName":"Hashimoto","suffix":""},{"id":594628969,"identity":"46ab2bf6-5a6f-48ce-bb91-f21b252f7d30","order_by":13,"name":"Masahisa Katsuno","email":"","orcid":"","institution":"Nagoya University Graduate School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Masahisa","middleName":"","lastName":"Katsuno","suffix":""},{"id":594628970,"identity":"2948d07c-4376-44e7-8cbe-f143475c9bad","order_by":14,"name":"Manabu Ikeda","email":"","orcid":"","institution":"The University of Osaka Graduate School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Manabu","middleName":"","lastName":"Ikeda","suffix":""},{"id":594628971,"identity":"f341f881-2c3c-4505-bd73-681dc8aa36aa","order_by":15,"name":"Gen Sobue","email":"","orcid":"","institution":"Aichi Medical University","correspondingAuthor":false,"prefix":"","firstName":"Gen","middleName":"","lastName":"Sobue","suffix":""},{"id":594628972,"identity":"4f32c41d-9463-4fcd-9656-fff2b8a07513","order_by":16,"name":"Hirohisa Watanabe","email":"","orcid":"","institution":"Fujita Health University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Hirohisa","middleName":"","lastName":"Watanabe","suffix":""}],"badges":[],"createdAt":"2025-10-22 08:08:49","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7917087/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7917087/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":103216344,"identity":"06feec3f-4a7b-4468-a0d5-e56c52125656","added_by":"auto","created_at":"2026-02-23 09:33:13","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":172707,"visible":true,"origin":"","legend":"\u003cp\u003eROC values for each fold. Left: Result from 5-fold CV. Right: Result from 7-fold CV.\u003cbr\u003e\nThe AUC performance of PWEL is evaluated using k-fold CV. The score shows the mean AUC across all folds.\u003c/p\u003e","description":"","filename":"figure1ROCcurverevised.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7917087/v1/7cf2c6291988c22b32f9f67a.jpeg"},{"id":103216342,"identity":"079ea0c9-ebba-4053-a26a-ae31971a9a5c","added_by":"auto","created_at":"2026-02-23 09:33:13","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":86799,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix of PWEL constructed by aggregating the predictions across all folds in 5-fold CV.\u003c/p\u003e","description":"","filename":"figure2confusionmatrixrevised.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7917087/v1/d917f9b04fd9315783a4f792.jpeg"},{"id":103216343,"identity":"3c67b205-c119-49f5-81dd-1868461ff973","added_by":"auto","created_at":"2026-02-23 09:33:13","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":181060,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of performance-weighted ensemble learning. Let 𝑇 be the number of selected tasks, where 1 ≤ 𝑇 ≤ 16.\u003c/p\u003e","description":"","filename":"figure3PWEL.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7917087/v1/1630ee49ec26c490a4b0448c.jpeg"},{"id":103216345,"identity":"d4d4e0a8-6f16-4c82-86af-1b81770b06db","added_by":"auto","created_at":"2026-02-23 09:33:13","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":304904,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of adaptive task selection. Let F and S be the values evaluated using leave-one-out cross-validation (LOOCV) performed on the training dataset.\u003c/p\u003e","description":"","filename":"figure4ATS.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7917087/v1/6f0db3f7f643037c258cb2f9.jpeg"},{"id":103506011,"identity":"560cafd1-0069-473e-82a3-130b79d32010","added_by":"auto","created_at":"2026-02-26 13:33:48","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2075517,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7917087/v1/3d79587c-4f32-42ca-9650-582aa1e1b334.pdf"},{"id":103216346,"identity":"322da0a5-ecf9-43a5-9bfa-287f3dc46598","added_by":"auto","created_at":"2026-02-23 09:33:14","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":32474,"visible":true,"origin":"","legend":"","description":"","filename":"supplementarymaterials20251014.docx","url":"https://assets-eu.researchsquare.com/files/rs-7917087/v1/c7aa6234136eaa42ce7e0476.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Performance-weighted ensemble learning for detecting patients with FTD and ALS from short reading tasks","fulltext":[{"header":"Introduction","content":"\u003cp\u003eFrontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS) belong to the same spectrum of neurodegenerative diseases and present with overlapping clinical manifestations and neuropathological features\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. FTD encompasses a heterogeneous group of clinical syndromes characterized by progressive behavioral changes, executive dysfunction, or language impairments. The three main FTD subtypes are behavioral variant frontotemporal dementia (bvFTD) and two primary progressive aphasia (PPA) subtypes, namely, progressive non-fluent aphasia (PNFA) and semantic dementia (SD). Notably, 10\u0026ndash;15% of individuals with FTD fulfill the diagnostic criteria for ALS \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. Furthermore, approximately half of patients with ALS exhibit varying degrees of cognitive and behavioral impairments, similar to those seen in FTD \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. There is currently no effective treatment to halt or delay disease progression in FTD. Its prevalence in patients with early-onset dementia (age\u0026thinsp;\u0026lt;\u0026thinsp;65 years) is comparable to that of Alzheimer\u0026rsquo;s disease, affecting approximately 15 per 100,000 individuals \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Meanwhile, ALS predominantly occurs in adults aged 45\u0026ndash;75 years, with annual prevalence rates of 4\u0026ndash;8 per 100,000 people \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. ALS results in substantial social and personal burdens to both patients and their caregivers.\u003c/p\u003e \u003cp\u003eDespite improvements in their clinical understanding, both FTD and ALS can be difficult to accurately diagnose, and thus delayed or incorrect diagnoses are common. In one study, 60% of patients initially diagnosed with bvFTD in community settings did not meet the diagnostic criteria when evaluated by specialists \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. This discrepancy underscores the complexity of FTD presentations and highlights the need for specialist involvement in the diagnostic workup. In ALS, certain clinical features, such as concomitant FTD, bulbar onset, rapid functional decline, significant weight loss, older age at onset, and reduced forced vital capacity, are associated with shorter survival \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Moreover, ALS is a clinically heterogeneous disease in which frontotemporal lobe involvement may be undetected, unless specifically assessed. Collectively, these challenges highlight the urgent need for a low-cost, non-invasive, quantitative, and time-efficient screening tool to identify patients with FTD and with ALS and reduce diagnostic delays.\u003c/p\u003e \u003cp\u003eAccumulating evidence suggests that speech and language changes arise early in FTD, and thus, they may be noninvasive parameters for screening\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. However, although prior machine learning research in this area focused primarily on linguistic features derived from spontaneous speech, these approaches often require lengthy and sometimes burdensome data collection protocols, as well as time-consuming manual transcription by trained annotators \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Clinically, patients with FTD or ALS may present with psychiatric or behavioral symptoms (e.g., apathy and disinhibition), speech stereotypies, and reduced verbal output, all of which can complicate lengthy interviews or spontaneous speech tasks. Similarly, specific variants such as patients with SD may feature pronounced anomia, repetitive speech, and behavioral symptoms that hamper extensive data collection \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Although existing methods are promising, they may be impractical for rapid screening or large-scale implementation.\u003c/p\u003e \u003cp\u003eTo overcome these challenges, recent investigations have shifted towards analyzing abnormal acoustic properties in the speech of patients with FTD and ALS \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e, leading to preliminary machine learning models that capture paralinguistic changes associated with emotional and cognitive deficits \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Building on these developments, we propose a machine learning approach focused on acoustic features extracted from a brief reading task to discriminate patients with FTD and with ALS from healthy controls (HCs). Our method requires only approximately 1 min of data collection: participants are asked to read single words or short sentences displayed on a screen without further prompts or examiner intervention. In addition to linguistic and temporal measurements, we emphasize the acoustic and paralinguistic features of the recorded speech. Furthermore, to facilitate scalability and reduce the examiner burden, our pipeline is designed for automated pre-processing and analysis, potentially allowing at-home recordings and adaptation to multiple languages.\u003c/p\u003e \u003cp\u003eIn this paper, we present our machine-learning framework for distinguishing patients with FTD and with ALS from healthy individuals based on acoustic, linguistic, and temporal features derived from a simple reading task.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDiagnostic accuracy of performance-weighted ensemble learning\u003c/h2\u003e \u003cp\u003eOur proposed model, performance-weighted ensemble learning (PWEL), showed superior performance compared to a simple ensemble learning model and task-wise classification, with an area under the curve (AUC) of 0.840 (95% confidence interval [CI]: 0.786\u0026ndash;0.891); sensitivity, 0.828ฏ (95% CI: 0.750\u0026ndash;0.894); specificity, 0.751 (95% CI: 0.667\u0026ndash;0.828); accuracy, 0.792 (95% CI: 0.731\u0026ndash;0.843); and macro-F1, 0.788 (95% CI: 0.730\u0026ndash;0.842) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the receiver operating characteristic (ROC) curves for 5-fold cross validation (CV). The AUC values for each fold ranged from 0.799 to 0.879. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the PWEL confusion matrix combining the results of 5-fold CV. The optimal values of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{k}^{*}\\)\u003c/span\u003e\u003c/span\u003e for each fold were [8, 1, 3, 3, and 12]. The optimal classification thresholds \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\theta\\)\u003c/span\u003e\u003c/span\u003e determined for each fold were [0.500, 0.500, 0.333, 0.500, and 0.292]. For weak learners, support vector machine (SVM) showed higher accuracy and macro-F1 values than did LightGBM and random forest. Therefore, we adopted SVM as the inner/outer weak learner in PWEL.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDiagnostic accuracy in PWEL (FTD and ALS vs. HC).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"10\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c7\" namest=\"c3\"\u003e \u003cp\u003eSensitivity for each disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSpec\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAcc\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003emacro-F1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFTD and ALS\u003c/p\u003e \u003cp\u003etotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePNFA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ebvFTD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eALS\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eEnsemble Learning\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePWEL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.828\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.700\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.933\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.812\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.751\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e0.792\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003e0.788\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMajority voting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.759\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1.000\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.700\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.867\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.715\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.784\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.772\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.770\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eTask-wise classification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.648\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.806\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.550\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.775\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.615\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.718\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.681\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.679\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandomForest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.681\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.808\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.575\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.841\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.631\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.659\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.671\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.667\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLightGBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.704\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.881\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.550\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.846\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.674\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.654\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.680\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.676\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"10\"\u003eThe task-wise classification indicates the mean classification performance across the 16 tasks.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"10\"\u003eHighest records are reported in bold type.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"10\"\u003ePWEL, Performance-Weighted Ensemble Learning; SVM, Support Vector Machine; Spec, Specificity; Acc, Accuracy.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eAblation study\u003c/h3\u003e\n\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents the comparison results for the four models, with and without under-bagging and adaptive task selection. Under-bagging or adaptive task selection improved specificity in the four models, and the combination of under-bagging and adaptive task selection applied to the PWEL model improved accuracy and macro-F1.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison results for the four models combined with and without under-bagging and ATS.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"10\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eunder-bagging\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eATS\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c7\" namest=\"c3\"\u003e \u003cp\u003eSensitivity for each disease\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSpec\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAcc\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003emacro-F1\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFTD an ALS\u003c/p\u003e \u003cp\u003etotal\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePNFA\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ebvFTD\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSD\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eALS\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e✓\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e✓\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.828\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.700\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.933\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.812\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.751\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e0.792\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003e0.788\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e✓\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.759\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1.000\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.700\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.867\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.715\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.784\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.772\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.770\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e✓\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.817\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1.000\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.600\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.867\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.810\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.719\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.771\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.769\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.865\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1.000\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.700\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.867\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.864\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.776\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.772\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"10\"\u003eHighest records are reported in bold type.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"10\"\u003eATS, Adaptive Task Selection; Spec, Specificity; Acc, Accuracy.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e\n\u003ch3\u003eAccuracy comparison among the three facilities\u003c/h3\u003e\n\u003cp\u003eA total of 140, 28, and 29 participants were recruited from one facility each in Nagoya, Toyoake, and Osaka, respectively (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). There was no significant difference in the accuracy rate of PWEL among these facilities (P\u0026thinsp;=\u0026thinsp;0.088).\u003c/p\u003e\n\u003ch3\u003eSpeech features of reading tasks selected through 5-fold CV\u003c/h3\u003e\n\u003cp\u003eAmong the tasks selected through adaptive task selection, three tasks (task 3, 8, and 15) were commonly selected across 4 of 5 folds. An examination of the top 10 features for each of the three tasks revealed that acoustic features were the most prevalent (96.7%), followed by temporal features (3.3%). No linguistic features were included in the top 10 features. MFCC\u0026mdash;representing the vocal tract transfer function\u0026mdash;was the most common acoustic feature, accounting for 66.7% (e.g., mfcc_sma_de[1]_maxPos, mfcc_sma[1]_min, mfcc_sma[4]_minPos). It was followed by Pcm zcr as a proxy for pitch at 13.3% (e.g., zcr_sma_de_kurtosis, zcr_sma_de_maxPos), RMS energy as an indicator of volume at 10.0% (e.g., RMSenergy_sma_linregc1, RMSenergy_sma_min), and voice probability as a measure of voicing at 6.7% (e.g., voiceProb_sma_de_skewness).\u003c/p\u003e\n\u003ch3\u003eDiscrimination performance by task and disease subtype\u003c/h3\u003e\n\u003cp\u003eThe discrimination performance by task and disease subtype is presented in Table S2. In patients with PNFA, the sensitivity to multiple word and sentence tasks reached 1.000. In patients with SD, the sensitivity to sentence tasks reached 1.000. Meanwhile, sensitivity was lower than 0.900 in all tasks in patients with bvFTD and ALS. The sensitivity for each task varied in patients with ALS, with the highest sensitivity being for the longest-sentence task (task No.15) (0.699), followed by that for three-letter words (No.3) (0.690). Therefore, task order (i.e., fatigue associated with performing tasks) did not affect diagnostic accuracy. In all study participants, accuracy and macro-F1 were the highest for the longest-sentence task (0.751 and 0.747, respectively) and lowest for the non-word task (0.594 and 0.591, respectively).\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eClinical characteristics of misclassified participants\u003c/h2\u003e \u003cp\u003eThere was no significant difference in demographic variables, cognitive scores, or language functions between the HC error group (HCs classified to have FTD and ALS) and the HC correct group (Table S3). Compared to the FTD and ALS correct groups, the FTD and ALS error groups (classified as HCs) were younger (P\u0026thinsp;=\u0026thinsp;0.042), had higher cognitive function (MMSE: P\u0026thinsp;=\u0026thinsp;0.047), and had milder ALS Functional Rating Scale-Revised (ALSFRS-R) bulbar symptoms (P\u0026thinsp;=\u0026thinsp;0.018). The 18 patients with misclassification in the disease group (false negatives in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) included 14 patients with ALS, 2 patients with bvFTD, 1 patient with SD, and 1 patient with PNFA. The ALS group had the highest misclassification rate among the disease groups, with 18.9% of patients classified as having HCs.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, PWEL, which included under-bagging and adaptive task selection, achieved good discrimination ability for identifying patients with FTD and ALS. It showed very high sensitivity for FTD subtypes and acceptable sensitivity for ALS. Further, PWEL demonstrated good generalization performance with high accuracy across 5-fold CV and consistent accuracy across recording sites.\u003c/p\u003e \u003cp\u003eThe strong performance of PWEL can be attributed to two key components of its design: (i) the correction of class imbalance through under-bagging and (ii) the use of ensemble learning with adaptive task weighting. First, under-bagging addresses the inherent class imbalance in clinical datasets, particularly the smaller number of patients with FTD than that of patients with ALS and HCs. Models trained on imbalanced data tend to favor the majority class, which can degrade the sensitivity for underrepresented groups. By implementing under-bagging, combining undersampling with bagging, PWEL ensures that each training subset maintains class balance, enabling better detection of rare subtypes such as FTD. Second, PWEL employs ensemble learning by integrating the predictions from multiple classifiers trained on different reading tasks. This approach leverages the strengths of individual classifiers and outperforms single learners in various domains \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Through adaptive task selection, the model dynamically identifies and weighs tasks that contribute the most to the classification accuracy. By combining task-level predictions with learned weights, PWEL enhances the overall robustness and sensitivity of the findings. As shown by the results of our ablation study, the combination of under-bagging and adaptive task selection yielded the highest accuracy and macro-F1 scores. Collectively, these design elements enable PWEL to achieve superior discriminative performance in comparison to conventional machine learning approaches, particularly in the context of clinically heterogeneous and imbalanced datasets.\u003c/p\u003e \u003cp\u003eAmong the 405 speech features used in this study, 384 were acoustic features derived from the INTERSPEECH 2009 Emotion Challenge (IS09) feature set. Numerous studies have adopted IS09 subsets for analyzing speech in affective and neurological contexts\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. However, to our best knowledge, this is the first study to apply a machine learning model that primarily leverages acoustic features to detect FTD and ALS in a relatively large cohort that involves a large number of patients with ALS. This emphasis on acoustic features may be especially relevant in ALS, where bulbar symptoms contribute to altered speech. In our analysis, patients with ALS who were misclassified as being HCs exhibited significantly milder bulbar symptoms, as measured using the ALSFRS-R subscores. This suggests that the model may detect bulbar motor involvement. However, PWEL had higher sensitivity for SD (0.933) as an FTD subtype than for ALS (0.812). Previous studies have suggested that the fundamental frequency (f0) range is associated with bulbar motor dysfunction in ALS, whereas increased pause duration and speech disruption are associated with cognitive deficits \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Moreover, patients with progressive apraxia of speech may exhibit articulatory distortions similar to those with motor neuron disease \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. These findings support the premise that PWEL can capture both dysarthria and FTD-like speech anomalies in ALS, thereby contributing to its diagnostic utility.\u003c/p\u003e \u003cp\u003ePWEL also demonstrated comparable performance to models that relied on linguistic features while offering notable practical advantages. Prior machine learning approaches for FTD and PPA have employed natural language processing on free speech (e.g., narrative descriptions or daily life conversations), reporting AUCs ranging from 0.700 to 0.900 and accuracies ranging from 80% to 90% \u003csup\u003e7,8,16\u0026ndash;18\u003c/sup\u003e. Although rich in diagnostic information, such approaches often require labor-intensive pre-processing (e.g., correction of phonological errors, agrammatism, and fillers) that is typically performed manually. Consequently, the results may be influenced by the accuracy and consistency of the transcription pipelines. In contrast, PWEL automatically extracts acoustic features from a short, structured reading task, thus minimizing examiner bias and transcription variability. This approach not only reduces the workload, but also enhances scalability and objectivity in real-world screening scenarios.\u003c/p\u003e \u003cp\u003eAlthough the IS09 was originally designed for emotion classification, its use here is justified by overlaps in speech alterations observed in both affective and neurological disorders. Particularly, prosodic and spectral features are effective in detecting dysarthria and motor speech anomalies. Our choice to prioritize acoustic features over linguistic features reflects a deliberate tradeoff between reducing the need for manual transcription and increasing model robustness in patients with low verbal output. Nevertheless, future studies are needed to systematically evaluate and optimize the feature sets for neurological disease classification.\u003c/p\u003e \u003cp\u003eWe employed a short reading task in PWEL that lasted approximately 1 min, using words and sentences from the WAB. We hypothesized that this approach would enhance task feasibility for patients with FTD and ALS while also improving the discriminative power of the machine learning model based on acoustic features. A short assessment duration is particularly important for patients with ALS who often experience fatigue due to motor dysfunction. The use of reading aloud tasks has several clinical advantages for FTD subtypes. Patients with PNFA typically exhibit effortful and halted speech due to articulatory difficulties; however, they often have mild and phonemic reading problems \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Meanwhile, patients with SD generally retain fluent speech, and, unless affected by surface dyslexia, their ability to read aloud is comparable to that of healthy individuals \u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. In patients with bvFTD, compulsive reading, a behavior associated with environmental dependency syndrome, may paradoxically increase engagement and feasibility, even in those with disinhibition or inertia. Moreover, as the tasks are presented visually, issues related to working memory or hearing impairments do not interfere with data collection.\u003c/p\u003e \u003cp\u003eThe enhanced classification performance observed with PWEL may be attributable to the emphasis on acoustic properties. By constraining the verbal content to standardized WAB items, we reduced linguistic variability and amplified interindividual differences in acoustic patterns. This may have enabled a more precise capture of paralinguistic features, such as emotional prosody, vocal effort, and articulation. Language independence is a notable advantage for international application, particularly in settings where linguistic resources and expert transcribers are limited. Moreover, reading tasks adopted from the WAB encompass a range of linguistic complexities that can indirectly reflect cognitive linguistic abilities. These include variations in word frequency, syllable count, sentence length, and grammatical structure \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. Importantly, the use of a structured reading task minimizes the burden on both participants and examiners and facilitates standardized, reproducible data collection across settings.\u003c/p\u003e \u003cp\u003eFurthermore, it is noteworthy that the PWEL reading task was developed based on the WAB. The WAB has been translated into at least 33 languages since its original development by Kertesz et al. \u003csup\u003e22\u003c/sup\u003e and is now widely used worldwide as a standardized tool for aphasia assessment. It has been proven effective in evaluating aphasia not only due to cerebrovascular disorders, but also due to neurodegenerative diseases \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. Given that the WAB has been linguistically and psychometrically validated across multiple languages, it may be possible to develop equivalent reading tasks with levels of difficulty and linguistic complexity matched to our study. However, cross-linguistic applicability is a future direction to be empirically evaluated.\u003c/p\u003e \u003cp\u003eIn this study, we discuss the generalization performance of PWEL from the following two perspectives: (i) the generalization performance of AUC scores across each fold in 5-fold CV and (ii) the percentage of correct responses by each institution. First, the mean AUC across the 5 folds was 0.840, with fold-specific values of 0.836, 0.832, 0.799, 0.856, and 0.879. Although 1 fold showed an AUC slightly below 0.800, the majority of folds achieved values close to or above this threshold, and the overall performance remained within an acceptable range for diagnostic applications \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Furthermore, the 95% CI width of 0.105 is relatively narrow, likely representing a precise estimate with limited sampling variation. Such an interval suggests that the PWEL model\u0026rsquo;s performance metric is more reliable for diagnosis in the sense that it has much less variability and edge cases.\u003c/p\u003e \u003cp\u003e Second, the generalizability of our approach was supported by consistent classification performance across the three participating institutions. Data were collected by three examiners in various clinical settings, including outpatient departments, inpatient wards, and affiliated facilities. As such, differences in background noise across recording locations are inevitable and can affect model training and performance. To mitigate this, we applied the spectral subtraction method \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e as a preprocessing step to reduce the background noise in all audio recordings. This technique helped standardize the acoustic quality of the input data and may have contributed to the consistency in model performance across sites. Overall, the findings support the generalizability and robustness of the proposed method. The development of robust, low-cost, and scalable tools, such as PWEL, has important implications for the early screening of ALS and FTD. Such tools can be deployed not only in clinical environments, but also in home-based or telemedicine settings, facilitating timely referral to specialists and reducing diagnostic delays.\u003c/p\u003e \u003cp\u003eThis study had limitations. First, the sample size, particularly for the FTD subtypes, was relatively small, limiting the statistical power and increasing the risk of inflated performance metrics in the subgroup analyses. To mitigate the risk of overfitting, we employed 5-fold CV and ablation studies as 5-fold CV is considered suitable for our moderate-sized datasets (n\u0026thinsp;=\u0026thinsp;197), offering a practical balance between bias and variance \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Importantly, AUC values were above 0.800 in 4 of the 5 folds, suggesting that the model maintained robustness even under limited data conditions. Second, although the acoustic features were derived from a validated and widely used feature set (IS09), the assumption that they are optimal for disease classification remains to be tested empirically. Future research should explore whether alternative or task-specific features can improve discriminative performance.\u003c/p\u003e \u003cp\u003eThird, missing linguistic features were imputed as zero in patients in whom automatic speech recognition (ASR) failed. However, ASR failure occurred in only 2/197 patients (1.0%), and the affected patients were evenly distributed across subtypes (1 patient with bvFTD and 1 patient with SD). To preserve input dimensionality, we imputed zero values, which arguably reflected severe speech impairment. Speech impairment is a potentially informative signal rather than random noise. Given the negligible frequency and distribution, this handling was unlikely to introduce systematic bias. However, more refined imputation and error-handling approaches must be developed. Finally, the high sensitivity observed in some FTD subtypes of each reading task (e.g., 1.000) may reflect the small number of test participants and should thus be interpreted cautiously. Although a separate external validation set was not used, speech data were collected from three distinct institutions across various clinical settings (inpatient, outpatient, and affiliated clinics). The consistency in the classification accuracy across these centers supports preliminary generalizability. However, the classification accuracy of PWEL could have been further improved by tailoring the noise-reduction strategies to the specific environmental characteristics of each facility. External validation using geographically and linguistically distinct datasets is planned in future studies.\u003c/p\u003e \u003cp\u003eIn conclusion, PWEL can be useful in distinguishing patients with FTD and ALS from healthy individuals through simply reading words and sentences for 1 min. Furthermore, noise reduction and the use of WAB tasks increase feasibility across different locations and languages, and automated machine learning with acoustic features improves time efficiency for the identification of patients with FTD and ALS. Collectively, our findings provide evidence that PWEL is a practical, scalable, and a promising screening tool adaptable to multiple sites that can be used for the clinical detection of neurodegenerative disorders.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eStudy design and participants\u003c/h2\u003e \u003cp\u003e This prospective study was approved by the Ethics Committees at Fujita Health University (approval number: HM19-216), Nagoya University (approval number: 2018\u0026thinsp;\u0026minus;\u0026thinsp;0243), The University of Osaka (approval number: 19183), and Nagoya Institute of Technology (approval number: 2019-011) and conformed to the Ethical Guidelines for Medical and Health Research Involving Human Subjects, as endorsed by the Japanese government. All patients provided written informed consent after the procedure was explained.\u003c/p\u003e \u003cp\u003ePatients diagnosed with FTD and ALS at the Department of Neurology at Fujita Health University and Nagoya University in accordance with the established diagnostic criteria for FTD\u003csup\u003e\u003cspan additionalcitationids=\"CR28\" citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e and ALS\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e were enrolled from November 2018 to February 2023. Patients diagnosed with FTD at the Department of Psychiatry, Osaka University were also recruited from December 2019 to February 2023. The exclusion criteria were as follows: (i) a history of other neurological diseases, stroke, or traumatic brain injury; (ii) unintelligible speech, defined as an ALSFRS-R speech subscore\u0026thinsp;\u0026le;\u0026thinsp;2; (iii) Mini-Mental State Examination (MMSE) score\u0026thinsp;\u0026lt;\u0026thinsp;26 \u003csup\u003e31\u003c/sup\u003e and Japanese version of Addenbrooke\u0026rsquo;s Cognitive Examination-Revised (ACE-R) score\u0026thinsp;\u0026lt;\u0026thinsp;89 \u003csup\u003e32\u003c/sup\u003e in HCs; and (iv) severe hearing and visual impairments that could interfere with task performance. A total of 104 patients (30 patients with FTD and 74 patients with ALS) were included. Among the 30 patients with FTD, 15, 7, and 8 patients had SD, PNFA, and bvFTD, respectively. We also included 93 age-matched HCs from the three institutions. All participants (n\u0026thinsp;=\u0026thinsp;197) were native Japanese speakers. The participants were matched by age but not by sex or educational attainment (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Consequently, age and sex were matched within each fold of the training and test datasets (Table S4), and year of education was considered in the analysis of the discrimination outcomes.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDemographics, clinical characteristics, and cognitive function of the participants.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c6\" namest=\"c2\"\u003e \u003cp\u003eFTD and ALS\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eHC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ep - value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal, \u003c/p\u003e \u003cp\u003en\u0026thinsp;=\u0026thinsp;104\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePNFA, \u003c/p\u003e \u003cp\u003en\u0026thinsp;=\u0026thinsp;7\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ebvFTD, \u003c/p\u003e \u003cp\u003en\u0026thinsp;=\u0026thinsp;8\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSD, \u003c/p\u003e \u003cp\u003en\u0026thinsp;=\u0026thinsp;15\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eALS, \u003c/p\u003e \u003cp\u003en\u0026thinsp;=\u0026thinsp;74\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003en\u0026thinsp;=\u0026thinsp;93\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e66.4 (10.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.1 (6.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e62.5 (9.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e69.6 (7.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e65.6 (10.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e67.9 (8.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.425\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003egender, M / F\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e54 / 50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3/4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3/5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6/9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e42/ 32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e35 / 58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.044\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eeducation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e12.2 (2.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12.0 (1.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e14.5 (3.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e12.5 (2.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e11.9 (2.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e13.4 (2.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMMSE\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e24.3 (5.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e23.0 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e21.4 (5.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16.5 (7.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e26.3 (3.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e28.8 (1.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eACE-R\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e76.1 (22.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e70.9 (13.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e61.9 (18.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e35.4 (14.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e86.5 (12.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e95.7 (3.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWAB, AQ\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e86.5 (14.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e77.2 (10.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e79.5 (11.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e60.2 (15.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e93.6 (4.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e97.1 (2.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSpontaneous Speech (/20)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e17.0 (2.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e13.9 (1.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e15.1 (3.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13.5 (2.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e18.2 (1.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.3 (0.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAuditory comprehension (/10)\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.9 (1.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8.7 (0.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7.7 (1.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6.6 (2.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e9.6 (0.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.8 (0.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRepetition (/10)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e9.3 (1.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.8 (2.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9.2 (1.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.3 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e9.8 (0.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.9 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNaming (/10)\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.1 (2.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8.3 (1.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7.7 (1.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2.6 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e9.3 (0.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.6 (0.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003edisease duration\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.1 (2.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2.2 (1.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2.5 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6.0 (2.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.3 (1.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eALSFRS, total\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e40.0 (3.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebulbar 1\u0026ndash;3 (/12)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e11.0 (1.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emotor 4\u0026ndash;9 (/24)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e17.5 (3.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003erespiratory 10\u0026ndash;12 (/12)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e11.5 (1.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003eData are mean (SD).\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003eFTD, frontotemporal dementia; ALS, amyotrophic lateral sclerosis; HC, healthy control; bvFTD, behavioral variant of frontotemporal dementia; PNFA, progressive non fluent aphasia; SD, semantic dementia; ALSFRS-R, ALS functional rating scale-revised; MMSE, Mini-Mental State Examination; ACE-R, Addenbrooke's Cognitive Examination-Revised; WAB, Western Aphasia Battery; AQ, Aphasia Quotient.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003ethe Mann-Whitney U test and the chi-square tests for demographic comparisons between the two independent study groups.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003e\u003csup\u003e1\u003c/sup\u003e The number of participants was different (for disease duration: SD, n\u0026thinsp;=\u0026thinsp;13; for ALSFRS: ALS, n\u0026thinsp;=\u0026thinsp;73; for MMSE, ACE-R: ALS, n\u0026thinsp;=\u0026thinsp;73; for WAB AQ, Auditory comprehension: ALS, n\u0026thinsp;=\u0026thinsp;71; for WAB Naming: ALS, n\u0026thinsp;=\u0026thinsp;73).\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eClinical data\u003c/h2\u003e \u003cp\u003eLanguage function and global cognition were evaluated through a series of neuropsychological assessments (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Language ability was assessed using the WAB \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e, while global cognitive function was measured using the MMSE and ACE-R. The AQ, a composite score based on spontaneous speech, auditory comprehension, repetition, and naming, was derived from the WAB. Disease duration was assessed in patients with both FTD and ALS. Functional disability at the time of testing was assessed using the ALSFRS-R in patients with ALS. To further characterize ALS symptoms, the ALSFRS-R was divided into three subscores according to previous studies \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e,\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e: bulbar subscore (maximum score\u0026thinsp;=\u0026thinsp;12, items 1\u0026ndash;3), motor subscore (maximum score\u0026thinsp;=\u0026thinsp;24, items 4\u0026ndash;9), and respiratory subscore (maximum\u0026thinsp;=\u0026thinsp;12, items 10\u0026ndash;12). In relation to the outputs of our machine learning model, the associations between classification performance and clinical variables, such as motor function and cognitive scores, were also examined.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eReading tasks\u003c/h2\u003e \u003cp\u003eWe applied the words and sentences of the repetition task from the WAB to our reading task. The WAB repetition task included various levels of language ability, such as frequent words with an increasing number of syllables, compound words, numbers, number-word combinations, common sentences, rare sentences, and sentences of increasing length, as well as grammatical complexity. Meanwhile, the WAB reading tasks were designed to assess reading comprehension, and the participants were asked to read aloud and act on the text. Therefore, we hypothesized that the content of the WAB repetition task would be suitable for assessing language ability in reading texts. Finally, one compound non-word was added to examine the influence of the semantic properties of language, and the reading task in this study consisted of 10 words and 6 sentences. Given that the Japanese language has three characters (i.e., Hiragana, Katakana, and Kanji), we added the Kana characters above Kanji characters to account for the differences between them. The participants were asked to read 16 tasks that were presented as a single line on a screen while their speech was digitally recorded. Voice recordings were collected using an AT9921 with a unidirectional microphone (Audio-Technica Corporation, Japan) and a PCM-A10 (SONY, Japan) placed approximately 30\u0026ndash;40 cm from the mouth. Voice samples were recorded in a linear PCM format (.wav) at a sampling rate of 44.1 kHz, with a 16-bit sample size.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eProposed method: PWEL\u003c/h2\u003e \u003cp\u003eThe system identified the optimal combination of reading tasks using PWEL. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates an overview of the proposed system, from preprocessing to output.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003ePreprocessing: noise reduction\u003c/h2\u003e \u003cp\u003eDenoising was applied to the audio data as a preprocessing step. The recordings were collected at multiple medical institutions. As background noise varied depending on the recording environment and could adversely affect model training, we applied the spectral subtraction method \u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e. In this approach, the observed signal is modeled as\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:x\\left(t\\right)=s\\left(t\\right)+n\\left(t\\right)$$\u003c/div\u003e\u003c/div\u003e,\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e as the recorded audio at an arbitrary time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:s\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e, the underlying clean speech signal; and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:n\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e is background noise\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:.\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003eAfter applying the discrete Fourier transform (DFT), the relationship becomes\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:X\\left(\\omega\\:\\right)=S\\left(\\omega\\:\\right)+N\\left(\\omega\\:\\right),$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:X\\left(\\omega\\:\\right)\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:S\\left(\\omega\\:\\right)\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\:\\left(\\omega\\:\\right)\\)\u003c/span\u003e\u003c/span\u003e denote the spectra of the observed signal, clean speech, and noise, respectively. The noise power spectrum \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\:\\left(\\omega\\:\\right)\\)\u003c/span\u003e\u003c/span\u003e was estimated from manually identified non-speech intervals that contained only background noise and no vocal activity. The denoised spectrum was then obtained by subtracting this estimate from \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\:\\left(\\omega\\:\\right)\\)\u003c/span\u003e\u003c/span\u003e, and the time-domain signal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{s}\\left(t\\right)\\:\\)\u003c/span\u003e\u003c/span\u003ewas reconstructed using the inverse DFT. Importantly, non-speech intervals were not removed from the recordings; they were used solely to estimate the noise statistics, thereby preserving pause-related temporal features.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eFeature extraction\u003c/h2\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003eDefinition of speech features\u003c/h2\u003e \u003cp\u003eSpeech features were extracted from voice recordings of 16 reading tasks. A total of 405 features, including 384 acoustic, 17 linguistic, and 4 temporal features, were extracted from the speech samples. Each feature type is detailed below.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eAcoustic features\u003c/h2\u003e \u003cp\u003eThe acoustic features used in this study were selected from the IS09 \u003csup\u003e37\u003c/sup\u003e set, a feature set widely used in affective computing and neurodegenerative disease studies \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. The IS09 feature set includes prosodic, spectral, and voice quality characteristics that reflect emotional states and motor impairments in speech. Although originally designed for emotion classification, it has also been applied in multiple neurological domains owing to its robustness and broad coverage of speech signal properties. The low-level descriptor features and statistical measures are shown in Tables\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, respectively. All 384 features were extracted by calculating 12 types of statistical measures for each of the 16 types of basic features related to vocal tract characteristics, volume, pitch, and voice probability. Feature extraction was performed using openSMILE software version 2.3.0 (audEERING GmbH, Gilching, Germany) \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eLLD and the 12 statistical measures calculated from LLD.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAcoustic Property\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLLD\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVolume\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRMSenergy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRoot-mean-square signal frame energy\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVocal Tract Transfer Function\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMFCC 1\u0026ndash;12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMel-Frequency cepstral coefficients 1\u0026ndash;12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePitch\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePcm zcr\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eZero-crossing rate of time signal (frame-based)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVoicing Probability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVoice Prob\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eThe voicing probability computed from the ACF\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePitch\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eThe fundamental frequency computed from the Cepstrum\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e*_de\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eThe first derivative of each of the above features\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003efunctionals\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003e\u003cb\u003eDescription\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emax\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe maximum value of the contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emin\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe minimum value of the contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003erange\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003emax-min\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emaxPos\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe absolute position of the maximum value (in frames)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eminPos\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe absolute position of the minimum value (in frames)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eamean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe arithmetic mean of the contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elinregc1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe slope (m) of a linear approximation of the contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elinregc2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe offset (t) of a linear approximation of the contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elinregerrO\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe quadratic error computed as the difference of the linear approximation and the actual contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003estddev\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe standard deviation of the values in the contour\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eskewness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe skewness (3rd order moment)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ekurtosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eThe kurtosis (4th order moment)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"3\"\u003eLLD, Low-Level Descriptors\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eLinguistic features\u003c/h2\u003e \u003cp\u003e Linguistic features were extracted from the participants\u0026rsquo; speech by first transcribing audio into text using the Azure Speech-to-Text system, a cloud-based ASR service provided by Microsoft Azure AI. In rare instances where ASR failed owing to incapable vocal intensity, all linguistic features were set to zero to maintain consistent input dimensionality across participants. Only 2 (1 patient with bvFTD and 1 patient with SD) of the 197 patients (1.0%) were affected by this imputation. Notably, no ASR failures occurred among participants with ALS. The transcribed text was further processed using MeCab\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e. Briefly, MeCab is an open-source morphological analysis engine tool for Japanese morphological analysis. Based on the morphological output, 17 linguistic features were extracted. Each feature is described in Table S5. The total number of words was denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\)\u003c/span\u003e\u003c/span\u003e, and the number of unique words was denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:V\\left(N\\right)\\)\u003c/span\u003e\u003c/span\u003e. The type-token ratio (TTR) and Simpson\u0026rsquo;s D value\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e were also calculated to assess lexical diversity, as follows:\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:TTR=\\:\\frac{V\\left(N\\right)}{N}.$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equd\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equd\" name=\"EquationSource\"\u003e\n$$\\:D=\\:\\sum\\:_{m\\:=\\:1}^{N}V\\left(m,N\\right)\\frac{m\\left(m-1\\right)}{N\\left(N-1\\right)}$$\u003c/div\u003e\u003c/div\u003e,\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:V\\:(m,\\:N)\\)\u003c/span\u003e\u003c/span\u003e is the number of unique words that appear m times in a text of length \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eSimpson\u0026rsquo;s D value is high when the same word is repeatedly used, suggesting low lexical diversity. Although the content of the reading task was predetermined, responses varied and included distortion or substitution of sounds due to effortful speech, mispronunciations of words, and repetitive speech. Therefore, we adopted an index to evaluate the lexical richness.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eTemporal features\u003c/h2\u003e \u003cp\u003eThe temporal features were measured manually. Table S6 lists the temporal features and their corresponding descriptions. The reaction time (i.e., the time between the monitor display the duration from speech onset to speech offset for and the start of the response) and speech time (i.e., the duration from speech onset to speech offset for each of the 16 tasks) were extracted. In addition, the speed and linguistic features \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:V\\left(N\\right)\\)\u003c/span\u003e\u003c/span\u003e were combined to define the two types of features related to speech speed.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eFeature selection\u003c/h2\u003e \u003cp\u003eConsidering that a total of 405 features were extracted, which was greater than the number of training data points, overlearning could occur. Therefore, feature selection was performed using a combination of filter and wrapper methods before training the model. The filter method selected characteristics using statistics such as variance and correlation coefficients \u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. The wrapper method used a subset of features to evaluate the performance of the model and select optimal features \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e. Although the filter method could select features in a short time, it could not consider the relationships between the features. In contrast, the wrapper method considered the relationships between features. Therefore, this study proposed a feature selection method that adapted the wrapper method to the features selected by the filter method. First, the filter method was applied to remove features that exhibited no variance across samples and those that showed high pairwise correlations (correlation coefficient\u0026thinsp;\u0026gt;\u0026thinsp;0.80). Second, in the wrapper method, a forward stepwise procedure guided by the Akaike Information Criterion was employed to sequentially add features and identify those contributing to the classification. The number of selected features was limited to less than the number of training data points (maximum of 157 or 158) to prevent model overtraining.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003ePWEL model\u003c/h2\u003e \u003cp\u003ePWEL is an ensemble-learning method that optimizes task combinations and weighting based on the classification performance of each disease subtype. It considers the continuity among subtypes. In clinical settings, the difficulty of data collection varies according to subtype. For rare subtypes, data scarcity makes it difficult to build machine learning models. Therefore, it is preferable to treat the subtypes as a single group and compare the classification performance between patients and HCs. PWEL evaluates the classification performance of each subtype. It assigns higher weights to weak learners with better performances. This approach effectively utilizes information from rare subtypes. The number of subtypes \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:d\\)\u003c/span\u003e\u003c/span\u003e and speech task types \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e can be adjusted. Weights are computed based on \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:d\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e. PWEL is composed of two methods: (i) under-bagging, which combines undersampling and bagging, and (ii) adaptive task selection, which consists of an inner weak learner, selecting tasks, and weighted ensemble learning (where the number of selected tasks is denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:T\\)\u003c/span\u003e\u003c/span\u003e). As a weak learner, a binary classifier (e.g., SVM, decision tree, and gradient boosting) was considered. The model with the highest discrimination accuracy was adopted for the adaptive task selection.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eAddressing unbalanced data by under-bagging\u003c/h2\u003e \u003cp\u003eThere were imbalances in the clinical data between the disease and non-disease groups. Directly inputting such imbalanced data into a machine learning model makes it difficult for the model to adequately learn from rare cases and correctly identify them. A recent systematic review on imbalanced medical datasets\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e shows that undersampling is frequently adopted in medical research because it avoids the risk of generating unrealistic synthetic data inherent to oversampling. Nevertheless, as undersampling inevitably discards part of the majority class, it is recommended to combine it with ensemble methods (e.g., bagging) to mitigate information loss and improve robustness. Based on this rationale, we applied an under-bagging strategy to balance the participant groups: the ALS group was undersampled to match the number of FTD samples, and the HC group was undersampled to match the combined total of the FTD and ALS groups.\u003c/p\u003e \u003cp\u003eTo assess the robustness of our analysis with respect to sample size, we employed both 5-fold stratified CV. In 5-fold CV, each training fold involved 157 participants (59 patients with ALS, 6 patients with PNFA, 6 patients with bvFTD, 12 patients with SD, and 74 HCs), whereas the corresponding test folds involved 40 participants (15 patients with ALS, 1 patient with PNFA, 2 patients with bvFTD, 3 patients with SD, and 19 HCs). Stratified 5-fold CV was conducted using mutually exclusive sample partitions across folds to mitigate data leakage and enhance the generalizability of the model. The slight variation in fold size reflects the inherent consequences of stratified partitioning with limited class frequencies and does not introduce systematic bias into model evaluation.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eAdaptive task selection\u003c/h2\u003e \u003cp\u003eAdaptive task selection dynamically calculated the ensemble weights to integrate the outputs of multiple models. The details of this approach were as follows.\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003e1.\u003c/b\u003e Adaptive task selection selects the optimal combination of speech tasks and their weights through CV within the training data. It applies leave-one-out cross-validation (LOOCV), in which one sample is left out, and the remaining data are used for training. This process is referred to as inner CV. In the current study, adaptive task selection was used for disease groups, including the FTD subtypes of PNFA, bvFTD, and SD and ALS. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e shows the detailed computational process.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003ea.\u003c/b\u003e Within the training data, the detection performance (macro-F1) is calculated for each disease subtype \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{d}_{j}\\:\\in\\:\\:D\\)\u003c/span\u003e\u003c/span\u003e in 16 tasks \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:({t}_{1},\\dots\\:,\\:{t}_{16})\\)\u003c/span\u003e\u003c/span\u003e. These values form the performance matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mathbf{F}\\:\\in\\:\\:{\\mathbf{R}}^{16\\times\\:4}\\)\u003c/span\u003e\u003c/span\u003e. Next, for each disease subtype \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{d}_{j}\\)\u003c/span\u003e\u003c/span\u003e, the tasks are sorted in descending order of macro-F1 scores, creating the matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\mathbf{S}}_{j\\:}=\\:\\text{a}\\text{r}\\text{g}\\text{s}\\text{o}\\text{r}\\text{t}\\left({\\mathbf{F}}_{:,j}\\right)\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eb.\u003c/b\u003e Based on \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{S}\\)\u003c/span\u003e\u003c/span\u003e, the weight of each task at threshold \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:k\\)\u003c/span\u003e\u003c/span\u003e is computed. The weight \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{w}_{l,k}\\:(l=\\:1,\\dots\\:,16)\\)\u003c/span\u003e\u003c/span\u003e is defined as the number of times task \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{t}_{l}\\)\u003c/span\u003e\u003c/span\u003e appears within the top \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:k\\)\u003c/span\u003e\u003c/span\u003e rankings across all disease subtypes.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003cdiv id=\"Eque\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Eque\" name=\"EquationSource\"\u003e\n$$\\:{w}_{l,k}=\\sum\\:_{{d}_{j}\\in\\:\\:D}\\mathbf{I}(\\text{r}\\text{a}\\text{n}\\text{k}\\left({t}_{l},\\:{d}_{j}\\right)\\leqq\\:k),$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mathbf{I}\\)\u003c/span\u003e\u003c/span\u003e is an indicator function, and rank\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{(t}_{l},\\:{d}_{j})\\)\u003c/span\u003e\u003c/span\u003e represents the rank of task \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{t}_{l}\\)\u003c/span\u003e\u003c/span\u003e for disease subtype \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{d}_{j}\\)\u003c/span\u003e\u003c/span\u003e. Using these weights, the prediction value \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{P}_{k}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e for a sample \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e is computed as follows:\u003cdiv id=\"Equf\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equf\" name=\"EquationSource\"\u003e\n$$\\:{P}_{k}\\left(x\\right)=\\:\\frac{{\\sum\\:}_{\\:l=\\:1}^{16}{w}_{l,\\:k}{f}_{l}\\left(x\\right)}{{\\sum\\:}_{l\\:=\\:1}^{16}{w}_{l,\\:k}}$$\u003c/div\u003e\u003c/div\u003e,\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{f}_{l}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the prediction value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e for task \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:l\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cb\u003ec.\u003c/b\u003e The optimal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{k}^{*}\\)\u003c/span\u003e\u003c/span\u003e is determined by performing LOOCV on the training data. The macro-F1 score is calculated for each \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:k\\)\u003c/span\u003e\u003c/span\u003e, and the value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:k\\)\u003c/span\u003e\u003c/span\u003e that maximizes the performance is selected.\u003cdiv id=\"Equg\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equg\" name=\"EquationSource\"\u003e\n$$\\:{k}^{*}={argmax}_{k\\in\\:\\left\\{1,\\dots\\:,16\\right\\}}\\text{m}\\text{a}\\text{c}\\text{r}\\text{o}‐\\text{F}1\\left({P}_{k}\\right).$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003e2.\u003c/b\u003e Weighted ensemble learning combines the predicted values of each task using the weights computed during adaptive task selection. The final prediction for a test sample \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{x}_{test}\\)\u003c/span\u003e\u003c/span\u003e is given by the following equation, using the optimal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{k}^{*}\\)\u003c/span\u003e\u003c/span\u003e:\u003cdiv id=\"Equh\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equh\" name=\"EquationSource\"\u003e\n$$\\:\\widehat{y}=\\:\\left\\{\\begin{array}{c}1\\:if\\:{P}_{{k}^{*}}\\left({x}_{test}\\right)\\geqq\\:\\theta\\:\\\\\\:0\\:otherwise\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\end{array}\\right.$$\u003c/div\u003e\u003c/div\u003e.\u003c/p\u003e \u003cp\u003eFor example, in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, if \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{k}^{*}=\\:2\\)\u003c/span\u003e\u003c/span\u003e, the task weights are \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{w}_{\\text{5,2}}\\:=\\:3,\\:{w}_{\\text{1,2}}=\\:2,\\:{w}_{\\text{4,2}}=\\:2,\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{w}_{\\text{3,2}}=\\:1\\)\u003c/span\u003e\u003c/span\u003e, while the weights of other weak learners are set to 0. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\theta\\:\\)\u003c/span\u003e\u003c/span\u003e represented the classification threshold, which was determined from the ROC curve constructed using the LOOCV-based predictions on the training data as the threshold that maximized the Youden index (TPR - FPR).\u003c/p\u003e \u003cp\u003eThis experiment evaluated the classification performances of weak learner and ensemble learning. In the weak learner method, SVM, random forest, and LightGBM were applied to identify which model performed the best. Table S7ฏ lists the parameters, search ranges, and methods for each model. In this experiment, the number of bagging cycles was set to nine. For ensemble learning, PWEL was compared with the majority voting method to confirm its effectiveness. Moreover, to examine the effectiveness of the two methods employed in PWEL, we compared their accuracy with and without (i) under-bagging and (ii) adaptive task selection.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eContinuous variables were compared using the Mann\u0026ndash;Whitney U test, Kruskal-Wallis test, chi-square test, or Fisher\u0026rsquo;s exact test, as appropriate. ROC analyses were performed to verify the diagnostic accuracy (AUC, sensitivity, specificity, accuracy, and macro-F1) of PWEL for FTD and ALS, and the optimal cut-off values for classification were determined based on the Youden index \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e,\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. As FTD was divided into three subtypes, sensitivity was evaluated separately for each subtype. All statistical analyses were performed using JMP Pro version 16.0 (SAS Institute Inc., Cary, NC, USA) and Python (version 3.8.12) and scikit-learn (version 1.0.2). Statistical significance was set at P\u0026thinsp;\u0026lt;\u0026thinsp;0.05 except for specific instances mentioned above.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding Declaration\u003c/h2\u003e \u003cp\u003eThis work was supported in part by the Ministry of Education, Culture, Sports, Science and Technology-Japan, Grant-in-Aid for Scientific Research under grant #JP24H00741, and part by the commissioned research by National Institute of Information and Communications Technology (NICT), JAPAN.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eYI, RO, SK, MM, and HW contributed to the conception and design of the study. RO, MM, MS, HW, TS, AO, KH, SS, YM, MH, MK, MI, and GS contributed to the acquisition of the data. YI, RO, KI, SK, and HW contributed the analysis of the data. YI, RO, SK, and HW contributed to the interpretation of the data. YI, RO, SK, and HW contributed to the draft of the article. YI, RO, SK, and HW revised the manuscript critically for important intellectual content. All authors read and approved the final version of the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated during and/or analysed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBurrell, J. R., Kiernan, M. C., Vucic, S. \u0026amp; Hodges, J. R. Motor neuron dysfunction in frontotemporal dementia. \u003cem\u003eBrain\u003c/em\u003e \u003cb\u003e134\u003c/b\u003e, 2582\u0026ndash;2594 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIazzolino, B. et al. Validation of the revised classification of cognitive and behavioural impairment in ALS. \u003cem\u003eJ. Neurol. Neurosurg. Psychiatry\u003c/em\u003e. \u003cb\u003e90\u003c/b\u003e, 734\u0026ndash;739 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRatnavalli, E., Brayne, C., Dawson, K. \u0026amp; Hodges, J. R. The prevalence of frontotemporal dementia. \u003cem\u003eNeurology\u003c/em\u003e \u003cb\u003e58\u003c/b\u003e, 1615\u0026ndash;1621 (2002).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMasrori, P. \u0026amp; Van Damme, P. Amyotrophic lateral sclerosis: a clinical review. \u003cem\u003eEur. J. Neurol.\u003c/em\u003e \u003cb\u003e27\u003c/b\u003e, 1918\u0026ndash;1929 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShinagawa, S., Catindig, J. A., Block, N. R., Miller, B. L. \u0026amp; Rankin, K. P. When a little knowledge can be dangerous: False-positive diagnosis of behavioral variant frontotemporal dementia among community clinicians. \u003cem\u003eDement. Geriatr. Cogn. Disord\u003c/em\u003e. \u003cb\u003e41\u003c/b\u003e, 99\u0026ndash;108 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCoppieters, R. et al. A systematic review of the quantitative markers of speech and language of the frontotemporal degeneration spectrum and their potential for cross-linguistic implementation. \u003cem\u003eNeurosci. Biobehav Rev.\u003c/em\u003e \u003cb\u003e167\u003c/b\u003e, 105909 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCho, S. et al. Automatic classification of AD pathology in FTD phenotypes using natural speech. \u003cem\u003eAlzheimer\u0026rsquo;s Dement.\u003c/em\u003e \u003cb\u003e20\u003c/b\u003e, 3416\u0026ndash;3428 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZimmerer, V. C. et al. Automated profiling of spontaneous speech in primary progressive aphasia and behavioral-variant frontotemporal dementia: an approach based on usage-frequency. \u003cem\u003eCortex\u003c/em\u003e \u003cb\u003e133\u003c/b\u003e, 103\u0026ndash;119 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeeley, W. W. et al. The natural history of temporal variant frontotemporal dementia. \u003cem\u003eNeurology\u003c/em\u003e \u003cb\u003e64\u003c/b\u003e, 1384\u0026ndash;1390 (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNevler, N. et al. Automated analysis of natural speech in amyotrophic lateral sclerosis spectrum disorders. \u003cem\u003eNeurology\u003c/em\u003e \u003cb\u003e95\u003c/b\u003e, E1629\u0026ndash;E1639 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVogel, A. P. et al. Motor speech signature of behavioral variant frontotemporal dementia: Refining the phenotype. \u003cem\u003eNeurology\u003c/em\u003e \u003cb\u003e89\u003c/b\u003e, 837\u0026ndash;844 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGong, Y. et al. Exploring emotion and emotional variability as digitalbiomarkers in frontotemporal dementia speech. \u003cem\u003eIEEE Access.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 71419\u0026ndash;71432 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSagi, O. \u0026amp; Rokach, L. Ensemble learning: a survey. \u003cem\u003eWiley Interdiscip Rev. Data Min. Knowl. Discov\u003c/em\u003e. \u003cb\u003e8\u003c/b\u003e, 1\u0026ndash;18 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKe, X., Mak, M. W. \u0026amp; Meng, H. M. Automatic selection of spoken language biomarkers for dementia detection. \u003cem\u003eNeural Netw.\u003c/em\u003e \u003cb\u003e169\u003c/b\u003e, 191\u0026ndash;204 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuffy, J. R., Peach, R. K. \u0026amp; Strand, E. A. Progressive apraxia of speech as a sign of motor neuron disease. \u003cem\u003eAm. J. Speech-Language Pathol.\u003c/em\u003e \u003cb\u003e16\u003c/b\u003e, 198\u0026ndash;208 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eda Cunha, P. L. et al. Automated free speech analysis reveals distinct markers of Alzheimer\u0026rsquo;s and frontotemporal dementia. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e19\u003c/b\u003e, 1\u0026ndash;19 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarrard, P., Rentoumi, V., Gesierich, B., Miller, B. \u0026amp; Gorno-Tempini, M. L. Machine learning approaches to diagnosis and laterality effects in semantic dementia discourse. \u003cem\u003eCortex\u003c/em\u003e \u003cb\u003e55\u003c/b\u003e, 122\u0026ndash;129 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFraser, K. C. et al. Automated classification of primary progressive aphasia subtypes from narrative speech transcripts. \u003cem\u003eCortex\u003c/em\u003e \u003cb\u003e55\u003c/b\u003e, 43\u0026ndash;60 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChristidi, F. et al. Gray matter and white matter changes in non-demented amyotrophic lateral sclerosis patients with or without cognitive impairment: a combined voxel-based morphometry and tract-based spatial statistics whole-brain analysis. \u003cem\u003eBrain Imaging Behav.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 547\u0026ndash;563 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrambati, S. M., Ogar, J., Neuhaus, J. \u0026amp; Miller, B. L. Gorno-Tempini, M. L. Reading disorders in primary progressive aphasia: A behavioral and neuroimaging study. \u003cem\u003eNeuropsychologia\u003c/em\u003e \u003cb\u003e47\u003c/b\u003e, 1893\u0026ndash;1900 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKertesz, A. \u003cem\u003eThe Western Aphasia Battery\u003c/em\u003e (Grune \u0026amp; Stratton, 1982).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKertesz, A. The Western Aphasia Battery: a systematic review of research and clinical applications. \u003cem\u003eAphasiology\u003c/em\u003e \u003cb\u003e36\u003c/b\u003e, 21\u0026ndash;50 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHanley, J. A. \u0026amp; McNeil, B. J. The meaning and use of the area under a receiver operating characteristic (ROC) curve. \u003cem\u003eRadiology\u003c/em\u003e \u003cb\u003e143\u003c/b\u003e, 29\u0026ndash;36 (1982).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAkobeng, A. K. Understanding diagnostic tests 3: Receiver operating characteristic curves. \u003cem\u003eActa Paediatr.\u003c/em\u003e \u003cb\u003e96\u003c/b\u003e, 644\u0026ndash;647 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJeancolas, L. et al. C,o,mparison of telephone recordings and professional microphone recordings for early detection of Parkinson\u0026rsquo;s disease, using mel-frequency cepstral coefficients with Gaussian mixture models to cite this version: HAL Id : hal-02474486 (2020). Comparison of Tele.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRodriguez, J. D., Perez, A. \u0026amp; Lozano, J. A. Sensitivity analysis of k-fold cross validation in prediction error estimation. \u003cem\u003eIEEE Trans. Pattern Anal. Mach. Intell.\u003c/em\u003e \u003cb\u003e32\u003c/b\u003e, 569\u0026ndash;575 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNeary, D. et al. Frontotemporal lobar degeneration: a consensus on clinical diagnostic criteria. \u003cem\u003eNeurology\u003c/em\u003e \u003cb\u003e51\u003c/b\u003e, 1546\u0026ndash;1554 (1998).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGorno-Tempini, M. L. et al. Classification of primary progressive aphasia and its variants. \u003cem\u003eNeurology\u003c/em\u003e \u003cb\u003e76\u003c/b\u003e, 1006\u0026ndash;1014 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRascovsky, K. et al. Sensitivity of revised diagnostic criteria for the behavioural variant of frontotemporal dementia. \u003cem\u003eBrain\u003c/em\u003e \u003cb\u003e134\u003c/b\u003e, 2456\u0026ndash;2477 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrooks, B. R., Miller, R. G., Swash, M. \u0026amp; Munsat, T. L. El Escorial revisited: revised criteria for the diagnosis of amyotrophic lateral sclerosis. \u003cem\u003eAmyotroph. Lateral Scler.\u003c/em\u003e \u003cb\u003e1\u003c/b\u003e, 293\u0026ndash;299 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFolstein, M. F., Folstein, S. E. \u0026amp; McHugh, P. R. Mini-mental state. A practical method for grading the cognitive state of patients for the clinician. \u003cem\u003eJ. Psychiatr Res.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 189\u0026ndash;198 (1975).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMioshi, E., Dawson, K., Mitchell, J., Arnold, R. \u0026amp; Hodges, J. R. The Addenbrooke\u0026rsquo;s Cognitive Examination Revised (ACE-R): a brief cognitive test battery for dementia screening. \u003cem\u003eInt. J. Geriatr. Psychiatry\u003c/em\u003e. \u003cb\u003e21\u003c/b\u003e, 1078\u0026ndash;1085 (2006).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKertesz, A. \u003cem\u003eAphasia and associated disorders: taxonomy, localization, and recovery\u003c/em\u003e (Grune \u0026amp; Stratton, 1979).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFranchignoni, F., Mora, G., Giordano, A., Volanti, P. \u0026amp; Chi\u0026ograve;, A. Evidence of multidimensionality in the ALSFRS-R Scale: a critical appraisal on its measurement properties using Rasch analysis. \u003cem\u003eJ. Neurol. Neurosurg. Psychiatry\u003c/em\u003e. \u003cb\u003e84\u003c/b\u003e, 1340\u0026ndash;1345 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRooney, J., Burke, T., Vajda, A., Heverin, M. \u0026amp; Hardiman, O. What does the ALSFRS-R really measure? A longitudinal and survival analysis of functional dimension subscores in amyotrophic lateral sclerosis. \u003cem\u003eJ. Neurol. Neurosurg. Psychiatry\u003c/em\u003e. \u003cb\u003e88\u003c/b\u003e, 381\u0026ndash;385 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoll, S. F. Suppression of acoustic noise in speech using spectral subtraction. \u003cem\u003eIEEE Trans. Acoust.\u003c/em\u003e \u003cb\u003e27\u003c/b\u003e, 113\u0026ndash;120 (1979).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchuller, B., Steidl, S. \u0026amp; Batliner, A. The INTERSPEECH 2009 emotion challenge. \u003cem\u003eProc. Annu. Conf. Int. Speech Commun. Assoc. INTERSPEECH\u003c/em\u003e 312\u0026ndash;315 (2009). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.21437/interspeech.2009-103\u003c/span\u003e\u003cspan address=\"10.21437/interspeech.2009-103\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchuller, B. et al. A Survey on perceived speaker traits: personality, likability, pathology, and the first challenge. \u003cem\u003eComput. Speech Lang.\u003c/em\u003e \u003cb\u003e29\u003c/b\u003e, 100\u0026ndash;131 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKudo, T. Applying conditional random fields to Japanese morphological analysis. \u003cem\u003eProc. 2004 Conf. Empir. Methods Nat. Lang. Process.\u003c/em\u003e 230\u0026ndash;237 (2004).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSIMPSON, E. H. Measurement of Diversity. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e163\u003c/b\u003e, 688\u0026ndash;688 (1949).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuyon, I. \u0026amp; Elisseeff, A. An introduction to variable and feature selection. \u003cem\u003eJ. Mach. Learn. Res.\u003c/em\u003e \u003cb\u003e3\u003c/b\u003e, 1157\u0026ndash;1182 (2003).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlum, A. L. \u0026amp; Langley, P. Selection of relevant features and examples in machine learning. \u003cem\u003eArtif. Intell.\u003c/em\u003e \u003cb\u003e97\u003c/b\u003e, 245\u0026ndash;271 (1997).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalmi, M., Atif, D., Oliva, D., Abraham, A. \u0026amp; Ventura, S. Handling imbalanced medical datasets: review of a decade of research. \u003cem\u003eArtif. Intell. Rev.\u003c/em\u003e \u003cb\u003e57\u003c/b\u003e, 273 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYouden, W. J. Index for rating diagnostic tests. \u003cem\u003eCancer\u003c/em\u003e \u003cb\u003e3\u003c/b\u003e, 32\u0026ndash;35 (1950).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFluss, R., Faraggi, D. \u0026amp; Reiser, B. Estimation of the Youden Index and its associated cutoff point. \u003cem\u003eBiometrical J.\u003c/em\u003e \u003cb\u003e47\u003c/b\u003e, 458\u0026ndash;472 (2005).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, Frontotemporal dementia, Amyotrophic lateral sclerosis, Diagnostic screening tool, Reading task, Speech features","lastPublishedDoi":"10.21203/rs.3.rs-7917087/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7917087/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eWe developed a novel machine learning model named performance-weighted ensemble learning (PWEL) to detect frontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS) using speech from word and sentence reading tasks. Overall, 197 participants (30 with FTD, 74 with ALS, and 93 healthy controls) were enrolled. The brief oral reading tasks consisted of 16 materials adapted from the Western Aphasia Battery and took 1 min to complete. For each task, 405 speech features (i.e., 384 acoustic, 17 linguistic, and 4 temporal) were extracted. Five-fold cross-validation highlighted acoustic features\u0026mdash;especially MFCCs\u0026mdash;as key discriminators of FTD and ALS from healthy controls. PWEL, which included under-bagging and adaptive task selection, achieved an area under the curve of 0.840, a sensitivity of 0.828, and an overall accuracy of 0.792. The sensitivity for FTD subtypes was 0.700-0.933, and that for ALS reached 0.812. Our proposed model surpassed both a simple ensemble learning model and task-wise classification in overall sensitivity, accuracy, and macro-F1. Its robust generalizability was demonstrated by consistent performance across stratified cross-validation folds and three recording sites. In conclusion, PWEL is practical, scalable, and generalizable across multiple sites, making it a promising tool for detecting neurodegenerative disorders.\u003c/p\u003e","manuscriptTitle":"Performance-weighted ensemble learning for detecting patients with FTD and ALS from short reading tasks","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-23 09:33:02","doi":"10.21203/rs.3.rs-7917087/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-02-18T17:27:40+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-10-31T10:15:07+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-10-29T05:38:19+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-10-29T05:37:50+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-10-21T13:30:17+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2c3c14fe-8bf6-4137-8a76-a97ea68c0998","owner":[],"postedDate":"February 23rd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":63282200,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":63282201,"name":"Health sciences/Neurology"},{"id":63282202,"name":"Biological sciences/Neuroscience"}],"tags":[],"updatedAt":"2026-05-15T13:38:33+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-23 09:33:02","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7917087","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7917087","identity":"rs-7917087","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00