Deep Learning Driven Evaluation of MR-guided Focused Ultrasound Ablation.

OA: closed
AI-generated summary by qwen3.7-flash, 2026-09-09

This study developed a deep learning framework using multi-parametric MRI to predict MRgFUS treatment efficacy in VX2 tumor model rabbits, achieving improved boundary accuracy with principal component analysis augmentation.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by qwen3.7-flash, 2026-09-09 · read from full text

This study developed a deep learning pipeline to predict the non-perfused volume following magnetic resonance-guided focused ultrasound ablation in a rabbit VX2 tumor model. Researchers utilized pre- and post-treatment multi-parametric MRI sequences, including T1-weighted, T2-weighted, and quantitative maps, to train neural networks against day-3 contrast-enhanced reference standards. The results demonstrated that these imaging biomarkers could accurately assess treatment efficacy without relying solely on real-time thermal monitoring or immediate post-procedural contrast enhancement. Relevance to endometriosis: listed as one indication for GnRH antagonists, though the paper's main focus is uterine fibroids.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

ObjectiveMagnetic resonance-guided focused ultrasound (MRgFUS) thermal therapy is a promising incisionless procedure for breast cancer treatment. for assessing treatment efficacy. Current approaches rely on thermal and vascular MRI-derived biomarkers to assess treatment efficacy. However, these techniques are not sufficiently accurate for real-time in vivo assessment. To address this challenge, we propose a deep-learning framework that utilizes multi-parametric MRI to predict treatment efficacy at the time of treatment.MethodsThe presented model utilizes qualitative T1- and T2-weighted images, MR temperature-derived metrics and quantitative T1 and T2 parametric maps to predict treatment efficacy. Model robustness was enhanced via extensive data augmentation techniques. To train and validate this approach, an MRgFUS ablation study was conducted on a dataset of VX2 tumor model rabbits (N=12). The predictive power of the multi-parametric MRI model was evaluated by comparing the trained model's predictions against the three-day post-treatment non-perfused volume.ResultsThe deep learning based biomarker trained with traditional augmentations (random rigid rotations, random cropping, label shifts and Gaussian noise) showed promising performance (Dice: 0.62, MDA: 3.7 mm). Adding principal component analysis based augmentation further refined boundary accuracy (Dice: 0.64, MDA: 3.0 mm).ConclusionsThe results highlight the potential of a deep learning multi-parametric MRI framework for accurately predicting MRgFUS treatment efficacy. While improvements would be required to reduce the acquisition time of the multi-parametric MRI protocol, inclusion of quantitative MRI data offers a pathway toward more accurate and immediate evaluation of MRgFUS therapy outcomes.
Full text 57,261 characters · extracted from pmc-nxml · 5 sections · click to expand

Methods

To develop a deep learning-based imaging biomarker for predicting MRgFUS treatment efficacy, a structured pipeline was created involving three key steps: (A) MRI data acquisition before, during, and after thermal ablation; (B) segmentation of the NPV at day 3 post-treatment as a reference standard; and (C) registration of multi-parametric MRI sequences across time points. These aligned datasets were then used to train and evaluate different neural networks architectures for accurate NPV prediction. The rest of this section describes each step of the workflow in detail. All experiments were conducted with the approval of the local Institutional Animal Care and Use Committee. A New Zealand White rabbit VX2 tumor model was utilized (N=12). VX2 tumor cells (approximately 1 million cells in 50% media/Matrigel solution) were injected into the quadriceps muscle of each rabbit. When the tumors reached approximately 2 cm in length, approximately 10 days after injection, they underwent MRgFUS thermal ablation in a 3T MRI scanner (PrismaFIT, Siemens, Erlangen, Germany). The ablation was conducted using an MRgFUS system (Image Guided Therapy, Inc.), equipped with a 256-element phased-array transducer (Imasonic, Voray-sur-l’Ognon, France, f=950kHz, 13 cm focal length, 15.4 cm diameter aperture. During the procedure, the animals were intubated, anesthetized with 1.0–2.5% isoflurane, and continuously monitored for vital signs. The table that supported the rabbit had an integrated 17 cm single-loop MR receive coil that encircled the targeted quadriceps. The parameters for MR image acquisition are detailed in Table I . MRI guidance facilitated tumor targeting (T2-weighted fat saturated sequence), real-time monitoring during treatment (MRTI), and pre- and post-procedural evaluation. The rabbit was positioned on the treatment table as illustrated in Figure 1 . High-resolution (0.5 mm isotropic) non-contrast T1-weighted images of the treated leg were acquired at multiple time points during the procedure to facilitate image registration including after initial positioning, before and immediately after ablation, and at the conclusion of the post-ablation imaging sequence. Volumetric Interpolated Breath-hold Examination (VIBE) is a 3D spoiled gradient-echo sequence that provides high-resolution T1-weighted images with excellent soft-tissue contrast and rapid volumetric coverage. Unlike conventional 2D T1-weighted sequences, VIBE acquires an entire 3D volume during a single breath-hold or short acquisition window, allowing isotropic voxel reconstruction and reduced motion artifacts. In this study, VIBE was used for contrast-enhanced T1-weighted imaging to delineate the NPV immediately after MRgFUS ablation. Representative images of the MR temperature imaging and T2 map data for one of the subjects, overlaid on the CE-T1w image obtained at the conclusion of the MRgFUS ablation, is shown in Fig. 2 . 3D MRTI data was acquired and 2D T1 and T2 maps were acquired that were interleaved to provide quasi-3D coverage, but in both cases the field of view was limited in the slice direction. Coronal multi-parametric MR images (T1w, T2w, T1 maps, T2 maps) of the targeted leg (shown in Figure 1 ) were obtained at multiple time points: pre-ablation and post-ablation. After the final sonication, CE-T1w MR images were captured following intravenous administration of gadolinium contrast agent (1 mL ProHance), after which the animals were allowed to recover. As transient effects of ablation, including edema and latent apoptosis, can persist for up to 72 hours post treatment [ 34 ], [ 35 ], the final MRgFUS treatment assessment was obtained three days after the MRgFUS procedure. At this time point, the rabbit was again anesthetized and placed on the MRgFUS table, and high-resolution T1w and CE-T1w images were again acquired. The rabbit was then euthanized. The hypointense region at the treatment site was considered, the NPV and measured through semiautomatic segmentation (Seg3D [ 36 ], Scientific Computing and Imaging Institute, University of Utah) of the CE-T1w images obtained at day of ablation (day 0) and at the 3-day post-treatment time point (day 3). This process involved thresholding the CE-T1w image intensity to identify the non-enhancing region, followed by manual adjustments. The NPV segmentations at day 3 were considered the clinical prediction of tissue necrosis and were used as the training labels for the neural network. Day-3 contrast-enhanced T1-weighted (CE-T1w) non-perfused volume (NPV) was used as the treatment efficacy reference standard in this study. While we acknowledge that Day-3 CE-T1w NPV is a surrogate and not a direct histological measure of necrosis, prior work in both VX2 rabbit and large-animal models has consistently demonstrated a strong correlation between non-perfused regions on CE-T1w MRI and histologically confirmed necrosis approximately 2–3 days post-treatment[ 7 ], [ 37 ], [ 38 ]. Specifically, Hynynen et al.[ 38 ] conducted a 28-rabbit brain MRgFUS experiment and performed MRI at 2 h, 2 d, 10 d, and 23 d, followed by whole-brain histology. They observed that non-enhancing zones on T1w MRI at 2 d corresponded precisely to the coagulated necrotic core seen microscopically, with enhancement confined to the viable rim, confirming MRI–histology alignment within 2 days of treatment. Similarly, Breen et al.[ 39 ] and Chung et al.[ 40 ] demonstrated that MR-visible lesion boundaries and histological necrosis volumes reach optimal concordance within 2–3 days after thermal therapy. To evaluate the reliability of the semi-automatic NPV segmentation used as the reference standard, inter-observer variability was assessed between two independent raters who manually refined the segmented NPV masks for all rabbit subjects. Each rater performed the segmentation using identical imaging data and guidelines within the Seg3D [ 36 ] software environment. Quantitative agreement between the two expert segmentations was evaluated using both the Dice Similarity Coefficient (Dice) and the Mean Distance to Agreement (MDA). Because MRgFUS thermal treatments result in targeted, localized ablation of treatments, it is imperative to achieve accurate alignment of all imaging data between all time points to evaluate the predictive power of the multiparametric MRI contrasts. A three-step image registration approach, comprised of rigid transformation, landmark-driven gradient flow, and intensity-based volume-penalizing registration, was employed to achieve accurate alignment between all time points while preserving volume. Three time points were registered for this analysis. Two time points were identified on the MRgFUS ablation day (day 0 pre- and post-ablation) and the remaining time point was obtained three days after the MRgFUS ablation (day 3). All images were ultimately mapped to the day 3 time point. All deformation fields were computed using the T1w 3D VIBE images and applied to all other image contrasts at the same timepoint. The registration pipeline implemented in this study was assembled by combining previously validated methods: the first two steps were adopted from Zimmerman et al.[ 5 ], while the final step was based on the approach introduced by Hinkle et al.[ 41 ]. The mathematical framework and the equations governing the image registration are presented in the Appendix . Image registration was evaluated using a total of 120 anatomically meaningful landmarks. Three of these landmarks were anatomically consistent and identifiable across all subjects, while an additional seven subject-specific landmarks were identified across all rabbits to capture local anatomical detail that could not be reliably matched across all animals. For every subject, transformed and target landmark coordinates were extracted, and registration accuracy was quantified using the L2 distance between corresponding points. To assess image registration robustness, we performed a bootstrap sampling analysis[ 42 ] in which subsets of 10, 20, up to all 120 landmarks were randomly selected with replacement 1000 times for each subset size. For every iteration, the mean and variability of the landmark-based registration error were computed. All MRI sequences were registered and resampled into the space of the high-resolution 3D T1-weighted (VIBE) images, which served as the common anatomical reference. Registration was performed exclusively using the T1-weighted images, as these primarily encode stable anatomical structure and do not strongly highlight tumor or ablation-related signal changes, thereby minimizing bias in deformation field estimation. The resulting deformation fields were then applied uniformly to all other MRI contrasts acquired at the same time point, including T2-weighted images, quantitative T1 and T2 maps, and MR thermometry–derived metrics. The trained neural network uses these resampled multiparametric volumes as input, enabling consistent voxel-wise integration across sequences while ensuring that any limitations related to anisotropic sampling or restricted coverage reflect acquisition constraints rather than registration inconsistencies. Three neural network architectures were evaluated in this work. In all cases, the network inputs were the registered imaging volumes acquired at the day 0 pre- and post-ablation time points detailed in Table I . The networks were trained in a supervised fashion using the registered day 3 NPV as the ground truth. A U-net segmentation architecture [ 43 ] was used to generate a volumetric treatment efficacy biomarker. The architecture uses 5 encoding decoding stages. The number of output channels at each encoder stage increases as follows: 8 → 16 → 32 → 64 → 128. Each encoder block consists of two 3×3×3 3D convolutional layers with stride 1 and padding 1, followed by batch normalization and ReLU activation ( Fig. 3 ). Max pooling with a 2×2×2 kernel and stride 2 was used between stages to perform downsampling. In the decoder, the channel dimensions after concatenating the skip connections are 256 (128+128), 128 (64+64) and so on. Each decoder block also uses two convolutional layers with batch normalization and ReLU activation. The design used stride-1 convolutions everywhere except during downsampling, ensuring that the spatial resolution only changes at pooling layers. No dropout layers were used. All convolutions had padding equal to 1 to maintain spatial size before pooling. Upsampling was handled by concatenating encoder feature maps at matching resolutions, without explicit transposed convolutions. A 1×1×1 convolution was used at the final layer to project features to a single output channels for treatment efficacy prediction task. The U-Net architecture contains approximately 2.51 million trainable parameters. The 3D U-Net architecture’s [ 44 ] symmetric design and use of skip connections make it highly effective for tasks requiring precise localization, as the network can combine global context from deeper layers with local details from shallower layers. This architecture extends the standard 3D U-Net by incorporating attention gates [ 45 ] to refine skip connections between encoder and decoder. The encoder path consisted of five stages with increasing channel dimensions: 8 → 16 → 32 → 64 → 128. Each encoder block had two consecutive 3×3×3 convolutions with stride 1 and padding 1, followed by batch normalization and ReLU activation. Downsampling was performed by 2×2×2 max pooling. The deepest feature map (enc6) is generated by two convolutional layers after the fifth pooling operation. Attention blocks were applied at the skip connections corresponding to feature levels with 128, 64, 32, and 16 channels, where each attention block used 1×1×1 convolutions to align feature dimensions before calculating attention maps. The decoder path progressively upsampled the feature maps using trilinear interpolation and concatenates the upsampled features with the attended skip connections. The decoder channel dimensions decrease as follows: 256 → 128 → 64 → 32 → 16. Each decoder block again used two 3×3×3 convolutional layers, batch normalization, and ReLU activation. The final output is produced by a 1×1×1 convolution that maps the feature channels (8) to a single output channel, suitable for treatment efficacy prediction. Attention mechanisms guided the network to focus on more relevant regions during feature fusion, improving segmentation performance, particularly in complex volumetric data. The U-Net Attention architecture contains approximately 2.53 million trainable parameters. The use of attention blocks helped the network selectively focus on the most relevant spatial regions during feature fusion. By weighting important features more heavily and suppressing irrelevant background information, attention blocks improved segmentation accuracy, especially in cases where target structures were small, complex, or had low contrast [ 46 ]. The implemented Vision Transformer (ViT) model [ 47 ] leveraged patch embeddings, transformer-based encoders, and a convolutional decoder for feature extraction, transformation, and output generation. The architecture was based on a 3D Vision Transformer (ViT) model. The input 3D volume was divided into non-overlapping patches of size 3×3×3, and each patch is embedded into a 128-dimensional feature space using a 3D convolutional layer with kernel size and stride both set to (3,3,3). After patch embedding, a learnable classification token was appended, and positional embeddings were added to incorporate spatial information. The model includes a dropout layer with a rate of 0.1 applied after positional embedding. Given that typical lesion diameters are on the order of ~2 cm, this patching strategy provides fine-grained local representation relative to lesion size, but may limit the capture of larger-scale spatial context compared to convolutional architectures. The transformer encoder consisted of 12 layers, each with multi-head attention using 2 heads and a feed-forward network with a hidden dimension of 1024. Each transformer block included layer normalization, residual connections, and dropout (0.1). After encoding, the classification token was discarded, and the output features are reshaped back to a 3D volume by transposing and reshaping operations. A 3D transposed convolution (deconvolution) with kernel size (3,3,3) and stride (3,3,3) was used to upsample the features back to the original input size, producing a single-channel output for treatment efficacy prediction task. The Vision Transformer architecture contains approximately 5.2 million trainable parameters. Overall, this model employed the strengths of transformers for global feature extraction and the flexibility of convolutional operations for spatial reconstruction, making it a powerful tool for analyzing and processing volumetric data. The Transformer-based UNET-R architecture [ 48 ] integrates a transformer encoder with a convolutional decoder to capture long-range spatial dependencies in 3D volumetric MRI data. The model takes a multi-channel MRI input volume of size 96 × 96 × 96, where each channel corresponds to a distinct MRI sequence. Instead of relying on conventional convolutional encoders, the input volume is partitioned into non-overlapping 16 × 16 × 16 patches. Each 3D patch is flattened and linearly projected into a 768-dimensional embedding, producing a sequence of patch tokens. Positional embeddings are then added to retain spatial context. The UNETR architecture contains approximately 13.3 million trainable parameters. The resulting token sequence is processed through a stack of 12 transformer encoder layers. Each layer contains a multi-head self-attention module with 12 attention heads and a feedforward MLP block with a hidden size of 3072, both wrapped with layer normalization and residual connections to promote stable training. Feature maps extracted from intermediate transformer layers corresponding to downsampled resolutions at 1/2, 1/4, 1/8, and 1/16 of the input scale are reshaped and connected to the decoder via skip connections. The decoder follows a conventional U-Net structure, employing trilinear interpolation for upsampling, concatenation with transformer-derived features, and two successive 3 × 3 × 3 convolutional layers with instance normalization and ReLU activation at each stage. As spatial resolution increases, the number of feature channels is reduced symmetrically from 128 to 16. A final 1 × 1 × 1 convolution projects the output to a single-channel prediction map, and a sigmoid activation function converts the result into voxel-wise probabilities for treatment efficacy prediction. Network architectures were compared via a repeated-measures ANOVA across architectures by treating cross-validation folds as paired observations. This was followed by paired Wilcoxon signed-rank post-hoc tests. For each architecture, Dice values were extracted on a per–cross-validation fold basis, ensuring identical data splits across models. We then conducted pairwise comparisons between U-Net, U-Net Attention, UNETR, and Vision Transformer models using a single paired test per comparison. To evaluate performance while maximizing data utility, a leave-one-subject-out cross-validation [ 49 ] scheme was implemented with a modified split; in each iteration, the model was trained on data from N=10 subjects, validated on N=1, and tested on the remaining subject (N=1). This strategy ensured that the model’s performance was assessed on completely unseen data while also incorporating a separate validation set for tuning and early stopping. The training process leveraged the binary cross-entropy (BCE) loss function, optimized using the Adam optimizer [ 50 ] with a learning rate of 0.001, allowing for stable and efficient convergence on this relatively small dataset. A batch size of 2 and 3000 training epochs(approximately 9 hours) were used for each model. Early stopping was applied based on convergence of training and validation accuracy, terminating training when the absolute difference between them remained below a defined threshold (1%) for a set number of consecutive epochs. Representative training curves from a single cross-validation fold for all four architectures are shown in Fig. 5 to assess training dynamics and potential overfitting. While all models exhibited stable optimization behavior, the Vision Transformer maintained consistently higher training loss throughout training, whereas the convolutional architectures converged to approximately the same final loss value. The network’s final layer used a sigmoid activation function to generate output probabilities ranging between 0 and 1, representing the likelihood of treatment efficacy. These outputs were thresholded at 0.5 to produce binary classifications, where values ≥ 0.5 indicated predicted efficacy. The deep learning framework was implemented with PyTorch 1.12.1 on an Ubuntu machine with NVIDIA A6000 GPU. Data augmentation was utilized to enhance the diversity of the training set. Random cropping, random rigid rotation, Gaussian noise addition, label shifting and principal component analysis deformable augmentation were used, as detailed below, to better simulate real-world variations, such as anatomical differences and instrumentation noise, allowing the models to generalize better by learning more robust features and reduce overfitting. All the MR image contrasts were resampled to a matrix size of 100 × 100 × 100) and random cropping was applied to these volumes to extract 80 × 80 × 80 subvolumes. For each volume, independent random rotation angles around the cardinal axes, θ x , θ y , θ z , were sampled uniformly from 0° 360°, In this study a maximum rotation angle was determined ( θ max = 180 ° ), allowing for complete rotational invariance across all three axes. The rotation was applied using standard 3D rotation matrices (see [ 51 ]), in the order X → Y → Z. To encourage robust training, artificial independent and identically distributed Gaussian noise with zero mean and a standard deviation of 10% of the maximum value of the signal was added to all images except the MRTI data. Label shifting was achieved by applying uniformly distributed random shifts of 1.5 mm in x, y, and z direction to the labels. Principal Component Analysis (PCA) [ 52 ], [ 53 ] on deformation fields for the different imaging time points (Day 0 pre- and post-ablation, Day 3) was used to generate the random deformations. By projecting the deformation data onto these principal components, dimensionality was reduced while preserving the most significant structural differences. This approach involved first computing deformation fields that register MRI scans taken at each time point and capturing the natural variations in anatomy and imaging conditions. PCA was then applied to these deformation fields to extract dominant modes of variation, enabling the synthesis of new, realistic deformations by sampling and recombining these principal components, this is similar to techniques shown in [ 54 ], [ 55 ]. These synthetic deformation fields were applied to training images, effectively increasing dataset diversity and exposing the network to a broader range of plausible anatomical variations. By learning to predict treatment efficacy under these augmented conditions, the model becomes more robust to real-world variations, improving generalization and performance. This augmentation method is performed as follows: We represented the 3D displacement field defining the deformation associated with each of the subjects as a single flattened vector r i ∈ R 3 H W D , for i = 1 , … , N , (in this study N=12) with three channels (x, y, z) over spatial dimensions H × W × D . The data was centered around the mean μ = 1 N ∑ i = 1 N r i and the collection of the centered displacement fields were stacked in to a matrix X c . SVD was applied X c = U 𝚺 V ⊤ , where the covariance eigenvalues are λ i = σ i 2 / ( N - 1 ) . The top- k components were selected as P = V 1 : k ⊤ ∈ R k × M , and sample weights w ~ 𝒩 ( 0,1 ) . The synthetic relative deformation was reconstructed as r synthetic = P ⊤ 𝚲 1 / 2 w + μ where 𝚲 1 / 2 = diag λ 1 , … , λ k and then reshaped: D synthetic = reshape r synthetic , ( 3 , H , W , D ) Finally, the deformation was applied to the image I ( x ) obtaining I ′ ( x ) = I x + D synthetic ( x ) To quantify how much of the total variability is captured by the first k − components, the percent variance explained by the principal components was computed, as shown in Figure 6 . A value of k = 8 principal components was chosen as it captures approximately 90% of the total variance in the deformation fields. The multiparametric MRI protocol detailed in Table I was acquired at all time points (Day 0 pre- and post-ablation and Day 3). To determine which combination of MR sequences provided the most informative biomarkers for MRgFUS ablation treatment assessment, the combinations detailed in Table II were evaluated in the neural network with these combinations as inputs. The first three combinations (0, 1, and 2) primarily included conventional T1-weighted and T2-weighted images with different variations. Combination 3 expanded on these by incorporating all typically used clinical sequences, while combination 4 included the additional quantitative T1 and T2 maps. Combination 5 removed thermal imaging sequences, focusing solely on anatomical and quantitative MR data. The quantitative MRI metrics, specifically MRTI (maximum temperature projection (MTP), and cumulative thermal dose (CTD)) and T1 and T2 maps were acquired with reduced number of slices due to acquisition time limitations. Because this limited the field of view in the anterior/posterior direction, in some subjects, these quantitative maps did not span the entire NPV resulting from the MRgFUS ablation treatment, as seen in Figure 2 . To evaluate the effect of the limited field of view of the T1/T2 maps and MRTI contrasts, a reduced treatment volume was defined that restricted the spatial regions where all data was available. All data outside this intersection was excluded from analysis. Neural networks were trained using the full acquired volumes, incorporating all acquired multiparametric data regardless of coverage gaps. However, model evaluation was limited to the reduced treatment volume defined by the spatial intersection of all MRI contrasts to ensure consistency across subjects. This strategy enabled us to assess the effect of limited field-of-view at test time, while still leveraging the maximum amount of data during training. Prior to inputting the MRI data into the neural networks, all imaging volumes underwent standardized pre-processing to ensure consistency. First, intensity clipping was applied to each volume by removing the top and bottom 1% of voxel intensities. This step reduced the impact of outliers and enhanced contrast in the remaining dynamic range. Following clipping, the intensities were normalized using min-max scaling to zero-center the data within a fixed range of −1 to 1. This normalization facilitated stable and efficient training by ensuring that all input channels contributed equally to the learning process, particularly important in deep neural networks sensitive to input scale. All training and inference experiments were conducted on a workstation equipped with an NVIDIA A6000 GPU. Model training required approximately 9 hours per cross-validation fold, and each experiment consisted of 12 folds, resulting in a total training time of approximately 108 hours per experiment. This computational cost reflects the optimization of deep neural networks on multi-parametric volumetric MRI data and was incurred entirely offline prior to deployment. Once trained, inference was computationally efficient, with prediction for a single treatment volume completed within 2 seconds, enabling practical use in an intra-procedural setting. The trained models required approximately 6 GB of GPU memory, indicating that deployment is feasible on almost any modern GPU hardware. While absolute runtimes are hardware-dependent, these results demonstrate that the proposed framework is compatible with near real-time clinical decision support. Accuracy of the predicted treatment region was compared to the NPV label generated by an expert segmentation of the Day 3 CE-T1w images. Accuracy of the prediction was compared through both Dice Similarity Coefficient and Mean Distance to Agreement, allowing determination of not only the total volume overlap but also quantified boundaries of prediction, which is representative of margin assessment that is typically performed as part of tumor excision. The Dice score is a similarity index used for discrete data that quantifies how well the predicted region aligns with the ground truth region. A Dice score of 1 indicates a perfect match between the volumes, whereas a score of 0 means no overlap between the two. It is similar to Intersection of Union (IoU) but gives more weight to true positives. It is defined as, Dice = 2 ⋅ | P ∩ G | | P | + | G | where | P | : Total number of pixels in the predicted volume. | G | : Total number of pixels in the ground truth volume. P ∩ G : Number of pixels common to both volumes. In all our experiments, we computed the Dice score between the NPV segmentation predicted by the network to the Day 3 NPV, which is close to the gold standard of histology.[ 12 ] The MDA is calculated by comparing the boundaries of the two segmentations [ 56 ]. The smaller the MDA, the higher the agreement between the segmentations. Unlike IoU or Dice, which focus on pixel overlap, MDA directly measures spatial disagreement in units of distance. MDA = 1 B P + B G ∑ x ∈ B P min y ∈ B G d ( x , y ) + ∑ y ∈ B G min x ∈ B P d ( y , x ) where B P : Boundary pixels of the predicted segmentation. B G : Boundary pixels of the ground truth segmentation. d ( x , y ) : Euclidean distance between boundary pixel x and y . B P : Number of boundary pixels in P . B G Number of boundary pixels in G . MDA of 0 implies perfect agreement between the boundaries. Higher values indicate greater disagreement between boundaries. In all our experiments, we compute the MDA between the NPV segmentation predicted by the network to the Day 3 NPV. The Mean Distance to Agreement (MDA) is reported as the average minimum distance between corresponding surface points of two segmentations. MDA provides a stable estimate of how closely the overall shapes align. While Hausdorff Distance is also a surface-based metric used to quantify spatial agreement between predicted and reference segmentations, it captures different aspects of geometric similarity. Specifically, the Hausdorff Distance represents the maximum surface discrepancy, essentially the largest deviation between any two corresponding points. For this study, MDA was thought to be a more robust measure of the spatial agreement between two contours.

Discussion

The experiments conducted in this study show the potential of non-contrast multiparametric MRI data to predict the treatment efficacy of MRgFUS ablation treatments using different neural network architectures and augmentation schemes. The results ( Table IV ) suggest that vision transformers, although powerful in many image-processing applications, may not be well-suited for MRgFUS treatment response prediction without significant modifications or larger-scale training data. U-Net models generally performed better than vision transformers, with the highest Dice score (0.64) achieved by both the standard and attention U-Net models when combined with PCA-based data augmentation. Additionally, these U-Net models showed lower MDA values (around 3 mm), indicating better spatial agreement. Data augmentation techniques are critical for the development of predictive models for conditions where large data sets are not available, such as was explored in this study. Traditional augmentation methods, including random cropping, random rigid rotations, gaussian noise, and label shifting, led to a clear improvement in both overlap and spatial agreement with the ground truth compared to using no augmentation. Adding PCA-DA further enhanced the model’s performance, achieving the best results among the tested methods. Overall, combining traditional augmentation with PCA-DA proved to be the most effective strategy for improving prediction accuracy for this data set. Notably, the integration of PCA-DA offered additional improvements in precision, making it a valuable enhancement to the augmentation pipeline. Our study underscored the critical role of temperature measurements in predicting treatment efficacy following MRgFUS treatment. MRTI monitoring of thermal changes during ablation provides critical real-time feedback to the provider, but is also critical as an input for the neural network. The experiments revealed that neural networks leveraging temperature data could more accurately predict treatment efficacy ( Table V ) by identifying regions where thermal energy achieved the desired therapeutic impact. Quantitative T1 and T2 maps are essential for evaluating MRgFUS ablation outcomes; however, their acquisition is time-intensive and limited field-of-view constrains full-volume coverage. When the trained network was applied to a reduced volume where all inputs were available, the Dice score improved from 0.64 to 0.72, indicating that performance is primarily limited by spatial coverage rather than model capability. These findings suggest that more complete quantitative coverage, potentially enabled by accelerated multi-parametric approaches such as MR fingerprinting or Multi-Pathway Multi-Echo (MPME), could further enhance prediction accuracy. While the observed improvement may be clinically meaningful for intra-procedural assessment of treatment efficacy, additional studies with expanded volumetric coverage are required to establish the robustness and generalizability of these gains. The U-Net Attention network architecture with both traditional augmentations and PCA-DA outperformed the standard U-Net, UNETR and Vision Transformer-based approaches, demonstrating that the inclusion of attention mechanisms in the U-Net framework enhances the prediction of MRgFUS ablation treatment efficacy. The attention mechanism likely improves feature selection, enabling the model to focus on relevant spatial regions, whereas PCA-DA refines feature extraction by reducing redundancy and improving generalization. This combination resulted in both high prediction accuracy and precise boundary alignment, making it the most effective approach among the tested architectures. As shown in Table VI , convolutional architectures consistently and significantly outperformed the Vision Transformer, with large median Dice improvements, while differences among U-Net based models were small and not statistically significant. These results suggest that strong spatial inductive bias inherent to convolutional designs is better suited for treatment efficacy prediction in MRgFUS imaging than vision transformer based approaches. An inherent limitation of this feasibility study is the modest sample size (N=12); however, this still represents a meaningful dataset for a resource-intensive preclinical MRgFUS model. Each VX2 rabbit experiment requires coordinated tumor implantation, MRgFUS treatment, MRI scanner access, and IACUC-regulated workflows, making large-scale cohort expansion challenging. Within these practical constraints, the dataset captures cross-subject anatomical and treatment variability, though the limited cohort size introduces some model uncertainty and limits generalizability. Although the cohort size in this feasibility study is modest (N = 12), model architectures were intentionally chosen to align capacity with available data. Convolutional U-Net–based models (~2.5M parameters) consistently outperformed higher-capacity transformer architectures, reflecting the advantage of strong spatial inductive bias in small-sample biomedical settings. In contrast, Vision Transformer models (~ 5.5M parameters) exhibited reduced performance, suggesting a mismatch between model complexity and dataset scale. The pure Vision Transformer underperformed likely due to the limited dataset size, as transformers generally require substantially more data than CNNs for stable optimization. ViT subdivides an image into patches and treats each patch as a “token.” This works well when the objects of interest occupy a reasonable area relative to patch size, but when lesions are small or fine-grained, a coarse patch size may mix lesion and background within the same patch diluting the discriminative signal or break a lesion across multiple patches, disrupting spatial continuity. A recent study demonstrated that ViT segmentation performance is highly sensitive to patch size: if the patch dimensions do not match the lesion scale, performance degrades, and determining an “optimal” patch size is both non-trivial and dataset-dependent.[ 57 ] Moreover, ViT’s limited spatial inductive bias unlike the strong locality bias encoded by CNNs further hampers performance in small-data medical settings.[ 58 ] In contrast, CNNs leverage these built-in spatial priors to generalize more effectively when training samples are scarce. Hybrid CNN–transformer architectures, such as UNETR, helped mitigate these limitations. By pairing convolutional feature extraction (capturing local structure) with transformer-based global contextual modeling, UNETR produced performance comparable to the U-Net with attention blocks. This suggests that hybrid designs are better suited for biomedical datasets of this scale than purely transformer-based models. Overall, while transformer models remain promising for MRgFUS treatment-efficacy prediction, they are more sensitive to data availability and architectural choices. Larger datasets, optimized patch configurations, and hybrid architectures are likely necessary for achieving robust and reliable performance in this clinical context. Similarly, in terms of imaging sequences, the best results were obtained when a standard U-Net was trained using T1-weighted, T2-weighted, T1 maps, T2 maps, and MRTI data. This combination yielded a Dice score of 0.62 and an MDA of 3.7 mm, indicating that incorporating both qualitative and quantitative MR-derived metrics significantly enhances predictive performance. The inclusion of T1 and T2 maps provides additional quantitative insights into tissue properties, and temperature data contributes critical information on treatment effects, collectively leading more accurate delineation of NPV. These results highlight that both model architecture and input data selection play a crucial role in optimizing prediction performance, with the U-Net model and multiparametric MR imaging providing the most reliable assessment of MRgFUS treatment efficacy. The inter-observer Dice score (0.85 ± 0.06) reflects annotation consistency for Day-0 NPV and does not represent the difficulty of the prediction task. In contrast, predicting Day-3 NPV from Day-0 imaging is inherently more challenging, with a baseline Dice of ~ 0.57 between Day-0 and Day-3 NPV. Therefore, the model Dice of ~ 0.64 exceeds this intrinsic agreement, indicating meaningful predictive performance. It is important to note that contrast-enhanced T1-weighted non-perfused volume (NPV) represents a surrogate marker of treatment effect based on sustained perfusion deficits rather than a direct histological measure of cellular necrosis. In the acute and subacute post-treatment period, vascular shutdown, edema, hemorrhage, and transient microvascular disruption may lead to non-enhancement in tissue that is not fully necrotic, while conversely, delayed cell death or partial reperfusion may occur in regions that demonstrate contrast enhancement [ 59 ], [ 60 ]. As a result, immediate post-procedural CE-T1w imaging does not reliably reflect durable tissue outcome and is therefore unsuitable as a reference for training outcome-predictive models. To reduce these transient physiological confounds, Day-3 post-treatment CE-T1w imaging was selected as the reference standard, as prior preclinical and large-animal studies have demonstrated strong concordance between non-perfused regions and histologically confirmed necrosis within a 2–3 day window following MR-guided focused ultrasound [ 61 ]–[ 64 ]. The literature[ 12 ] supported a correlation between the Day 3 NPV and histology findings support its use as a reference, but inherent limitations such as delayed tissue response and potential underestimation or overestimation of actual nonviable regions must be acknowledged. In a clinical setting, adapting the image registration component of the pipeline from the preclinical rabbit model to human subjects would require several modifications. Human anatomy presents greater variability and scale, necessitating more advanced, potentially multi-resolution registration strategies. The pipeline discussed here relies on a three-step process - rigid transformation, landmark-driven gradient flow, and intensity-based volume-penalizing registration validated on relatively uniform rabbit datasets. For clinical translation, this approach would need to be adapted to account for larger anatomical differences, more motion artifacts, and diverse imaging conditions encountered in patient data. Human-compatible registration must robustly align longitudinal data while preserving anatomical accuracy. To make this feasible in real time, the registration process must also be automated, optimized for speed, and integrated with clinical MRI software environments. Although the multiparametric protocol employed in this feasibility study was designed to evaluate the relative contribution of diverse MRI contrasts, we acknowledge that the current MRI protocol is not directly practical for routine clinical use. The T1 and T2 mapping scans used in this work required a ~46-minute acquisition time. However, more efficient quantitative MRI techniques [ 65 ]–[ 67 ] have been developed that significantly reduce acquisition time while preserving multiparametric information. Our recent work [ 68 ] incorporating the MPME MRI sequence demonstrates that volumetric multi-contrast data can be acquired in approximately 13 minutes using the MRI coil configuration in this study, compared to over an hour for conventional protocols. This approach enables full-volume coverage of quantitative parameters (T1, T2, T2*, B0, B1+, and proton density [ 69 ]) within a single 3D acquisition, eliminating the need for multiple sequential scans. While further optimization is required for full clinical translation, these advances provide a clear pathway toward integrating quantitative MRI into the MRgFUS workflow without exceeding practical time constraints. Notably, even when restricted to clinically available inputs (T1- and T2-weighted MRI with temperature data), the model achieved Dice = 0.57±0.03 and MDA = 4.16±0.36 mm, indicating that clinically feasible protocols retain strong predictive value. Incorporation of accelerated quantitative sequences such as MPME therefore enables a practical and scalable implementation of the proposed biomarker in real-world settings.

Conclusions

The study demonstrates that combining U-Net architectures, particularly the attention-based variant with PCA-DA augmentation and multi-parametric MRI inputs yields the most accurate and spatially precise predictions of MRgFUS treatment efficacy. Temperature data and quantitative T1/T2 maps were shown to be vital for capturing both the thermal and structural components of ablation, collectively improving Dice scores and spatial agreement metrics. Hybrid CNN–transformer models like UNETR also exhibited promising results, confirming the value of architectures that balance local spatial awareness and global context modeling. The full training process for the proposed model requires approximately 9 hours on a single GPU, but inference is extremely fast, with each prediction generated in only 0.2 seconds. Consequently, computational burden is not a limiting factor, with model inference operating in real-time. The primary bottleneck in clinical deployment remains MRI acquisition time, which dominates the overall latency within the MRgFUS workflow. While the Day 3 NPV serves as a practical imaging-based reference standard for treatment outcome, future validation using histology or clinical outcomes is needed to confirm the biomarker’s true accuracy is predicting the gold standard clinical outcome. For clinical translation, robust registration methods and efficient multi-contrast acquisitions, such as MPME[ 67 ] or MR Fingerprinting[ 65 ] will be essential. Overall, this study provides a framework for non-contrast, data-driven prediction of MRgFUS treatment efficacy. The integration of optimized architectures, efficient imaging protocols, and validated biomarkers could ultimately enable real-time, non-invasive assessment of treatment success within the MRgFUS workflow.

Experimental

As summarized in Figure 7 , the bootstrap analysis demonstrates that image registration accuracy is stable across different landmark subset sizes. Across all evaluations, the overall target registration error was 1.42 ± 0.08 mm. An example registration between the post-ablation Day 0 and Day 3 timepoints for one subject is shown in Figure 8 , illustrating qualitative continuity of anatomical structures, particularly within the non-perfused volume. Across all subjects, the mean inter-observer Dice was 0.85 ± 0.06 , indicating high volumetric overlap between raters. The corresponding MDA was 1.18 ± 0.49 mm , reflecting 3 voxel boundary consistency. This was also chosen as the length for random shift augmentation for the ground truth label. While these results demonstrate strong inter-rater consistency, supporting the robustness of the semi-automatic segmentation approach used to define the ground-truth NPV, inter-observer variability in the manual delineation of non-perfused volume (mean Dice 0.85 ± 0.06) introduces inherent uncertainty in the reference labels used for model training and evaluation. This variability is primarily concentrated at lesion boundaries, resulting in inconsistent voxel-level labels across expert segmentations and effectively setting an upper bound on achievable overlap-based performance metrics. Consequently, a portion of the model error reflects irreducible annotation noise rather than prediction failure, particularly near the perfusion–viability transition zone. Table III compares the effectiveness of different data augmentation strategies on the performance of a standard U-Net architecture for predicting day 3 post-ablation NPV. The performance is evaluated using both Dice and MDA. Results show that when no augmentation is applied during training, the model achieves a Dice score of 0.49 and an MDA of 5.5 mm. When traditional augmentation methods are applied, including Gaussian noise, random rotation, random crop, and label shift, the performance improves significantly (p-value < 0.001), with the Dice score increasing to 0.62, and the MDA decreasing to 3.7 mm. Adding PCA-based data augmentation to the traditional methods further improves the boundary precision: the Dice score remains similar at 0.64, but the MDA decreases from 3.7 mm to 3.0 mm. These results are shown for both the training and test data, compared to the Day 0 NPV in Figures 9 and 10 . Model performance was compared against clinically used baselines, including Day-0 NPV (average Dice = 0.60 across runs) and cumulative thermal dose thresholded at 240 culumulative minutes (average Dice = 0.28 across runs), both of which showed lower agreement with Day-3 NPV. A paired t-test found the inclusion of PCA-DA led to a statistically significant improvement in model performance for predicting NPV compared to the traditional augmentation methods (p-value = 0.006). Figure 9 shows that training and test accuracy both improve when augmentation is applied to the same underlying dataset. This simultaneous improvement indicates that augmentation reduces the train–test performance gap, demonstrating improved generalization rather than memorization, and therefore directly mitigates overfitting. The standard U-Net architecture, with a combination of augmentations (Gaussian noise, label shift, random crop, random rotation, and PCA-DA), achieves a Dice score of 0.62 and an MDA of 3.7 mm (see Table IV ). Adding PCA-DA slightly improves the Dice score to 0.64 and reduces the MDA to 3.0 mm. The U-Net Attention architecture performance was comparable. With traditional augmentations, the Dice score is 0.63, slightly higher than the baseline U-Net, but the MDA improves to 2.9 mm. When PCA-DA is introduced, the Dice score further increased to 0.64, matching the best performance observed with the standard U-Net, but the MDA remains similar at 3.0 mm. In contrast, the vision transformer architecture performs worse than both U-Net variants. With traditional augmentations, the Dice score drops to 0.45, and the MDA increases to 7.7 mm. When PCA-DA is applied, the Dice score further decreases to 0.38, and the MDA increases to 7.9 mm ( Table IV ). UNETR architecture performs similar to U-Net variants. With traditional augmentations, the Dice score is 0.60, and the MDA is 3.4 mm. When PCA-DA is applied, the Dice score improves to 0.62, and the MDA decreases to 3.3 mm ( Table IV ). Comparisons for all network architectures and use of traditional and PCA-DA augmentations techniques are shown in Figures 13 and 14 . Results from training a U-Net with the traditional augmentation methods using the MRI input combinations are outlined in Table II . Using only T1-weighted and T2-weighted images (Combination 0) leads to a Dice score of 0.50 and an MDA of 4.8 mm. When T1-weighted images are paired with MRTI data, the Dice score improves to 0.55, and the MDA decreases slightly to 4.6 mm. When T2-weighted images are combined with temperature data, the Dice score further improves to 0.58, and the MDA decreases to 3.9 mm. Surprisingly, when all three clinically utilized inputs, T1-weighted, T2-weighted, and, MRTI are used together, the Dice score remains at 0.57, with an MDA of 4.1 mm. These results are shown graphically and compared to the Day 0 NPV in Figures 15 and 16 . The best performance is achieved when additional quantitative maps are added and the model includes T1-weighted, T2-weighted, T1 maps, T2 maps, and MRTI data. This combination results in a Dice score of 0.62 and an MDA of 3.7 mm. Removing MRTI data while retaining T1-weighted, T2-weighted, T1 maps, and T2 maps slightly reduces performance, yielding a Dice score of 0.60 and an MDA of 3.8 mm. The best performing model architecture in this study was the U-Net Attention with both traditional augmentations and PCA-based data augmentation (PCA-DA), achieving a Dice score of 0.64 and the lowest mean deviation of agreement (MDA) of 3.0 mm. The spatial visualization of the prediction examples using this network in one of the subjects is shown in Figure 11 . A prediction example using all sequences with the U-Net architecture with attention blocks is shown in Figure 12 . When evaluating performance using all data and a reduced data volume that covered the truncated treatment region, the combination 4 data input achieved an overall Dice of 0.62 and a reduced data volume Dice of 0.66. Similarly, UNet-A reached a Dice score of 0.64 when using all data and a significant improvement of 0.72 (p < 0.0001) when evaluating the reduced data volume. A one-factor repeated-measures ANOVA revealed a significant effect of network architecture on Dice performance across cross-validation folds ( F 3 , 33 = 43.41 , p = 1.48 × 10 - 11 ), indicating that at least one architecture differed significantly from the others. Subsequent post-hoc paired Wilcoxon signed-rank tests were therefore performed to identify specific pairwise differences. The comparisons included U-Net vs. U-Net Attention, U-Net Attention vs. UNETR, and U-Net vs. Vision Transformer. This paired analysis directly tests whether performance differences are systematic across folds rather than driven by split-specific variability. Paired Wilcoxon signed-rank test results for all model comparisons, including median paired Dice differences and corresponding p -values, are summarized in Table VI .

Introduction

Magnetic resonance-guided focused ultrasound (MRgFUS) is emerging as a promising minimally invasive treatment for thermal therapies [ 1 ], [ 2 ]. This technique utilizes magnetic resonance imaging (MRI) to plan, monitor, and evaluate treatments, allowing for precise ultrasound-induced tumor destruction [ 3 ], [ 4 ]. Since MRgFUS is incisionless and the treated tumor is not excised after the MRgFUS treatment, imaging biomarkers are essential [ 5 ] for assessing tumor viability to ensure effective therapy. One of the key powers of MRI is that it provides different types of information about tissue properties and changes induced by interventions, including MRgFUS thermal ablation [ 6 ]. Conventional MR sequences, such as T1-weighted and T2-weighted imaging, are fundamental for anatomical visualization. Images acquired before the treatment help establish a baseline for tissue characteristics, while post-treatment images highlight changes caused by the therapy. Contrast-enhanced T1-weighted (CE-T1w) imaging is useful for assessing vascular changes and identifying necrotic tissue regions and T2-weighted (T2w) imaging is particularly effective in detecting edema and inflammation [ 7 ]–[ 12 ]. Quantitative MR images such as T1 and T2 maps offer quantifiable measures of tissue composition by quantitatively assessing relaxation times [ 13 ]–[ 15 ]. Unlike standard T1w and T2w images, T1 and T2 maps provide numerical values that can ideally be compared across different time points, scanners, and sites, to longitudinally assess physiological changes. These MR sequences are particularly useful for evaluating treatment-induced alterations in perfusion, cellular structure, and water content. Although the features from the reconstruced T1 and T2 maps are typically not used in MRgFUS assessment due to the long acquisition times, incorporation of T1 and T2 maps into biomarker analysis could potentially evaluate tissue response with greater precision, improving the ability to distinguish between viable and non-viable tissue immediately after treatment[ 13 ]. Magnetic resonance temperature imaging (MRTI) is critical for monitoring and quantifying the subject-specific response to MRgFUS ablation [ 16 ]. This real-time acquired data determines whether sufficient thermal energy was delivered to achieve therapeutic goals while minimizing damage to surrounding tissue. In clinical applications, MRgFUS treatments typically rely on MRTI and T2w and T1w MRI (with and without gadolinium contrast) to monitor and assess the treatment effects. Most MRgFUS ablative therapies are evaluated using imaging biomarkers that utilize MRTI data and CE-T1w to reflect temperature changes [ 17 ], [ 18 ] or vascular effects [ 10 ] induced by the treatment. Thermal dose is cumulatively assessed using MRTI obtained during the ablation process [ 19 ], [ 20 ], while vascular effects are evaluated at the conclusion of treatment using CE-T1w MRI. However, thermal dose has been shown to underestimate the extent of tissue perfusion loss and necrosis in large treatment areas [ 21 ], and CE-T1w images, when used to assess immediate post-treatment tissue perfusion (within 1 hour as compared to several days after ablation), are complicated by transient physiological responses such as tissue edema, hemorrhage, and cell membrane disruption [ 22 ]. Although other MR contrasts are sensitive to acute physiological changes resulting from MRgFUS treatment, the observed changes have been inconsistent in the literature [ 22 ]–[ 24 ]. The sensitivity of multiple MR parameters to MRgFUS thermal ablation suggests that a non-contrast multiparametric imaging biomarker for predicting treatment efficacy immediately after MRgFUS could be feasible. Deep learning methods are transforming various aspects of MRI [ 25 ]–[ 28 ], and this study explores the potential of using a deep learning neural network imaging biomarker to predict outcomes of MRgFUS ablation treatments. A recent study [ 29 ] has demonstrated that radiomics models based on non-contrast MRI can predict treatment outcomes for focused-ultrasound ablation. This model successfully predicted non-perfused volume ratio after FUS ablation of uterine leiomyomas, achieving strong discrimination on test sets. Similarly, a dual-sequence (T1 + T2) MRI-based radiomics model has been shown to predict ablation efficacy in adenomyosis patients with promising accuracy [ 30 ]. Work by Johnson et. al. [ 7 ] demonstrated that machine learning applied to non-contrast enhanced MRI can effectively predict tissue necrosis, offering a promising alternative to contrast-based assessment methods. Recent work has increasingly combined magnetic resonance guided focused ultrasound with machine learning and deep learning methods to automate quantitative assessment and improve treatment outcome prediction across multiple clinical indications. In uterine fibroids, deep learning and radiomics approaches have been applied to MRI for automated lesion segmentation, non-perfused volume quantification, and prediction of post-treatment outcomes and recurrence risk from pre-treatment imaging [ 31 ]. In functional neurosurgery, MR-guided focused ultrasound (MRgFUS) thalamotomy for essential tremor is supported by a substantial clinical evidence base, with emerging data-driven methods increasingly applied to characterize treatment effects, optimize targeting, and improve understanding of network-level changes using advanced imaging and connectivity analyses [ 32 ]. In musculoskeletal oncology, MRgFUS is an established modality for pain palliation in bone metastases [ 33 ]; however, despite its widespread clinical use and recognition as an imaging-guided precision therapy, dedicated deep learning–based approaches for treatment planning and monitoring in this indication remain comparatively limited. This work goes beyond these above efforts, including additional quantitative MRI metrics in the biomarker and identifying which components of the multiparametric imaging biomarker most accurately predict treatment efficacy. This information is crucial, as it could enable MRgFUS treatment providers to optimize data acquisition during MRgFUS ablative therapies, thereby enhancing the clinical workflow and overall efficiency of the procedure. Rigorous image registration techniques were employed to ensure accurate correspondence of all data at each of the assessment points. In addition, several data augmentation techniques were employed and evaluated to determine their effect in overall image biomarker prediction accuracy.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-09-13T09:25:22.628771+00:00