Rapid evaluation of the comprehensive quality of Paeonol/Cyclodextrin supramolecular complexes using CASSA based on near-infrared spectroscopy combined with artificial intelligence algorithm

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Preparing volatile component/cyclodextrin supramolecular complexes is a common method for enhancing the stability of volatile components. However, the quality assessment of supramolecular complexes is highly complex. This study first prepared paeonol/cyclodextrin supramolecular complexes and evaluated their overall quality. Then, Fourier Transform Near-Infrared Spectroscopy (FT-NIR) was combined with artificial intelligence (AI) to build a Support Vector Machine Classification (SVM) model and a Partial Least Squares Regression (PLSR) model. The SVM classification model reached 100% accuracy, whereas the PLSR quantitative model demonstrated R² > 0.90 for both calibration and prediction sets. Results confirm that integrating FT-NIR with AI improves the accuracy and reliability of qualitative/quantitative models. Via the Comprehensive Analysis of Single Spectral Acquisition (CASSA) method, dual detection of the complexes’ formation state and concentration was achieved, enabling rapid, comprehensive quality evaluation. This study demonstrates the excellent prospects of combining FT-NIR technology with the preparation of supramolecular complexes.
Full text 151,263 characters · extracted from preprint-html · click to expand
Rapid evaluation of the comprehensive quality of Paeonol/Cyclodextrin supramolecular complexes using CASSA based on near-infrared spectroscopy combined with artificial intelligence algorithm | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Rapid evaluation of the comprehensive quality of Paeonol/Cyclodextrin supramolecular complexes using CASSA based on near-infrared spectroscopy combined with artificial intelligence algorithm Bo-yao Zhang, Xin-li Li, Cheng Qian, Si-min Xue, Zhi-tong Zhang, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9167929/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Preparing volatile component/cyclodextrin supramolecular complexes is a common method for enhancing the stability of volatile components. However, the quality assessment of supramolecular complexes is highly complex. This study first prepared paeonol/cyclodextrin supramolecular complexes and evaluated their overall quality. Then, Fourier Transform Near-Infrared Spectroscopy (FT-NIR) was combined with artificial intelligence (AI) to build a Support Vector Machine Classification (SVM) model and a Partial Least Squares Regression (PLSR) model. The SVM classification model reached 100% accuracy, whereas the PLSR quantitative model demonstrated R² > 0.90 for both calibration and prediction sets. Results confirm that integrating FT-NIR with AI improves the accuracy and reliability of qualitative/quantitative models. Via the Comprehensive Analysis of Single Spectral Acquisition (CASSA) method, dual detection of the complexes’ formation state and concentration was achieved, enabling rapid, comprehensive quality evaluation. This study demonstrates the excellent prospects of combining FT-NIR technology with the preparation of supramolecular complexes. FT-NIR Artificial Intelligence Algorithms Cyclodextrin CASSA supramolecular complexes Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1. Introduction Volatile components are an important active ingredient in traditional Chinese medicine. However, because its primary active ingredients are volatile constituents, this property often results in poor stability and susceptibility to oxidation after extraction when used as a medicinal product, making it difficult to control its pharmaceutical quality. To address these issues, an effective approach involves using β-cyclodextrin to package volatile components into supramolecular complexes (also known as inclusion complexes). This method ensures the stability of volatile components while enabling their absorption by the human body; moreover, it is widely used owing to its simple process and low cost. Examples include 2-hydroxypropyl-β-cyclodextrin (HP-β-CD) encapsulating nutmeg (Xi et al., 2023 ). Yet for questions including the contents of volatile components in supramolecular complexes and whether inclusion has been fully achieved, solutions primarily depend on techniques such as scanning electron microscopy, X-ray diffraction, differential scanning calorimetry (Xu et al., 2017 ), and HPLC—all of which are often cumbersome, time-consuming, and require both sophisticated equipment and highly skilled operators. Additionally, it remained difficult to determine whether free volatile monomer components were present during HPLC-based content analysis of the complex, thus leading to issues such as inconsistent and unreliable content results. Therefore, for the inclusion of volatile components today, there is an urgent need for a rapid and reliable technique to detect the inclusion state and content of supramolecular complexes. Near-infrared spectroscopy (NIRS) is a technique that reflects absorption information from the harmonic and sum frequencies of molecular vibrations in chemical groups containing hydrogen elements (such as C-H, O-H, S-H, N-H, etc.). Consequently, this spectroscopic analysis covers nearly all organic compounds and mixtures (Kato et al., 2023 ; Liu et al., 2023 ). Furthermore, NIRS technology enables rapid scanning of samples with minimal pretreatment requirements, yielding stable, reliable, and accurate data. Currently, near-infrared technology is widely used in various fields such as rapid quantitative analysis of target components and identification of food and pharmaceutical products (Beć et al., 2022 ). For example, near-infrared technology can be used to rapidly detect adulteration in ginger powder and quantify its content (Yu et al., 2022 ). NIRS technology can also be used for rapid quantitative analysis of herbal medicine components. Artificial intelligence algorithms can establish models based on a series of data, gradually arriving at optimal solutions through repeated trial-and-error learning (Yang et al., 2024 ). Therefore, leveraging the vast data provided by near-infrared scanning technology and continuously optimizing models through artificial intelligence learning ultimately yields a relatively reliable and accurate mathematical model, which is a great solution. Peony root bark is the dried root bark of Paeonia suffruticosa Andr. , a plant belonging to the Ranunculaceae family, possessing rich edible and medicinal value. Its primary component, paeonol, is a volatile compound exhibiting antibacterial, anti-inflammatory, antioxidant, and anticancer effects (Xu et al., 2024 ; Bai et al., 2022 ; Yu et al., 2018 ). Owing to its instability, this study formed a Paeonol/Cyclodextrin supramolecular complex (P-CD complex ) by cyclodextrin inclusion of paeonol. The quality of the complex was evaluated using techniques such as X-ray diffraction and HPLC. Then, by integrating NIRS technology with artificial intelligence algorithms, a classification model was established capable of distinguishing three types: fully encapsulated paeonol, partially encapsulated paeonol, and free monomer groups (i.e., paeonol completely unencapsulated by cyclodextrin). Using preprocessing methods and feature wavelength extraction techniques, establish a PLSR quantitative model for the complexes in the fully encapsulated group. B = y combining NIRS technology with artificial intelligence algorithms, this approach enables rapid assessment of inclusion complexes formed between cyclodextrin and paeonol, as well as determination of the complex's content. Our institute has developed a method that is rapid and simple to operate, yielding accurate and reliable results with broad applicability. Also, it enables comprehensive quality evaluation through a single spectral acquisition, known as CASSA. This approach provides technical guidance and conceptual insights for comprehensive studies on the properties of supramolecular complexes. 2. Materials and Methods 2.1 Instruments and Reagents Antaris II FT-NIR Spectrometer (Thermo Fisher Scientific, USA); Waters e2695 HPLC with a 2998PDA detector (Waters, USA) were utilized. Chromatography-grade methanol was purchased from Tedia, USA; X-ray diffractometer (Bruker Corporation, Germany, Model: D8 Advance); Field-emission scanning electron microscope (Hitachi High-Tech Corporation, Japan, Model: SU8600); Fourier transform infrared spectrometer (Thermo Fisher Scientific, U.S.A., Model: Nicolet iS5); Differential scanning calorimeter (TA Instruments, U.S.A., Model: DSC 2500); Thermogravimetric analyzer (TA Instruments, U.S.A., Model: TGA 55);ultrapure water was prepared using a Millipore Milli-Q ultrapure water system; paeonol-rich bark extract monomer (or: Paeonia suffruticosa bark extract monomer, paeonol-containing) was purchased from Ruimao Biotechnology Co., Ltd.; paeonol-rich bark extract reference standard (or: Paeonia suffruticosa bark extract reference standard) was obtained from the China National Institute for Food and Drug Control (Batch No.: 110708–202309, purity ≥ 99.9%); β-cyclodextrin was purchased from Anhui Shanhe Pharmaceutical Excipients Co., Ltd.; all other reagents were of analytical grade and purchased from Sinopharm Chemical Reagent Co., Ltd. 2.2 Sample Preparation 2.2.1 Preparation of completely encapsulated paeonol samples Take an appropriate amount of paeonol extract, precisely weigh the mass, dissolve in an appropriate amount of anhydrous ethanol, add to a saturated β-cyclodextrin aqueous solution at 50°C, stir for 30 minutes, filter under vacuum after refrigerating for 24 hours, wash the filter cake with ethanol, dry at 60°C for 15 hours, and grind through a No. 4 sieve to obtain the paeonol complex. Add cyclodextrin passed through the No. 4 sieve and mix thoroughly. Prepare 160 portions of the complex (I). As shown in Fig. 1 (I). 2.2.2 Preparation of Partially Encapsulated paeonol Samples Take the complex prepared in Section 2.2.1 , add the ground paeonol monomer (passed through a No. 4 sieve), and mix thoroughly. Prepare 120 samples of partially encapsulated paeonol (II)—i.e., samples where the complex and paeonol coexist—with the results shown in Fig. 1 (II). 2.2.3 Preparation of Paeonol Monomer and Cyclodextrin Monomer Samples Take cyclodextrin passed through a No. 4 sieve and ground paeonol monomer passed through a No. 4 sieve separately, then mix them thoroughly. Prepare 40 portions of the mixture of paeonol monomer and cyclodextrin monomer (III). As shown in Fig. 1 (III). 2.2.4 Solution preparation Sample Solution Preparation Approximately 0.02 g of the paeonol complex was accurately weighed and placed in a stoppered Erlenmeyer flask. Precisely 100 mL of methanol was added, the flask was sealed tightly, and its weight was recorded. After sonication for 20 minutes, the weight loss was replenished with methanol. The mixture was filtered, and 1 mL of the subsequent filtrate was precisely transferred to a 10 mL volumetric flask. Methanol was added to dilute the solution to the mark, and the mixture was thoroughly mixed to obtain the test solution. Preparation of Reference Solution Take an appropriate amount of paeonol, precisely weigh the mass, place it in a volumetric flask, dilute to the mark with methanol, and mix thoroughly to prepare a solution containing 20 µg of Paeonol per 1 mL. 2.2.5 Chromatographic conditions A Hedera ODS-2 column (4.6 mm × 250 mm, 5 µm) was employed. The mobile phase was composed of methanol and water (60:40, v/v). Detection was carried out at a wavelength of 274 nm. The column temperature was maintained at 30°C. The injection volume was set to 10 µL. The flow rate was adjusted to 0.8 mL/min. 2.2.6 Methodological examination To verify the reliability of this method, validation tests including linearity, precision, repeatability, stability, and spiked recovery were performed in accordance with the guidelines of the International Conference on Harmonisation (ICH). A detailed description of the method is available in Section 1 of the Supplementary Materials. 2.3 Characterization of P-CD complex The successful formation of the Paeonol/Cyclodextrin supramolecular complex (P-CD complex) was confirmed via scanning electron microscopy (SEM), X-ray diffraction (XRD), Fourier Transform Infrared Spectroscopy(FT-IR) ,and thermogravimetric analysis (TGA). 2.3.1 Differential Scanning Calorimetry Approximately 5–10 mg of the sample was hermetically sealed in an aluminum pan and analyzed using a differential scanning calorimeter (DSC, TA Instruments, model: DSC 2500). The measurement was conducted from 30 to 295°C at a heating rate of 10°C/min under a constant flow of nitrogen gas (e.g., 50 mL/min) to prevent oxidation. An empty aluminum pan was used as the reference. 2.3.2 X-ray Diffraction An X-ray diffractometer (Bruker, Germany; model: D8 Advance) was operated at room temperature under the following conditions: Cu Kα radiation, voltage of 40 kV, current of 40 mA, scanning speed of 2°/min, and a scanning range of 2–50°. The obtained patterns were baseline-corrected and smoothed using MDI Jade 6.5, generating the final XRD patterns. 2.3.3 Thermogravimetric Analysis Thermogravimetric analysis (TGA) was performed on a thermogravimetric analyzer (TA Instruments, model: TGA 55). The measurements were carried out from 30 to 600°C at a heating rate of 20°C/min under a constant nitrogen purge (e.g., 50 mL/min) to prevent sample oxidation. 2.3.4 Fourier Transform Infrared Spectroscopy An accurate amount (2 mg) of the sample was mixed thoroughly with pure KBr at a mass ratio of 1:100 (sample: KBr), and the mixture was placed in a mold and pressed into a transparent pellet under a hydraulic press. The Fourier transform infrared (FT-IR) spectrum was recorded on an infrared spectrometer (Thermo Fisher Scientific, Inc., Model: Nicolet iS5) over the range of 400–4000 cm⁻¹ at a resolution of 4 cm⁻¹, with 32 scans. 2.3.5 Scanning Electron Microscope A small amount of the sample was evenly spread on a piece of conductive adhesive tape, and any loosely attached particles were removed using a rubber bulb. The sample was then gold sputter-coated for approximately 60 seconds. SEM images were acquired using a field emission scanning electron microscope (Hitachi High-Technologies Corporation, model: SU8600) in the secondary electron (SE) mode at an acceleration voltage of 5 kV. 2.4 Near-infrared spectroscopy acquisition The samples were scanned and collected using an Antaris II FT-NIR spectrometer (Thermo Fisher Scientific, USA) in diffuse reflection mode. Samples were loaded into Fisher Shell Type 1 Glass quartz sample cells (19 × 51 mm, 2 dram). Spectra were acquired over the wavelength range of 12,000–4,000 cm⁻¹ with a resolution of 16 cm⁻¹. Each spectrum was scanned 32 times. For each sample, one background scan was performed followed by three parallel scans. The averaged spectrum was used as the experimental data for subsequent analysis. 2.5 Establishment of Classification Models The average spectral data acquired via near-infrared spectroscopy (NIRS) exhibits high accuracy without the need for preprocessing or characteristic wavelength selection. Therefore, this study adopted a direct modeling approach that omits data processing, splitting the raw spectral data into a training set and a prediction set at a ratio of 7:3. The data were categorized into three groups: the fully encapsulated group (I), the partially encapsulated group (II), and the free monomer group (III). The optimal classification model was selected using training accuracy and prediction accuracy as evaluation metrics. 2.5.1 Classification Model Selection (1) Bayesian Optimization is a highly effective global optimization algorithm aimed at finding the global optimum solution. It efficiently addresses the classic machine-intelligent problem in sequential decision theory: determining the next evaluation location based on information gathered about the unknown objective function f, thereby achieving the optimum solution most rapidly (Lorenz et al., 2017 ). The Bayesian optimization framework can obtain the optimal solution for complex objective functions with a minimal number of evaluations. Essentially, the framework first employs a surrogate model to approximate the true objective function. It then proactively selects the most "promising" evaluation points for sampling based on this approximation, thereby avoiding unnecessary sampling (Dastin-van Rijn & Widge, 2025 ). Consequently, Bayesian Optimization is also termed active optimization. Additionally, the framework effectively leverages complete historical information to enhance search efficiency. (2) Long Short-Term Memory (LSTM) and Recurrent Neural Networks (RNN) are primarily used for processing sequential data. Their defining feature is that the output of a neuron at one time step can be fed back as input to the same neuron in the next time step. This recurrent architecture is highly suited for time series data, as it preserves temporal dependencies within the data. When unfolding an RNN, repetitive structures emerge, and parameters in the architecture are shared—this significantly reduces the number of neural network parameters requiring training. On the other hand, shared parameters also enable the model to scale to data of varying lengths, allowing RNN inputs to be sequences of arbitrary length. For example, when training a fixed-length sentence: a feedforward neural network would assign a separate parameter to each input feature, whereas a recurrent neural network can share the same weight parameters across time steps. Although RNNs were originally designed to learn long-term dependencies, extensive practice has shown that standard RNNs often struggle to retain information over the long term. To address this long-term dependency problem, Hochreiter et al. proposed the LSTM network (Graves & Schmidhuber, 2005 ) to enhance traditional recurrent neural network models. Today, LSTM has become one of the most effective sequence models in practical applications. Compared to the hidden units in RNNs, those in LSTM have a more complex internal structure: as information flows through the network, the addition of linear interventions enables LSTM to selectively amplify or diminish information intensity (Malashin et al., 2024 ). (3) Convolutional Neural Network (CNN). The basic structure of CNN consists of an input layer, convolutional layers, pooling layers (also called sampling layers), fully connected layers, and an output layer. Typically, multiple convolutional and pooling layers are employed, arranged alternately—i.e., one convolutional layer followed by one pooling layer, and so forth. In the convolutional layer, each neuron in the output feature map forms local connections with its local inputs from the previous layer. The neuron's input value is calculated by weighting and summing these local inputs using corresponding connection weights, then adding a bias term. This process is fundamentally equivalent to the convolution operation, from which CNN gets its name (Li et al., 2022 ; Zang et al., 2021 ; Sun et al., 2020 ). (4) Extreme Learning Machine (ELM) randomly selects the input weights and hidden layer biases of the network, and derives output weights through analytical computation. This effectively overcomes the limitations of traditional SLFN learning algorithms and has been widely applied in various fields such as disease diagnosis, traffic sign recognition, and image quality assessment. Initially limited to single-hidden-layer feedforward neural networks, it was later extended to RBF neural networks, recurrent neural networks, generalized single-hidden-layer feedforward neural networks, and multi-hidden-layer feedforward neural networks. From a design perspective, ELM aims to unify research problems in machine learning—including regression, classification, clustering, compression, and feature extraction—within a single framework (Huang et al., 2015 ; Chen et al., 2014 ). In terms of learning efficiency, ELM is simple to implement, has an extremely fast learning speed, and requires minimal human intervention. (5) Support Vector Machine (SVM) (Hsieh & Yeh, 2011 ) is a supervised learning algorithm based on statistical learning theory, widely applied in classification and regression tasks. Its core principle involves finding an optimal decision hyperplane that maximizes the margin (i.e., the distance between classes) between samples of different classes, thereby enhancing the model's generalization capability. The key to SVM lies in minimizing structural risk—not only minimizing low error rates on training data but also reducing model complexity by maximizing the classification margin to prevent overfitting. For linearly separable data, SVM employs hard margin optimization to directly determine a hyperplane that perfectly classifies all samples. In contrast, for linearly inseparable data, it introduces slack variables and a penalty parameter (soft margin, to tolerate minor misclassifications), allowing some samples to be misclassified to enhance model robustness. The SVM optimization problem ultimately transforms into a convex quadratic programming problem, solvable via the Lagrange multiplier method. Its solution is determined by only a few support vectors, and thus exhibits sparsity and computational efficiency. Furthermore, SVM excels in scenarios with small sample sizes and high-dimensional data (e.g., text classification, image recognition), making it one of the classic algorithms in machine learning. 2.6 Establishment of Quantitative Models Given the extremely high requirements for reliability and accuracy in quantitative models, such models require preprocessing methods—including spectral data preprocessing and characteristic wavelength extraction—to perform dimensionality reduction and optimization on sample data, thereby preventing overfitting. Using the content data of the complex in the fully encapsulated group as the independent variable, a quantitative model was established for the paeonol complex content in this group. 2.6.1 Preprocessing Method This study used The Unscrambler X 10.4 software (CAMO, Inc., Texas, USA) to preprocess raw spectra. Nine preprocessing methods were applied to the spectra, including first derivative (1st derivative, 1D), second derivative (2nd derivative, 2D), Savitzky-Golay (SG), standard normal variate (SNV), multiplicative scatter correction (MSC), median filter (MF), SNV+1D, MSC+1D, and SG + 1D. These nine methods were evaluated to select the one that yielded the highest model fitting accuracy. 2.6.2 Characteristic Wavelength Selection (1) Competitive Adaptive Reweighted Sampling (CARS) is a feature selection method based on full-spectrum partial least squares (PLS) regression. By mimicking Darwinian evolution algorithms, it selects the optimal combination of effective variables in the spectrum: specifically, it retains wavelength points with larger absolute regression coefficients in the PLS model while removing those with smaller weights in the same model. Through iterative and competitive operations, it selects N wavelength subsets in each iteration and identifies the subset with the lowest root mean square error of cross-validation (RMSECV) value, which corresponds to the optimal number of variables (Li et al., 2009 ). (2) The Iterative Retaining Informative Variables (IRIV) method (Yun et al., 2014 ) is a feature selection algorithm based on the Binary Matrix Shuffle Filter (BMSF). It comprehensively accounts for the importance of each variable, and after optimization, further screens feature wavelength variables to determine the optimal number of features. As an effective method for selecting optimal feature wavelengths in near-infrared spectroscopy (NIRS), IRIV adopts a strategy of randomly combining variables to account for potential interactions between them, thereby classifying all variables into four categories: strong informative variables, weak informative variables, non-informative variables, and interfering variables. Through multiple iterations—each aimed at retaining strong and weak informative variables while eliminating non-informative and interfering variables—the optimal variable set is ultimately obtained via backward elimination. (3) Interval Combination Optimization (ICO) is one of the commonly used interval variable selection methods, as it can provide more reasonable wavelength intervals and enhance the predictive capability of models. Proposed within the framework of Model Portfolio Analysis (MPA) combined with Weighted Bootstrap Sampling (WBS), ICO leverages the strengths of both approaches. As a general framework for variable selection and evaluation, MPA offers comprehensive insights from a large number of submodels; its core concept is to establish data analysis methods by statistically analyzing the distribution of these submodels, making interval selection methods more advantageous than single-wavelength selection approaches. Weighted Bootstrap Sampling (WBS) is a random sampling technique in which different objects are assigned weights based on their occurrence frequencies. As an interval selection method, ICO searches through the interval weights across the entire variable space and selects useful subintervals (Song et al., 2016 ). (4) The Successive Projections Algorithm (SPA) is a forward selection method—an iterative forward technique that minimizes multicollinearity—traditionally used for variable selection in multivariate calibration. It starts with a single wavelength and adds a new wavelength in each iteration until a specified number (N) of wavelengths is reached. The optimal number of variables is determined by calculating the root mean square error of cross-validation (RMSECV) for the N-wavelength subset using multiple linear regression (MLR) (Canova et al., 2023 ; Xiaobo et al., 2010 ). 2.6.3 Establishment of the PLSR Quantitative Model Partial Least Squares Regression (PLSR) analysis was performed using MATLAB 2024a software (MathWorks Inc., Natick, MA, USA) to enable quantitative prediction of the compound. The Kennard and Stone (KS) algorithm was employed to divide all samples into a training set and a test set at a 7:3 ratio, establishing the PLSR model for performance prediction. $$\:{R}^{2}=1-\frac{\sum\:{\left(Yi-\widehat{Yi}\right)}^{2}}{\sum\:\left(Yi-Yn\right)}\:$$ $$\:RMSE=\sqrt{\frac{\sum\:{(\widehat{Yi}-Yi)}^{2}}{n}}$$ $$\:RPD=\frac{SD}{RMSEP}$$ In the above formulas, Yi represents the ith reference value, ̂ \(\:\widehat{Yi}\) represents the ith predicted value, and Yn represents the mean of n referencevalue. 3. Results and discussion 3.1 Characterization of P-CD complex 3.1.1Differential Scanning Calorimetry In Fig. 2 (A), an endothermic peak for β-cyclodextrin at 92°C, as well as two endothermic peaks for paeonol at 50°C and 224°C, are observed in the free monomer group, indicating no changes after physical mixing. When guest molecules are incorporated into the β-CD cavity, their melting and sublimation points may shift to different temperatures or disappear (Nerome et al., 2013 ). In the DSC curve of the inclusion complex, the endothermic peak of paeonol at 50°C vanishes, confirming P-CD complex formation. 3.1.2 X-ray Diffraction As shown in Fig. 2 (B), paeonol exhibits strong diffraction peaks in the range of 10° < 2θ < 30°, with notably high intensity at 11.72°, 16.41°, 20.95°, 23.58°, and 25.54°. These peaks are characteristic of paeonol, indicating that it is a crystalline compound. The XRD pattern of the free monomer group (physical mixture of paeonol and β-cyclodextrin) represents the superposition of the patterns of paeonol monomer and β-cyclodextrin, confirming that it is merely a simple physical mixture without chemical reactions or structural modifications. The absence of paeonol’s characteristic peaks in the inclusion complex confirms that paeonol is encapsulated within the β-cyclodextrin cavity to form an inclusion complex, and consequently loses its original crystalline structure (Cao et al., 2023 ; Betlejewska-Kielak et al., 2021 ; Guo et al., 2011 ). 3.1.3 Thermogravimetric Analysis The TGA curves for each substance are shown in Fig. 2 (C). Paeonol exhibits poor thermal stability, with a total mass loss of over 85% at relatively low temperatures (50–177°C). The mass loss of the P-CD complex occurs in three stages: the first stage accounts for approximately 12% mass loss, attributable to the evaporation of surface and internal water; the second stage occurs between 300–328°C, with a mass loss of approximately 77%, representing the main weight loss phase, which is associated with the thermal decomposition of the P-CD complex during this stage. The thermal stability of paeonol was significantly enhanced due to interactions between the guest molecules and the inner cavity of β-cyclodextrin, indicating the formation of the inclusion complex (Li et al., 2016 ). 3.1.4 Fourier Transform Infrared Spectroscopy As shown in Fig. 2 (D), the infrared spectrum of β-cyclodextrin exhibits four characteristic peaks: a broad hydroxyl (-OH) stretching vibration peak near 3,440 cm⁻¹; methylene (-CH₂) and methyl (-CH₃) stretching vibration peaks near 2,920 cm⁻¹; a peak near 1,642 cm⁻¹, which is attributed to the bending vibration of adsorbed water or hydroxyl groups (not a carbonyl stretch, as β-cyclodextrin contains no carbonyl groups); and an ether bond (-C-O-C) stretching vibration peak near 1,020 cm⁻¹. In the infrared spectrum of paeonol, the broad peak at 2,976 cm⁻¹ corresponds to the stretching vibration of the phenolic hydroxyl (-OH) group; the high-intensity absorption peak at 1,620 cm⁻¹ represents the stretching vibration of the aromatic ketone carbonyl group, which undergoes a significant bathochromic shift to lower wavenumbers due to conjugation between the benzene ring and the carbonyl group; the strong absorption peak at 1,207 cm⁻¹ corresponds to the stretching vibration of the carbon-oxygen (C-O) bond in the aromatic ether. The aromatic ring skeletal stretching vibrations observed in the 1,650–1,430 cm⁻¹ region further confirm the presence of the benzene ring. These analytical results are consistent with the theoretical structure of paeonol. The infrared spectrum of the physical mixture exhibits characteristic peaks similar to those of β-cyclodextrin and paeonol, representing a simple superposition of both compounds. This indicates no significant intermolecular interactions during physical mixing (Cao et al., 2023 ). The disappearance of the characteristic peaks of paeonol in the infrared spectrum of the paeonol/β-cyclodextrin supramolecular complexes indicates the formation of the inclusion complex (Mansi et al., 2025). 3.1.5 Scanning Electron Microscope As shown in Fig. 3 (A-D), β-cyclodextrin (A) appears as irregularly shaped lumps, while paeonol (B) consists of fine crystalline particles. The physical mixture (free monomer group) exhibits the characteristic features of both β-cyclodextrin and paeonol (C). In contrast, the P-CD complex (D) appears as polygonal aggregates with smooth surfaces and a layered structure. Its morphology differs from that of both β-cyclodextrin and paeonol, indicating the formation of the P-CD complex (Zhu et al., 2020 ). 3.2 Classification of Raw Near-Infrared Spectral Analysis As shown in Fig. 4 (A), no significant differences were observed among the 320 sets of raw near-infrared spectral data—including 160 fully encapsulated samples, 120 incompletely encapsulated samples, and 40 mixed monomer samples—making it difficult to distinguish between fully encapsulated and non-fully encapsulated samples. A principal component analysis (PCA) model was established using all raw spectral data, and the results are presented in Fig. 4 (B). Although the three sample categories could be preliminarily classified—with Principal Component 1 (PC1) contributing 88.8% and the total cumulative contribution reaching 99.1%—complete differentiation of the three categories remained unattainable. Furthermore, raw near-infrared spectra often contain extensive sample information, including irrelevant data that may interfere with quantitative spectral analysis. Therefore, we preprocessed the raw spectral data and combined it with artificial intelligence algorithms for characteristic wavelength selection, ultimately establishing a partial least squares regression (PLSR) model to detect the content of the P-CD complex. 3.3 Classification Model Development The experiment employed LSTM-Bayesian optimization, CNN-Bayesian optimization, ELM, and an optimizable SVM model to establish classification models using raw spectral data. The results are shown in Table 1 : Table 1 Classification Model Accuracy Model Training set accuracy Prediction Set Accuracy Class Ⅰ Class Ⅱ Class Ⅲ Class Ⅰ Class Ⅱ Class Ⅲ LSTM 100% 98.9% 92.9% 100% 100% 87.5% CNN 98.1% 98.9% 92.9% 100% 100% 93.3% ELM 100% 97.9% 100% 96.4% 100% 100% SVM 100% 100% 100% 100% 100% 100% Note: Bold text indicates the best model for each category. Among these, the optimal SVM model obtained after 30 iterations achieved a confusion matrix for the three-class classification shown in Fig. 5 . Each class attained 100% accuracy, with identical accuracy rates for both the training and prediction sets, indicating no overfitting. Among the four selected classification models, the optimizable SVM model demonstrated the highest accuracy and good reliability. The other models failed to completely and correctly classify the three categories. Therefore, this study proposes to adopt the SVM model as the classification model. 3.4 Preprocessing of raw spectral data 3.4.1 Preprocessing of Fully Enclosed Group Spectral Data By nine preprocessing methods, PLSR models were established. Among them, the four methods—1D, 2D, MSC, and SNV—exhibited higher R² values as shown in Table-2, demonstrating more pronounced effects. The preprocessed spectra are depicted in Fig. 6 . Table 2 Complete Enclosure Group Preprocessing Preprocessing A R 2 T R 2 P RPD 1D 13 1.00 0.94 4.08 2D 12 0.99 0.95 4.54 MF 13 1.00 0.90 3.67 MSC 9 0.97 0.92 3.71 MSC+1D 39 1.00 0.42 1.45 SG 14 1.00 0.93 3.94 SG+1D 25 1.00 0.38 1.46 SNV 23 1.00 0.91 3.52 SNV+1D 39 1.00 -0.13 1.10 Note: Bold text indicates the selected preprocessing method. 3.5 Characteristic Wavelength Selection After preprocessing near-infrared hyperspectral data, feature wavelengths are selected using multi-intelligence algorithms to extract relevant spectral variables. Models are constructed based on CARS-I RIV, CARS-SPA, and ICO feature wavelength algorithms, and their predictive capabilities are evaluated. 3.5.1 Fully Enclosed Group Characteristic Wavelength For the fully encapsulated sample group, the results of three feature wavelength extraction methods—CARS-SPA, CARS-IRIV, and ICO—are presented in Table 4. Among these methods, the combination of ICO (feature wavelength extraction) and SNV (spectral preprocessing) yielded the following model parameters: R² T =0.9757, R² P =0.9476, RMSEP/RMSEC = 1.04, and RPD = 4.43. These parameter results indicate that the model is accurate and reliable, with excellent predictive performance and a reasonable fitting level. Preliminary dimensionality reduction of the raw spectral data via preprocessing simplified data complexity. Subsequently, during the iterations of the ICO algorithm, the weight coefficients of each wavelength band changed as the number of iterations increased. A yellower color indicates a weight coefficient closer to 1, while a bluer color indicates a weight coefficient closer to 0; if the color is between blue and yellow, the weight coefficient falls between 0 and 1. The width of the finally selected wavelength intervals (Fig. 7 I-A) was automatically optimized using a local search strategy integrated into the ICO algorithm. A total of 465 feature wavelengths were selected (Fig. 7 I-C), which significantly reduced information complexity compared to the original 1557 wavelengths and further reduced data dimensionality. This indicates that the ICO algorithm not only rapidly extracts spectral data from complex compound samples but also refines spectral information based on this foundation. In summary, selecting the ICO + SNV algorithm as the PLSR quantitative model for establishing a complete complex is the most effective and reliable approach. Table 3 Evaluation of PLSR Parameters for Complete Inclusion Complexes Method Preprocessing A Variables Calibration Prediction RPD RMSEP/RMSEC RSMEC R 2 T RMSEP R 2 P ICO 1D 7 310 4.18 0.9648 4.84 0.9021 3.20 1.16 2D 37 366 0.00 1.0000 4.40 0.9167 3.53 82309.25 MSC 26 456 0.01 1.0000 5.48 0.8734 2.83 743.23 SNV 11 465 3.46 0.9757 3.59 0.9476 4.43 1.04 IRIV 1D 3 4 7.80 0.8801 9.12 0.6123 1.61 1.17 2D 6 7 7.04 0.9023 8.71 0.6466 1.68 1.24 MSC 22 38 1.86 0.9931 2.74 0.9658 5.46 1.48 SNV 5 7 3.66 0.9733 5.85 0.8415 2.52 1.60 SPA 1D 4 9 7.03 0.9025 8.45 0.6488 1.71 1.20 2D 6 9 7.00 0.9034 9.40 0.5547 1.52 1.34 MSC 20 29 2.64 0.9862 3.37 0.9471 4.35 1.28 SNV 6 9 4.09 0.9670 4.76 0.8943 3.12 1.16 Note: Bold text indicates the optimal PLSR model results. 3.6 Visualization of PLSR Quantitative Models 3.6.1 Complete Enclosure Group PLSR Analysis As shown in Fig. 8 , the scatter plots of the PLSR models for the P-CD complex (Fig. 8 A&B) reveal that all samples cluster closely around the regression lines—indicating the successful establishment of the regression models and strong predictive capability. Additionally, the high degree of overlap between the red and blue lines demonstrates the close agreement between the predicted values and reference values of the samples (Fig. 8 D&E). Collectively, these results confirm that the model established for the fully encapsulated group samples (following preprocessing and characteristic wavelength selection) is reliable and possesses robust predictive capability. Therefore, this study ultimately adopted the SNV + ICO-PLSR method to establish a quantitative model for the complex content in the fully encapsulated complex group. 4. Conclusion In this study, the combination of spectral data obtained via FT-NIR rapid scanning and deep learning multi-intelligence algorithms not only addressed the difficulty in evaluating the inclusion state of Paeonol after cyclodextrin encapsulation but also resolved inaccuracies in complex content detection. This approach enabled comprehensive analysis of single spectral acquisition (CASSA) while improving both reliability and accuracy. First of all, categorical analysis of the raw NIR spectral data successfully achieved the classification of three inclusion states: complete inclusion, incomplete inclusion, and free monomer. Second, the raw spectral data were subjected to dimensionality reduction and optimization using preprocessing and characteristic wavelength selection methods, which enhanced the predictive capability and accuracy of the regression model. A PLSR quantitative model for P-CD complex content was established for the fully encapsulated group. Based on the model parameters, this established model exhibits good linearity and predictive capability, demonstrating high reliability. In conclusion, the FT-NIR technology combined with artificial intelligence algorithms employed in this study establishes predictive models for cyclodextrin-included volatile components. Compared to traditional methods for complex characterization and content detection, this approach enables comprehensive, rapid, and accurate multi-parameter evaluation of complex quality in a single measurement. It facilitates rapid tracking of complexes, ensures their pharmaceutical stability, and provides a more efficient solution for establishing quality standards for these compounds. CRediT authorship contribution statement Bo-yao Zhang : Conceptualization, Investigation, Methodology, Visualization, Writing – original draft. Xin-li Li : Data curation, Formal analysis, Software, Project administration. Cheng Qian : Methodology, Validation, Visualization. Si-min Xue : Conceptualization, Investigation, Data curation. Zhi-tong Zhang : Project administration, Methodology. Si-yi Liu : Visualization. Guo Zhang : Data curation. Guo-ping Peng : Supervision, Writing – review and editing. Ning Ding : Supervision, Formal analysis. Yun-feng Zheng : Writing – review and editing, Methodology, Funding acquisition. Declarations Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Funding The author(s) declare that financial support was received for the research and/or publication of this article. This research study was supported by the National Natural Science Foundation of China (No. 81473390),and Jiangsu “333” project. Author Contribution Bo-yao Zhang: Conceptualization, Investigation, Methodology, Visualization, Writing – original draft. Xin-li Li: Data curation, Formal analysis, Software, Project administration. Cheng Qian: Methodology, Validation, Visualization. Si-min Xue: Conceptualization, Investigation, Data curation. Zhi-tong Zhang: Project administration, Methodology. Si-yi Liu: Visualization. Guo Zhang: Data curation. Guo-ping Peng: Supervision, Writing – review and editing. Ning Ding: Supervision, Formal analysis. Yun-feng Zheng: Writing – review and editing, Methodology, Funding acquisition. Data availability The data that has been used is confidential. References Bai D, Li X, Wang S, Zhang T, Wei Y, Wang Q, Dong W, Song J, Gao P, Li Y, Wang S, Dai L (2022) Advances in extraction methods, chemical constituents, pharmacological activities, molecular targets and toxicology of volatile oil from Acorus calamus var. angustatus Besser. Front Pharmacol 13:1004529 Beć KB, Grabska J, Huck CW (2022) Miniaturized NIR Spectroscopy in Food Analysis and Quality Control: Promises, Challenges, and Perspectives. Foods, 11 (10) Betlejewska-Kielak K, Bednarek E, Budzianowski A, Michalska K, Maurin JK (2021) Comprehensive Characterisation of the Ketoprofen-β-Cyclodextrin Inclusion Complex Using X-ray Techniques and NMR Spectroscopy. Molecules, 26 (13) Canova LDS, Vallese FD, Pistonesi MF, de Araújo Gomes A (2023) An improved successive projections algorithm version to variable selection in multiple linear regression. Anal Chim Acta 1274:341560 Cao C, Deng C, Xuan F, Zhou Y (2023) Structural characterization and molecular dynamics simulation of inclusion complex of large ring cyclodextrin and alpha-tocopherol. China Oils Fats 48(5):49–55 Article 1003–7969(2023)48:5 Chen H, Peng J, Zhou Y, Li L, Pan Z (2014) Extreme learning machine for ranking: generalization analysis and applications. Neural Netw 53:119–126 Dastin-van Rijn EM, Widge AS (2025) Failure modes and mitigations for Bayesian optimization of neuromodulation parameters. J Neural Eng, 22 (3) Graves A, Schmidhuber J (2005) Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Netw 18(5–6):602–610 Guo P, Su Y, Cheng Q, Pan Q, Li H (2011) Crystal structure determination of the β-cyclodextrin-p-aminobenzoic acid inclusion complex from powder X-ray diffraction data. Carbohydr Res 346(7):986–990 Hsieh TJ, Yeh WC (2011) Knowledge discovery employing grid scheme least squares support vector machines based on orthogonal design bee colony algorithm. IEEE Trans Syst Man Cybern B Cybern 41(5):1198–1212 Huang G, Huang GB, Song S, You K (2015) Trends in extreme learning machines: a review. Neural Netw 61:32–48 Kato M, Kim J, Oh J, Shimizu D, Fukui N, Shinokubo H (2023) Near-Infrared-Responsive Hydrocarbons Designed by π-Extension of Indeno[1,2,3,4-pgra] perylene at the 1,2,12-Positions. Chemistry, 29(23), e202300249 Li H, Liang Y, Xu Q, Cao D (2009) Key wavelengths screening using competitive adaptive reweighted sampling method for multivariate calibration. Anal Chim Acta 648(1):77–84 Liu CH, He Z, Ruchlin C, Che Y, Somers K, Perepichka DF (2023) Thiele's Fluorocarbons: Stable Diradicaloids with Efficient Visible-to-Near-Infrared Fluorescence from a Zwitterionic Excited State. J Am Chem Soc 145(29):15702–15707 Li Y, Li F, Cai HY, Chen X, Sun W, Shen WY (2016) Structural characterization of inclusion complex of arbutin and hydroxypropyl-β-cyclodextrin. Trop J Pharm Res 15(10):2227–2233 Li Z, Liu F, Yang W, Peng S, Zhou J (2022) A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects. IEEE Trans Neural Netw Learn Syst 33(12):6999–7019 Lomarat P, Phechkrajang C, Sunghad P, Anantachoke N (2024) Raman spectroscopy coupled with the PLSR model: A rapid method for analyzing gamma-oryzanol content in rice bran oil. Food Chem X 24:101923 Lorenz R, Hampshire A, Leech R (2017) Neuroadaptive Bayesian Optimization and Hypothesis Testing. Trends Cogn Sci 21(3):155–167 Malashin I, Tynchenko V, Gantimurov A, Nelyub V, Borodulin A (2024) Applications of Long Short-Term Memory (LSTM) Networks in Polymeric Sciences: A Review. Polym (Basel), 16 (18) Mansi, Khanna P, Yadav S, Singh A, Khanna L (2025) Inclusion complexes of novel formyl chromone Schiff bases with β-Cyclodextrin: Synthesis, characterization, DNA binding studies and in-vitro release study. Carbohydr Polym, 347 , 17, Article 122667. Nerome H, Machmudah S, Wahyudiono, Fukuzato R, Higashiura T, Youn YS, Lee YW, Goto M (2013) Nanoparticle formation of lycopene/β-cyclodextrin inclusion complex using supercritical antisolvent precipitation. J Supercrit Fluids 83:97–103 Sankom A, Mahakarnchanakul W, Rittiron R, Sajjaanantakul T, Thongket T (2021) Detection of Profenofos in Chinese Kale, Cabbage, and Chili Spur Pepper Using Fourier Transform Near-Infrared and Fourier Transform Mid-Infrared Spectroscopies. ACS Omega 6(40):26404–26415 Song X, Huang Y, Yan H, Xiong Y, Min S (2016) A novel algorithm for spectral interval combination optimization. Anal Chim Acta 948:19–29 Sun Y, Xue B, Zhang M, Yen GG, Lv J (2020) Automatically Designing CNN Architectures Using the Genetic Algorithm for Image Classification. IEEE Trans Cybern 50(9):3840–3854 Wu K, Ma F, Wei C, Gan F, Du C (2023) Rapid Determination of Nitrate Nitrogen Isotope in Water Using Fourier Transform Infrared Attenuated Total Reflectance Spectroscopy (FTIR-ATR) Coupled with Deconvolution Algorithm. Molecules, 28 (2) Xiaobo Z, Jiewen Z, Povey MJ, Holmes M, Hanpin M (2010) Variables selection methods in near-infrared spectroscopy. Anal Chim Acta 667(1–2):14–32 Xi X, Huang J, Zhang S, Lu Q, Fang Z, Li C, Zhang Q, Liu Y, Chen H, Liu A, Liu S, Wang C, Li S, Hu B (2023) Preparation and characterization of inclusion complex of Myristica fragrans Houtt. (nutmeg) essential oil with 2-hydroxypropyl-β-cyclodextrin. Food Chem 423:136316 Xu F, Yang Q, Wu L, Qi R, Wu Y, Li Y, Tang L, Guo DA, Liu B (2017) Investigation of Inclusion Complex of Patchouli Alcohol with β-Cyclodextrin. PLoS ONE, 12(1), e0169578 Xu Y, Bian S, Shang L, Wang X, Bai X, Zhang W (2024) Phytochemistry, pharmacological effects and mechanism of action of volatile oil from Panax ginseng C.A.Mey: a review. Front Pharmacol 15:1436624 Xu Y, Liu J, Sun Y, Chen S, Miao X (2023) Fast detection of volatile fatty acids in biogas slurry using NIR spectroscopy combined with feature wavelength selection. Sci Total Environ 857(Pt 1):159282 Yang J, Ma X, Guan H, Yang C, Zhang Y, Li G, Li Z, Lu Y (2024) A quality detection method of corn based on spectral technology and deep learning model. Spectrochim Acta Mol Biomol Spectrosc 305:123472 Yu DX, Guo S, Zhang X, Yan H, Zhang ZY, Chen X, Chen JY, Jin SJ, Yang J, Duan JA (2022) Rapid detection of adulteration in powder of ginger (Zingiber officinale Roscoe) by FT-NIR spectroscopy combined with chemometrics. Food Chem X 15:100450 Yu X, Sun S, Guo Y, Liu Y, Yang D, Li G, Lü S (2018) Citri Reticulatae Pericarpium (Chenpi): Botany, ethnopharmacology, phytochemistry, and pharmacology of a frequently used traditional Chinese medicine. J Ethnopharmacol 220:265–282 Yun YH, Wang WT, Tan ML, Liang YZ, Li HD, Cao DS, Lu HM, Xu QS (2014) A strategy that iteratively retains informative variables for selecting optimal variable subset in multivariate calibration. Anal Chim Acta 807:36–43 Zang B, Ding L, Feng Z, Zhu M, Lei T, Xing M, Zhou X (2021) CNN-LRP: Understanding Convolutional Neural Networks Performance for Target Recognition in SAR Images. Sensors (Basel) , 21 (13) Zhu W, Wu J, Guo X, Sun X, Li Q, Wang J, Chen L (2020) Development and physicochemical characterization of chitosan hydrochloride/sulfobutyl ether-β-cyclodextrin nanoparticles for cinnamaldehyde entrapment. J Food Biochem, 44(6), e13197 Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterials.docx Supplementary material Supplementary data 1. Information on data related to the highperformance liquid chromatography Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 09 Apr, 2026 Reviews received at journal 03 Apr, 2026 Reviews received at journal 01 Apr, 2026 Reviewers agreed at journal 31 Mar, 2026 Reviewers agreed at journal 31 Mar, 2026 Reviewers agreed at journal 30 Mar, 2026 Reviewers invited by journal 30 Mar, 2026 Editor assigned by journal 21 Mar, 2026 Submission checks completed at journal 21 Mar, 2026 First submitted to journal 19 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9167929","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":615464615,"identity":"1c14da5c-ca1f-4880-9bef-801494bc029d","order_by":0,"name":"Bo-yao Zhang","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Bo-yao","middleName":"","lastName":"Zhang","suffix":""},{"id":615464616,"identity":"630f1de7-918e-44aa-b7f6-09850c02043b","order_by":1,"name":"Xin-li Li","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Xin-li","middleName":"","lastName":"Li","suffix":""},{"id":615464617,"identity":"66b67703-7bf0-41bc-94e0-b98e6c6151bb","order_by":2,"name":"Cheng Qian","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Cheng","middleName":"","lastName":"Qian","suffix":""},{"id":615464618,"identity":"04beaa47-79ae-4299-bcf8-07c073dbdbb7","order_by":3,"name":"Si-min Xue","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Si-min","middleName":"","lastName":"Xue","suffix":""},{"id":615464619,"identity":"69d86668-9fc3-495a-8a2d-d5706158ca4f","order_by":4,"name":"Zhi-tong Zhang","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Zhi-tong","middleName":"","lastName":"Zhang","suffix":""},{"id":615464620,"identity":"1ff5e9ea-554c-4283-8425-afdefbc06b86","order_by":5,"name":"Si-yi Liu","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Si-yi","middleName":"","lastName":"Liu","suffix":""},{"id":615464623,"identity":"89b15603-25af-4cdc-8f75-4e3c40be2845","order_by":6,"name":"Guo Zhang","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Guo","middleName":"","lastName":"Zhang","suffix":""},{"id":615464624,"identity":"1023e4d7-af9e-4137-9588-bf5c867ba777","order_by":7,"name":"Guo-ping Peng","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Guo-ping","middleName":"","lastName":"Peng","suffix":""},{"id":615464627,"identity":"e2d3e86e-ae75-4171-accf-6c56474f36da","order_by":8,"name":"Ning Ding","email":"","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Ning","middleName":"","lastName":"Ding","suffix":""},{"id":615464628,"identity":"aa24f969-3970-4676-97e3-c95f0868feff","order_by":9,"name":"Yun-feng Zheng","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA80lEQVRIiWNgGAWjYLCCCiDmZ2BIgPAOEKPlDBBLNpCsxQCukpAW+Rk5Zg8OVNyx23y74ZnEzxwGOb4bCYyfC/BoYZyRY25w4Myz5G13DqRJ9m5jMJa8kcAsPQOPFmaJHDPpj22Hk81uJKRJM25jSNxwI4GNmQePFjagFomD/w4nG8+AaKknqIUHrKXhsJ2BBERLggEhLRI8z8okDhw7nCBxIyHZsnebhOHMMw+bpfFpkW9P3iZxoOawPf+MnMQbP7fZyPMdTz74GZ8WBoEEMJXYwMADYkkAMWMDPg3AhHIATNkzMLAfwK9yFIyCUTAKRiwAACIbTyW1KsLMAAAAAElFTkSuQmCC","orcid":"","institution":"Nanjing University of Chinese Medicine","correspondingAuthor":true,"prefix":"","firstName":"Yun-feng","middleName":"","lastName":"Zheng","suffix":""}],"badges":[],"createdAt":"2026-03-19 09:53:44","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9167929/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9167929/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105956956,"identity":"f745b51f-4eb5-4286-a865-e94e69c5c9ac","added_by":"auto","created_at":"2026-04-01 20:40:10","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":110560,"visible":true,"origin":"","legend":"\u003cp\u003eInclusion Complex Preparation\u003c/p\u003e\n\u003cp\u003eⅠ:Complete Inclusion Group Ⅱ: Incomplete Enclosure Group Ⅲ: Free Monomer Group\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/ea4ee364363f95d62fb96fe5.jpg"},{"id":105956957,"identity":"8c16ccbb-4764-4ada-8d40-ccd80ab5c3e1","added_by":"auto","created_at":"2026-04-01 20:40:10","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":239628,"visible":true,"origin":"","legend":"\u003cp\u003eCharacterization of Supramolecular Complexes\u003c/p\u003e\n\u003cp\u003eA: Differential Scanning Calorimetry \u0026nbsp;\u0026nbsp;B: X-ray diffraction\u003c/p\u003e\n\u003cp\u003eC: Thermogravimetric Analysis \u0026nbsp;\u0026nbsp;\u0026nbsp;D: Fourier Transform Infrared Spectroscopy\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/8f1d7cf21a919b718ab0ffa4.jpg"},{"id":105956964,"identity":"36d6dad9-076b-427e-b360-54a963a26d67","added_by":"auto","created_at":"2026-04-01 20:40:11","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":464541,"visible":true,"origin":"","legend":"\u003cp\u003eScanning Electron Microscope\u003c/p\u003e\n\u003cp\u003eA: β-Cyclodextrin \u0026nbsp;B:Paeonol \u0026nbsp;C:Free Monomer Group \u0026nbsp;D:Supramolecular Complex\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/8099ed67e88f4b8ce36a6a4c.jpg"},{"id":105956962,"identity":"6f0f2105-0533-4893-a66d-b206b25bb15a","added_by":"auto","created_at":"2026-04-01 20:40:11","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":166993,"visible":true,"origin":"","legend":"\u003cp\u003eRaw Spectral Analysis\u003c/p\u003e\n\u003cp\u003e(A): Raw Spectra\u003c/p\u003e\n\u003cp\u003e(B): Raw Spectra Principal Component Analysis (PCA)\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/f5c3c789cc034a9c70063344.jpg"},{"id":106093338,"identity":"5ab3e040-b56c-4ee2-9428-8b232875f08e","added_by":"auto","created_at":"2026-04-03 11:36:53","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":197325,"visible":true,"origin":"","legend":"\u003cp\u003eClassification Confusion Matrix Diagram\u003c/p\u003e\n\u003cp\u003eA&E SVM Model B&F LSTM- Bayesian Model\u003c/p\u003e\n\u003cp\u003eC&G CNN- Bayesian Model D&H ELM Model\u003c/p\u003e","description":"","filename":"Picture5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/c777bdba8305e0d795fff072.jpg"},{"id":105956959,"identity":"e11d4342-b838-4ca7-b0ff-4c6e5a441238","added_by":"auto","created_at":"2026-04-01 20:40:10","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":164319,"visible":true,"origin":"","legend":"\u003cp\u003eSpectral Data Preprocessing\u003c/p\u003e","description":"","filename":"Picture6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/823f5d502d991bf0b0c5b6ca.jpg"},{"id":106093354,"identity":"b13e29a4-c9fc-483b-b235-193a4e011f63","added_by":"auto","created_at":"2026-04-03 11:36:55","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":158201,"visible":true,"origin":"","legend":"\u003cp\u003e(I) ICO+SNV Characteristic Wavelength Selection Results: A: Sampling weights for each feature interval during optimization. B: Feature intervals selected by the ICO algorithm.\u003c/p\u003e\n\u003cp\u003eC: Variation in RMSECV.\u003c/p\u003e","description":"","filename":"Picture7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/f81f45b05b9f198242bbf99d.jpg"},{"id":105956961,"identity":"59d11c5b-d919-464d-99df-3f8deb62ea06","added_by":"auto","created_at":"2026-04-01 20:40:10","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":193870,"visible":true,"origin":"","legend":"\u003cp\u003eComplete Inclusion Group Inclusion Complex ICO+SNV PLSR Model\u003c/p\u003e\n\u003cp\u003eA: Training Set Visualization \u0026nbsp;B: Prediction Set Visualization\u003c/p\u003e\n\u003cp\u003eC: Training Set Results Comparison \u0026nbsp;D: Prediction Set Results Comparison Chart\u003c/p\u003e","description":"","filename":"Picture8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/ff680dd62757062cba9a4960.jpg"},{"id":106401927,"identity":"5b8db68d-b0b7-4db4-86be-7d9c783732a2","added_by":"auto","created_at":"2026-04-08 09:10:14","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3277556,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/fce5603e-657b-4cf4-83c0-e84cf50abdee.pdf"},{"id":106093238,"identity":"0e5ab479-4007-468d-b748-dfdf295d9794","added_by":"auto","created_at":"2026-04-03 11:36:15","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":418029,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary material\u003c/p\u003e\n\u003cp\u003eSupplementary data 1. Information on data related to the highperformance liquid chromatography\u003c/p\u003e","description":"","filename":"SupplementaryMaterials.docx","url":"https://assets-eu.researchsquare.com/files/rs-9167929/v1/c55befd8483ab4421d0f5ac6.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Rapid evaluation of the comprehensive quality of Paeonol/Cyclodextrin supramolecular complexes using CASSA based on near-infrared spectroscopy combined with artificial intelligence algorithm","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eVolatile components are an important active ingredient in traditional Chinese medicine. However, because its primary active ingredients are volatile constituents, this property often results in poor stability and susceptibility to oxidation after extraction when used as a medicinal product, making it difficult to control its pharmaceutical quality. To address these issues, an effective approach involves using β-cyclodextrin to package volatile components into supramolecular complexes (also known as inclusion complexes). This method ensures the stability of volatile components while enabling their absorption by the human body; moreover, it is widely used owing to its simple process and low cost. Examples include 2-hydroxypropyl-β-cyclodextrin (HP-β-CD) encapsulating nutmeg (Xi et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Yet for questions including the contents of volatile components in supramolecular complexes and whether inclusion has been fully achieved, solutions primarily depend on techniques such as scanning electron microscopy, X-ray diffraction, differential scanning calorimetry (Xu et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), and HPLC\u0026mdash;all of which are often cumbersome, time-consuming, and require both sophisticated equipment and highly skilled operators. Additionally, it remained difficult to determine whether free volatile monomer components were present during HPLC-based content analysis of the complex, thus leading to issues such as inconsistent and unreliable content results. Therefore, for the inclusion of volatile components today, there is an urgent need for a rapid and reliable technique to detect the inclusion state and content of supramolecular complexes.\u003c/p\u003e \u003cp\u003eNear-infrared spectroscopy (NIRS) is a technique that reflects absorption information from the harmonic and sum frequencies of molecular vibrations in chemical groups containing hydrogen elements (such as C-H, O-H, S-H, N-H, etc.). Consequently, this spectroscopic analysis covers nearly all organic compounds and mixtures (Kato et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Liu et al., \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Furthermore, NIRS technology enables rapid scanning of samples with minimal pretreatment requirements, yielding stable, reliable, and accurate data. Currently, near-infrared technology is widely used in various fields such as rapid quantitative analysis of target components and identification of food and pharmaceutical products (Beć et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). For example, near-infrared technology can be used to rapidly detect adulteration in ginger powder and quantify its content (Yu et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). NIRS technology can also be used for rapid quantitative analysis of herbal medicine components. Artificial intelligence algorithms can establish models based on a series of data, gradually arriving at optimal solutions through repeated trial-and-error learning (Yang et al., \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Therefore, leveraging the vast data provided by near-infrared scanning technology and continuously optimizing models through artificial intelligence learning ultimately yields a relatively reliable and accurate mathematical model, which is a great solution.\u003c/p\u003e \u003cp\u003ePeony root bark is the dried root bark of \u003cem\u003ePaeonia suffruticosa Andr.\u003c/em\u003e, a plant belonging to the Ranunculaceae family, possessing rich edible and medicinal value. Its primary component, paeonol, is a volatile compound exhibiting antibacterial, anti-inflammatory, antioxidant, and anticancer effects (Xu et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Bai et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Yu et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Owing to its instability, this study formed a Paeonol/Cyclodextrin supramolecular complex (P-CD complex\u003cb\u003e)\u003c/b\u003e by cyclodextrin inclusion of paeonol. The quality of the complex was evaluated using techniques such as X-ray diffraction and HPLC. Then, by integrating NIRS technology with artificial intelligence algorithms, a classification model was established capable of distinguishing three types: fully encapsulated paeonol, partially encapsulated paeonol, and free monomer groups (i.e., paeonol completely unencapsulated by cyclodextrin). Using preprocessing methods and feature wavelength extraction techniques, establish a PLSR quantitative model for the complexes in the fully encapsulated group. B\u0026thinsp;=\u0026thinsp;y combining NIRS technology with artificial intelligence algorithms, this approach enables rapid assessment of inclusion complexes formed between cyclodextrin and paeonol, as well as determination of the complex's content. Our institute has developed a method that is rapid and simple to operate, yielding accurate and reliable results with broad applicability. Also, it enables comprehensive quality evaluation through a single spectral acquisition, known as CASSA. This approach provides technical guidance and conceptual insights for comprehensive studies on the properties of supramolecular complexes.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Instruments and Reagents\u003c/h2\u003e \u003cp\u003eAntaris II FT-NIR Spectrometer (Thermo Fisher Scientific, USA); Waters e2695 HPLC with a 2998PDA detector (Waters, USA) were utilized. Chromatography-grade methanol was purchased from Tedia, USA; X-ray diffractometer (Bruker Corporation, Germany, Model: D8 Advance); Field-emission scanning electron microscope (Hitachi High-Tech Corporation, Japan, Model: SU8600); Fourier transform infrared spectrometer (Thermo Fisher Scientific, U.S.A., Model: Nicolet iS5); Differential scanning calorimeter (TA Instruments, U.S.A., Model: DSC 2500); Thermogravimetric analyzer (TA Instruments, U.S.A., Model: TGA 55);ultrapure water was prepared using a Millipore Milli-Q ultrapure water system; paeonol-rich bark extract monomer (or: Paeonia suffruticosa bark extract monomer, paeonol-containing) was purchased from Ruimao Biotechnology Co., Ltd.; paeonol-rich bark extract reference standard (or: Paeonia suffruticosa bark extract reference standard) was obtained from the China National Institute for Food and Drug Control (Batch No.: 110708\u0026ndash;202309, purity\u0026thinsp;\u0026ge;\u0026thinsp;99.9%); β-cyclodextrin was purchased from Anhui Shanhe Pharmaceutical Excipients Co., Ltd.; all other reagents were of analytical grade and purchased from Sinopharm Chemical Reagent Co., Ltd.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Sample Preparation\u003c/h2\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1 Preparation of completely encapsulated paeonol samples\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTake an appropriate amount of paeonol extract, precisely weigh the mass, dissolve in an appropriate amount of anhydrous ethanol, add to a saturated β-cyclodextrin aqueous solution at 50\u0026deg;C, stir for 30 minutes, filter under vacuum after refrigerating for 24 hours, wash the filter cake with ethanol, dry at 60\u0026deg;C for 15 hours, and grind through a No. 4 sieve to obtain the paeonol complex. Add cyclodextrin passed through the No. 4 sieve and mix thoroughly. Prepare 160 portions of the complex (I). As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e (I).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e2.2.2 Preparation of Partially Encapsulated paeonol Samples\u003c/h2\u003e \u003cp\u003eTake the complex prepared in Section \u003cspan refid=\"Sec5\" class=\"InternalRef\"\u003e2.2.1\u003c/span\u003e, add the ground paeonol monomer (passed through a No. 4 sieve), and mix thoroughly. Prepare 120 samples of partially encapsulated paeonol (II)\u0026mdash;i.e., samples where the complex and paeonol coexist\u0026mdash;with the results shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e (II).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.2.3 Preparation of Paeonol Monomer and Cyclodextrin Monomer Samples\u003c/h2\u003e \u003cp\u003eTake cyclodextrin passed through a No. 4 sieve and ground paeonol monomer passed through a No. 4 sieve separately, then mix them thoroughly. Prepare 40 portions of the mixture of paeonol monomer and cyclodextrin monomer (III). As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e(III).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e2.2.4 Solution preparation\u003c/h2\u003e \u003cp\u003e \u003cstrong\u003eSample Solution Preparation\u003c/strong\u003e \u003cp\u003eApproximately 0.02 g of the paeonol complex was accurately weighed and placed in a stoppered Erlenmeyer flask. Precisely 100 mL of methanol was added, the flask was sealed tightly, and its weight was recorded. After sonication for 20 minutes, the weight loss was replenished with methanol. The mixture was filtered, and 1 mL of the subsequent filtrate was precisely transferred to a 10 mL volumetric flask. Methanol was added to dilute the solution to the mark, and the mixture was thoroughly mixed to obtain the test solution.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003ePreparation of Reference Solution\u003c/strong\u003e \u003cp\u003eTake an appropriate amount of paeonol, precisely weigh the mass, place it in a volumetric flask, dilute to the mark with methanol, and mix thoroughly to prepare a solution containing 20 \u0026micro;g of Paeonol per 1 mL.\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e2.2.5 Chromatographic conditions\u003c/h2\u003e \u003cp\u003eA Hedera ODS-2 column (4.6 mm \u0026times; 250 mm, 5 \u0026micro;m) was employed. The mobile phase was composed of methanol and water (60:40, v/v). Detection was carried out at a wavelength of 274 nm. The column temperature was maintained at 30\u0026deg;C. The injection volume was set to 10 \u0026micro;L. The flow rate was adjusted to 0.8 mL/min.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e2.2.6 Methodological examination\u003c/h2\u003e \u003cp\u003eTo verify the reliability of this method, validation tests including linearity, precision, repeatability, stability, and spiked recovery were performed in accordance with the guidelines of the International Conference on Harmonisation (ICH). A detailed description of the method is available in Section \u003cspan refid=\"Sec1\" class=\"InternalRef\"\u003e1\u003c/span\u003e of the Supplementary Materials.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Characterization of P-CD complex\u003c/h2\u003e \u003cp\u003eThe successful formation of the Paeonol/Cyclodextrin supramolecular complex (P-CD complex) was confirmed via scanning electron microscopy (SEM), X-ray diffraction (XRD), Fourier Transform Infrared Spectroscopy(FT-IR) ,and thermogravimetric analysis (TGA).\u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e \u003ch2\u003e2.3.1 Differential Scanning Calorimetry\u003c/h2\u003e \u003cp\u003eApproximately 5\u0026ndash;10 mg of the sample was hermetically sealed in an aluminum pan and analyzed using a differential scanning calorimeter (DSC, TA Instruments, model: DSC 2500). The measurement was conducted from 30 to 295\u0026deg;C at a heating rate of 10\u0026deg;C/min under a constant flow of nitrogen gas (e.g., 50 mL/min) to prevent oxidation. An empty aluminum pan was used as the reference.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e2.3.2 X-ray Diffraction\u003c/h2\u003e \u003cp\u003eAn X-ray diffractometer (Bruker, Germany; model: D8 Advance) was operated at room temperature under the following conditions: Cu Kα radiation, voltage of 40 kV, current of 40 mA, scanning speed of 2\u0026deg;/min, and a scanning range of 2\u0026ndash;50\u0026deg;. The obtained patterns were baseline-corrected and smoothed using MDI Jade 6.5, generating the final XRD patterns.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003e2.3.3 Thermogravimetric Analysis\u003c/h2\u003e \u003cp\u003eThermogravimetric analysis (TGA) was performed on a thermogravimetric analyzer (TA Instruments, model: TGA 55). The measurements were carried out from 30 to 600\u0026deg;C at a heating rate of 20\u0026deg;C/min under a constant nitrogen purge (e.g., 50 mL/min) to prevent sample oxidation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section3\"\u003e \u003ch2\u003e2.3.4 Fourier Transform Infrared Spectroscopy\u003c/h2\u003e \u003cp\u003eAn accurate amount (2 mg) of the sample was mixed thoroughly with pure KBr at a mass ratio of 1:100 (sample: KBr), and the mixture was placed in a mold and pressed into a transparent pellet under a hydraulic press. The Fourier transform infrared (FT-IR) spectrum was recorded on an infrared spectrometer (Thermo Fisher Scientific, Inc., Model: Nicolet iS5) over the range of 400\u0026ndash;4000 cm⁻\u0026sup1; at a resolution of 4 cm⁻\u0026sup1;, with 32 scans.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e2.3.5 Scanning Electron Microscope\u003c/h2\u003e \u003cp\u003eA small amount of the sample was evenly spread on a piece of conductive adhesive tape, and any loosely attached particles were removed using a rubber bulb. The sample was then gold sputter-coated for approximately 60 seconds. SEM images were acquired using a field emission scanning electron microscope (Hitachi High-Technologies Corporation, model: SU8600) in the secondary electron (SE) mode at an acceleration voltage of 5 kV.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Near-infrared spectroscopy acquisition\u003c/h2\u003e \u003cp\u003eThe samples were scanned and collected using an Antaris II FT-NIR spectrometer (Thermo Fisher Scientific, USA) in diffuse reflection mode. Samples were loaded into Fisher Shell Type 1 Glass quartz sample cells (19 \u0026times; 51 mm, 2 dram). Spectra were acquired over the wavelength range of 12,000\u0026ndash;4,000 cm⁻\u0026sup1; with a resolution of 16 cm⁻\u0026sup1;. Each spectrum was scanned 32 times. For each sample, one background scan was performed followed by three parallel scans. The averaged spectrum was used as the experimental data for subsequent analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Establishment of Classification Models\u003c/h2\u003e \u003cp\u003eThe average spectral data acquired via near-infrared spectroscopy (NIRS) exhibits high accuracy without the need for preprocessing or characteristic wavelength selection. Therefore, this study adopted a direct modeling approach that omits data processing, splitting the raw spectral data into a training set and a prediction set at a ratio of 7:3. The data were categorized into three groups: the fully encapsulated group (I), the partially encapsulated group (II), and the free monomer group (III). The optimal classification model was selected using training accuracy and prediction accuracy as evaluation metrics.\u003c/p\u003e \u003cdiv id=\"Sec19\" class=\"Section3\"\u003e \u003ch2\u003e2.5.1 Classification Model Selection\u003c/h2\u003e \u003cp\u003e(1) Bayesian Optimization is a highly effective global optimization algorithm aimed at finding the global optimum solution. It efficiently addresses the classic machine-intelligent problem in sequential decision theory: determining the next evaluation location based on information gathered about the unknown objective function f, thereby achieving the optimum solution most rapidly (Lorenz et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). The Bayesian optimization framework can obtain the optimal solution for complex objective functions with a minimal number of evaluations. Essentially, the framework first employs a surrogate model to approximate the true objective function. It then proactively selects the most \"promising\" evaluation points for sampling based on this approximation, thereby avoiding unnecessary sampling (Dastin-van Rijn \u0026amp; Widge, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Consequently, Bayesian Optimization is also termed active optimization. Additionally, the framework effectively leverages complete historical information to enhance search efficiency.\u003c/p\u003e \u003cp\u003e(2) Long Short-Term Memory (LSTM) and Recurrent Neural Networks (RNN) are primarily used for processing sequential data. Their defining feature is that the output of a neuron at one time step can be fed back as input to the same neuron in the next time step. This recurrent architecture is highly suited for time series data, as it preserves temporal dependencies within the data. When unfolding an RNN, repetitive structures emerge, and parameters in the architecture are shared\u0026mdash;this significantly reduces the number of neural network parameters requiring training. On the other hand, shared parameters also enable the model to scale to data of varying lengths, allowing RNN inputs to be sequences of arbitrary length. For example, when training a fixed-length sentence: a feedforward neural network would assign a separate parameter to each input feature, whereas a recurrent neural network can share the same weight parameters across time steps. Although RNNs were originally designed to learn long-term dependencies, extensive practice has shown that standard RNNs often struggle to retain information over the long term. To address this long-term dependency problem, Hochreiter et al. proposed the LSTM network (Graves \u0026amp; Schmidhuber, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2005\u003c/span\u003e) to enhance traditional recurrent neural network models. Today, LSTM has become one of the most effective sequence models in practical applications. Compared to the hidden units in RNNs, those in LSTM have a more complex internal structure: as information flows through the network, the addition of linear interventions enables LSTM to selectively amplify or diminish information intensity (Malashin et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e(3) Convolutional Neural Network (CNN). The basic structure of CNN consists of an input layer, convolutional layers, pooling layers (also called sampling layers), fully connected layers, and an output layer. Typically, multiple convolutional and pooling layers are employed, arranged alternately\u0026mdash;i.e., one convolutional layer followed by one pooling layer, and so forth. In the convolutional layer, each neuron in the output feature map forms local connections with its local inputs from the previous layer. The neuron's input value is calculated by weighting and summing these local inputs using corresponding connection weights, then adding a bias term. This process is fundamentally equivalent to the convolution operation, from which CNN gets its name (Li et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Zang et al., \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Sun et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e(4) Extreme Learning Machine (ELM) randomly selects the input weights and hidden layer biases of the network, and derives output weights through analytical computation. This effectively overcomes the limitations of traditional SLFN learning algorithms and has been widely applied in various fields such as disease diagnosis, traffic sign recognition, and image quality assessment. Initially limited to single-hidden-layer feedforward neural networks, it was later extended to RBF neural networks, recurrent neural networks, generalized single-hidden-layer feedforward neural networks, and multi-hidden-layer feedforward neural networks. From a design perspective, ELM aims to unify research problems in machine learning\u0026mdash;including regression, classification, clustering, compression, and feature extraction\u0026mdash;within a single framework (Huang et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Chen et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). In terms of learning efficiency, ELM is simple to implement, has an extremely fast learning speed, and requires minimal human intervention.\u003c/p\u003e \u003cp\u003e(5) Support Vector Machine (SVM) (Hsieh \u0026amp; Yeh, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2011\u003c/span\u003e) is a supervised learning algorithm based on statistical learning theory, widely applied in classification and regression tasks. Its core principle involves finding an optimal decision hyperplane that maximizes the margin (i.e., the distance between classes) between samples of different classes, thereby enhancing the model's generalization capability. The key to SVM lies in minimizing structural risk\u0026mdash;not only minimizing low error rates on training data but also reducing model complexity by maximizing the classification margin to prevent overfitting. For linearly separable data, SVM employs hard margin optimization to directly determine a hyperplane that perfectly classifies all samples. In contrast, for linearly inseparable data, it introduces slack variables and a penalty parameter (soft margin, to tolerate minor misclassifications), allowing some samples to be misclassified to enhance model robustness. The SVM optimization problem ultimately transforms into a convex quadratic programming problem, solvable via the Lagrange multiplier method. Its solution is determined by only a few support vectors, and thus exhibits sparsity and computational efficiency. Furthermore, SVM excels in scenarios with small sample sizes and high-dimensional data (e.g., text classification, image recognition), making it one of the classic algorithms in machine learning.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Establishment of Quantitative Models\u003c/h2\u003e \u003cp\u003eGiven the extremely high requirements for reliability and accuracy in quantitative models, such models require preprocessing methods\u0026mdash;including spectral data preprocessing and characteristic wavelength extraction\u0026mdash;to perform dimensionality reduction and optimization on sample data, thereby preventing overfitting. Using the content data of the complex in the fully encapsulated group as the independent variable, a quantitative model was established for the paeonol complex content in this group.\u003c/p\u003e \u003cdiv id=\"Sec21\" class=\"Section3\"\u003e \u003ch2\u003e2.6.1 Preprocessing Method\u003c/h2\u003e \u003cp\u003eThis study used The Unscrambler X 10.4 software (CAMO, Inc., Texas, USA) to preprocess raw spectra. Nine preprocessing methods were applied to the spectra, including first derivative (1st derivative, 1D), second derivative (2nd derivative, 2D), Savitzky-Golay (SG), standard normal variate (SNV), multiplicative scatter correction (MSC), median filter (MF), SNV+1D, MSC+1D, and SG\u0026thinsp;+\u0026thinsp;1D. These nine methods were evaluated to select the one that yielded the highest model fitting accuracy.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section3\"\u003e \u003ch2\u003e2.6.2 Characteristic Wavelength Selection\u003c/h2\u003e \u003cp\u003e(1) Competitive Adaptive Reweighted Sampling (CARS) is a feature selection method based on full-spectrum partial least squares (PLS) regression. By mimicking Darwinian evolution algorithms, it selects the optimal combination of effective variables in the spectrum: specifically, it retains wavelength points with larger absolute regression coefficients in the PLS model while removing those with smaller weights in the same model. Through iterative and competitive operations, it selects N wavelength subsets in each iteration and identifies the subset with the lowest root mean square error of cross-validation (RMSECV) value, which corresponds to the optimal number of variables (Li et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2009\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e(2) The Iterative Retaining Informative Variables (IRIV) method (Yun et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2014\u003c/span\u003e) is a feature selection algorithm based on the Binary Matrix Shuffle Filter (BMSF). It comprehensively accounts for the importance of each variable, and after optimization, further screens feature wavelength variables to determine the optimal number of features. As an effective method for selecting optimal feature wavelengths in near-infrared spectroscopy (NIRS), IRIV adopts a strategy of randomly combining variables to account for potential interactions between them, thereby classifying all variables into four categories: strong informative variables, weak informative variables, non-informative variables, and interfering variables. Through multiple iterations\u0026mdash;each aimed at retaining strong and weak informative variables while eliminating non-informative and interfering variables\u0026mdash;the optimal variable set is ultimately obtained via backward elimination.\u003c/p\u003e \u003cp\u003e(3) Interval Combination Optimization (ICO) is one of the commonly used interval variable selection methods, as it can provide more reasonable wavelength intervals and enhance the predictive capability of models. Proposed within the framework of Model Portfolio Analysis (MPA) combined with Weighted Bootstrap Sampling (WBS), ICO leverages the strengths of both approaches. As a general framework for variable selection and evaluation, MPA offers comprehensive insights from a large number of submodels; its core concept is to establish data analysis methods by statistically analyzing the distribution of these submodels, making interval selection methods more advantageous than single-wavelength selection approaches. Weighted Bootstrap Sampling (WBS) is a random sampling technique in which different objects are assigned weights based on their occurrence frequencies. As an interval selection method, ICO searches through the interval weights across the entire variable space and selects useful subintervals (Song et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2016\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e(4) The Successive Projections Algorithm (SPA) is a forward selection method\u0026mdash;an iterative forward technique that minimizes multicollinearity\u0026mdash;traditionally used for variable selection in multivariate calibration. It starts with a single wavelength and adds a new wavelength in each iteration until a specified number (N) of wavelengths is reached. The optimal number of variables is determined by calculating the root mean square error of cross-validation (RMSECV) for the N-wavelength subset using multiple linear regression (MLR) (Canova et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Xiaobo et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2010\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003e2.6.3 Establishment of the PLSR Quantitative Model\u003c/h2\u003e \u003cp\u003ePartial Least Squares Regression (PLSR) analysis was performed using MATLAB 2024a software (MathWorks Inc., Natick, MA, USA) to enable quantitative prediction of the compound. The Kennard and Stone (KS) algorithm was employed to divide all samples into a training set and a test set at a 7:3 ratio, establishing the PLSR model for performance prediction.\u003c/p\u003e \u003cp\u003e \u003cdiv id=\"Equa\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:{R}^{2}=1-\\frac{\\sum\\:{\\left(Yi-\\widehat{Yi}\\right)}^{2}}{\\sum\\:\\left(Yi-Yn\\right)}\\:$$\u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Equb\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:RMSE=\\sqrt{\\frac{\\sum\\:{(\\widehat{Yi}-Yi)}^{2}}{n}}$$\u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Equc\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:RPD=\\frac{SD}{RMSEP}$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn the above formulas, \u003cem\u003eYi\u003c/em\u003e represents the ith reference value, ̂\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{Yi}\\)\u003c/span\u003e\u003c/span\u003e represents the ith predicted value, and \u003cem\u003eYn\u003c/em\u003e represents the mean of n referencevalue.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"3. Results and discussion","content":"\u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Characterization of P-CD complex\u003c/h2\u003e \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003e3.1.1Differential Scanning Calorimetry\u003c/h2\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003e(A), an endothermic peak for β-cyclodextrin at 92\u0026deg;C, as well as two endothermic peaks for paeonol at 50\u0026deg;C and 224\u0026deg;C, are observed in the free monomer group, indicating no changes after physical mixing. When guest molecules are incorporated into the β-CD cavity, their melting and sublimation points may shift to different temperatures or disappear (Nerome et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). In the DSC curve of the inclusion complex, the endothermic peak of paeonol at 50\u0026deg;C vanishes, confirming P-CD complex formation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section3\"\u003e \u003ch2\u003e3.1.2 X-ray Diffraction\u003c/h2\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003e(B), paeonol exhibits strong diffraction peaks in the range of 10\u0026deg; \u0026lt; 2θ\u0026thinsp;\u0026lt;\u0026thinsp;30\u0026deg;, with notably high intensity at 11.72\u0026deg;, 16.41\u0026deg;, 20.95\u0026deg;, 23.58\u0026deg;, and 25.54\u0026deg;. These peaks are characteristic of paeonol, indicating that it is a crystalline compound. The XRD pattern of the free monomer group (physical mixture of paeonol and β-cyclodextrin) represents the superposition of the patterns of paeonol monomer and β-cyclodextrin, confirming that it is merely a simple physical mixture without chemical reactions or structural modifications. The absence of paeonol\u0026rsquo;s characteristic peaks in the inclusion complex confirms that paeonol is encapsulated within the β-cyclodextrin cavity to form an inclusion complex, and consequently loses its original crystalline structure (Cao et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Betlejewska-Kielak et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Guo et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2011\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section3\"\u003e \u003ch2\u003e3.1.3 Thermogravimetric Analysis\u003c/h2\u003e \u003cp\u003eThe TGA curves for each substance are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003e(C). Paeonol exhibits poor thermal stability, with a total mass loss of over 85% at relatively low temperatures (50\u0026ndash;177\u0026deg;C). The mass loss of the P-CD complex occurs in three stages: the first stage accounts for approximately 12% mass loss, attributable to the evaporation of surface and internal water; the second stage occurs between 300\u0026ndash;328\u0026deg;C, with a mass loss of approximately 77%, representing the main weight loss phase, which is associated with the thermal decomposition of the P-CD complex during this stage. The thermal stability of paeonol was significantly enhanced due to interactions between the guest molecules and the inner cavity of β-cyclodextrin, indicating the formation of the inclusion complex (Li et al., \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2016\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec29\" class=\"Section3\"\u003e \u003ch2\u003e3.1.4 Fourier Transform Infrared Spectroscopy\u003c/h2\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003e(D), the infrared spectrum of β-cyclodextrin exhibits four characteristic peaks: a broad hydroxyl (-OH) stretching vibration peak near 3,440 cm⁻\u0026sup1;; methylene (-CH₂) and methyl (-CH₃) stretching vibration peaks near 2,920 cm⁻\u0026sup1;; a peak near 1,642 cm⁻\u0026sup1;, which is attributed to the bending vibration of adsorbed water or hydroxyl groups (not a carbonyl stretch, as β-cyclodextrin contains no carbonyl groups); and an ether bond (-C-O-C) stretching vibration peak near 1,020 cm⁻\u0026sup1;. In the infrared spectrum of paeonol, the broad peak at 2,976 cm⁻\u0026sup1; corresponds to the stretching vibration of the phenolic hydroxyl (-OH) group; the high-intensity absorption peak at 1,620 cm⁻\u0026sup1; represents the stretching vibration of the aromatic ketone carbonyl group, which undergoes a significant bathochromic shift to lower wavenumbers due to conjugation between the benzene ring and the carbonyl group; the strong absorption peak at 1,207 cm⁻\u0026sup1; corresponds to the stretching vibration of the carbon-oxygen (C-O) bond in the aromatic ether. The aromatic ring skeletal stretching vibrations observed in the 1,650\u0026ndash;1,430 cm⁻\u0026sup1; region further confirm the presence of the benzene ring. These analytical results are consistent with the theoretical structure of paeonol. The infrared spectrum of the physical mixture exhibits characteristic peaks similar to those of β-cyclodextrin and paeonol, representing a simple superposition of both compounds. This indicates no significant intermolecular interactions during physical mixing (Cao et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The disappearance of the characteristic peaks of paeonol in the infrared spectrum of the paeonol/β-cyclodextrin supramolecular complexes indicates the formation of the inclusion complex (Mansi et al., 2025).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec30\" class=\"Section3\"\u003e \u003ch2\u003e3.1.5 Scanning Electron Microscope\u003c/h2\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003e(A-D), β-cyclodextrin (A) appears as irregularly shaped lumps, while paeonol (B) consists of fine crystalline particles. The physical mixture (free monomer group) exhibits the characteristic features of both β-cyclodextrin and paeonol (C). In contrast, the P-CD complex (D) appears as polygonal aggregates with smooth surfaces and a layered structure. Its morphology differs from that of both β-cyclodextrin and paeonol, indicating the formation of the P-CD complex (Zhu et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec31\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Classification of Raw Near-Infrared Spectral Analysis\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e(A), no significant differences were observed among the 320 sets of raw near-infrared spectral data\u0026mdash;including 160 fully encapsulated samples, 120 incompletely encapsulated samples, and 40 mixed monomer samples\u0026mdash;making it difficult to distinguish between fully encapsulated and non-fully encapsulated samples. A principal component analysis (PCA) model was established using all raw spectral data, and the results are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e(B). Although the three sample categories could be preliminarily classified\u0026mdash;with Principal Component 1 (PC1) contributing 88.8% and the total cumulative contribution reaching 99.1%\u0026mdash;complete differentiation of the three categories remained unattainable. Furthermore, raw near-infrared spectra often contain extensive sample information, including irrelevant data that may interfere with quantitative spectral analysis. Therefore, we preprocessed the raw spectral data and combined it with artificial intelligence algorithms for characteristic wavelength selection, ultimately establishing a partial least squares regression (PLSR) model to detect the content of the P-CD complex.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec32\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Classification Model Development\u003c/h2\u003e \u003cp\u003eThe experiment employed LSTM-Bayesian optimization, CNN-Bayesian optimization, ELM, and an optimizable SVM model to establish classification models using raw spectral data. The results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eClassification Model Accuracy\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003eTraining set accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c8\" namest=\"c6\"\u003e \u003cp\u003ePrediction Set Accuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClass Ⅰ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eClass Ⅱ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eClass Ⅲ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eClass Ⅰ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eClass Ⅱ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClass Ⅲ\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLSTM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e98.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e92.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e87.5%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e98.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e98.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e92.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e93.3%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eELM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e97.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e96.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e100%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e100%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e100%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e100%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e100%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e100%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003e\u003cb\u003eNote: Bold text indicates the best model for each category.\u003c/b\u003e\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAmong these, the optimal SVM model obtained after 30 iterations achieved a confusion matrix for the three-class classification shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. Each class attained 100% accuracy, with identical accuracy rates for both the training and prediction sets, indicating no overfitting. Among the four selected classification models, the optimizable SVM model demonstrated the highest accuracy and good reliability. The other models failed to completely and correctly classify the three categories. Therefore, this study proposes to adopt the SVM model as the classification model.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec33\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Preprocessing of raw spectral data\u003c/h2\u003e \u003cdiv id=\"Sec34\" class=\"Section3\"\u003e \u003ch2\u003e3.4.1 Preprocessing of Fully Enclosed Group Spectral Data\u003c/h2\u003e \u003cp\u003eBy nine preprocessing methods, PLSR models were established. Among them, the four methods\u0026mdash;1D, 2D, MSC, and SNV\u0026mdash;exhibited higher R\u0026sup2; values as shown in Table-2, demonstrating more pronounced effects. The preprocessed spectra are depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComplete Enclosure Group Preprocessing\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePreprocessing\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003csub\u003eT\u003c/sub\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003csub\u003eP\u003c/sub\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRPD\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e1D\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e13\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e1.00\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.94\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e4.08\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e2D\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e12\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.99\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.95\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e4.54\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3.67\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMSC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e9\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.97\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.92\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e3.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMSC+1D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.45\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSG\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3.94\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSG+1D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSNV\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e23\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e1.00\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.91\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e3.52\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSNV+1D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e-0.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eNote: Bold text indicates the selected preprocessing method.\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec35\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Characteristic Wavelength Selection\u003c/h2\u003e \u003cp\u003eAfter preprocessing near-infrared hyperspectral data, feature wavelengths are selected using multi-intelligence algorithms to extract relevant spectral variables. Models are constructed based on CARS-I RIV, CARS-SPA, and ICO feature wavelength algorithms, and their predictive capabilities are evaluated.\u003c/p\u003e \u003cdiv id=\"Sec36\" class=\"Section3\"\u003e \u003ch2\u003e3.5.1 Fully Enclosed Group Characteristic Wavelength\u003c/h2\u003e \u003cp\u003eFor the fully encapsulated sample group, the results of three feature wavelength extraction methods\u0026mdash;CARS-SPA, CARS-IRIV, and ICO\u0026mdash;are presented in Table\u0026nbsp;4. Among these methods, the combination of ICO (feature wavelength extraction) and SNV (spectral preprocessing) yielded the following model parameters: R\u0026sup2;\u003csub\u003eT\u003c/sub\u003e=0.9757, R\u0026sup2;\u003csub\u003eP\u003c/sub\u003e=0.9476, RMSEP/RMSEC\u0026thinsp;=\u0026thinsp;1.04, and RPD\u0026thinsp;=\u0026thinsp;4.43. These parameter results indicate that the model is accurate and reliable, with excellent predictive performance and a reasonable fitting level.\u003c/p\u003e \u003cp\u003ePreliminary dimensionality reduction of the raw spectral data via preprocessing simplified data complexity. Subsequently, during the iterations of the ICO algorithm, the weight coefficients of each wavelength band changed as the number of iterations increased. A yellower color indicates a weight coefficient closer to 1, while a bluer color indicates a weight coefficient closer to 0; if the color is between blue and yellow, the weight coefficient falls between 0 and 1. The width of the finally selected wavelength intervals (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eI-A) was automatically optimized using a local search strategy integrated into the ICO algorithm. A total of 465 feature wavelengths were selected (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eI-C), which significantly reduced information complexity compared to the original 1557 wavelengths and further reduced data dimensionality. This indicates that the ICO algorithm not only rapidly extracts spectral data from complex compound samples but also refines spectral information based on this foundation.\u003c/p\u003e \u003cp\u003eIn summary, selecting the ICO\u0026thinsp;+\u0026thinsp;SNV algorithm as the PLSR quantitative model for establishing a complete complex is the most effective and reliable approach.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eEvaluation of PLSR Parameters for Complete Inclusion Complexes\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"10\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ePreprocessing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eVariables\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e \u003cp\u003eCalibration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c8\" namest=\"c7\"\u003e \u003cp\u003ePrediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eRPD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eRMSEP/RMSEC\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRSMEC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003csub\u003eT\u003c/sub\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eRMSEP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003csub\u003eP\u003c/sub\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eICO\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e310\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9648\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.9021\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e3.20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e366\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.9167\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e3.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e82309.25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMSC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e456\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.8734\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e2.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e743.23\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eSNV\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e11\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e465\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e3.46\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.9757\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e3.59\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.9476\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e4.43\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003e1.04\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eIRIV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.8801\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.6123\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e1.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e8.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.6466\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e1.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMSC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9931\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.9658\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e5.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSNV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9733\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.8415\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e2.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eSPA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e8.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.6488\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e1.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9034\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.5547\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e1.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.34\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMSC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9862\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e3.37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.9471\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e4.35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSNV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4.09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9670\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.8943\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e3.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e1.16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"10\"\u003e\u003cb\u003eNote: Bold text indicates the optimal PLSR model results.\u003c/b\u003e\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec37\" class=\"Section2\"\u003e \u003ch2\u003e3.6 Visualization of PLSR Quantitative Models\u003c/h2\u003e \u003cdiv id=\"Sec38\" class=\"Section3\"\u003e \u003ch2\u003e3.6.1 Complete Enclosure Group PLSR Analysis\u003c/h2\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, the scatter plots of the PLSR models for the P-CD complex (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eA\u0026amp;B) reveal that all samples cluster closely around the regression lines\u0026mdash;indicating the successful establishment of the regression models and strong predictive capability. Additionally, the high degree of overlap between the red and blue lines demonstrates the close agreement between the predicted values and reference values of the samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eD\u0026amp;E). Collectively, these results confirm that the model established for the fully encapsulated group samples (following preprocessing and characteristic wavelength selection) is reliable and possesses robust predictive capability. Therefore, this study ultimately adopted the SNV\u0026thinsp;+\u0026thinsp;ICO-PLSR method to establish a quantitative model for the complex content in the fully encapsulated complex group.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Conclusion","content":"\u003cp\u003eIn this study, the combination of spectral data obtained via FT-NIR rapid scanning and deep learning multi-intelligence algorithms not only addressed the difficulty in evaluating the inclusion state of Paeonol after cyclodextrin encapsulation but also resolved inaccuracies in complex content detection. This approach enabled comprehensive analysis of single spectral acquisition (CASSA) while improving both reliability and accuracy. First of all, categorical analysis of the raw NIR spectral data successfully achieved the classification of three inclusion states: complete inclusion, incomplete inclusion, and free monomer. Second, the raw spectral data were subjected to dimensionality reduction and optimization using preprocessing and characteristic wavelength selection methods, which enhanced the predictive capability and accuracy of the regression model. A PLSR quantitative model for P-CD complex content was established for the fully encapsulated group. Based on the model parameters, this established model exhibits good linearity and predictive capability, demonstrating high reliability.\u003c/p\u003e \u003cp\u003eIn conclusion, the FT-NIR technology combined with artificial intelligence algorithms employed in this study establishes predictive models for cyclodextrin-included volatile components. Compared to traditional methods for complex characterization and content detection, this approach enables comprehensive, rapid, and accurate multi-parameter evaluation of complex quality in a single measurement. It facilitates rapid tracking of complexes, ensures their pharmaceutical stability, and provides a more efficient solution for establishing quality standards for these compounds.\u003c/p\u003e \u003cp\u003e \u003cb\u003eCRediT authorship contribution statement\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eBo-yao Zhang\u003c/b\u003e: Conceptualization, Investigation, Methodology, Visualization, Writing \u0026ndash; original draft. \u003cb\u003eXin-li Li\u003c/b\u003e: Data curation, Formal analysis, Software, Project administration. \u003cb\u003eCheng Qian\u003c/b\u003e: Methodology, Validation, Visualization. \u003cb\u003eSi-min Xue\u003c/b\u003e: Conceptualization, Investigation, Data curation. \u003cb\u003eZhi-tong Zhang\u003c/b\u003e: Project administration, Methodology. \u003cb\u003eSi-yi Liu\u003c/b\u003e: Visualization. \u003cb\u003eGuo Zhang\u003c/b\u003e: Data curation. \u003cb\u003eGuo-ping Peng\u003c/b\u003e: Supervision, Writing \u0026ndash; review and editing. \u003cb\u003eNing Ding\u003c/b\u003e: Supervision, Formal analysis. \u003cb\u003eYun-feng Zheng\u003c/b\u003e: Writing \u0026ndash; review and editing, Methodology, Funding acquisition.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eDeclaration of competing interest\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThe author(s) declare that financial support was received for the research and/or publication of this article. This research study was supported by the National Natural Science Foundation of China (No. 81473390),and Jiangsu \u0026ldquo;333\u0026rdquo; project.\u003c/p\u003e\n\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\n\u003cp\u003eBo-yao Zhang: Conceptualization, Investigation, Methodology, Visualization, Writing \u0026ndash; original draft. Xin-li Li: Data curation, Formal analysis, Software, Project administration. Cheng Qian: Methodology, Validation, Visualization. Si-min Xue: Conceptualization, Investigation, Data curation. Zhi-tong Zhang: Project administration, Methodology. Si-yi Liu: Visualization. Guo Zhang: Data curation. Guo-ping Peng: Supervision, Writing \u0026ndash; review and editing. Ning Ding: Supervision, Formal analysis. Yun-feng Zheng: Writing \u0026ndash; review and editing, Methodology, Funding acquisition.\u003c/p\u003e\n\u003ch2\u003eData availability\u003c/h2\u003e\n\u003cp\u003eThe data that has been used is confidential.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBai D, Li X, Wang S, Zhang T, Wei Y, Wang Q, Dong W, Song J, Gao P, Li Y, Wang S, Dai L (2022) Advances in extraction methods, chemical constituents, pharmacological activities, molecular targets and toxicology of volatile oil from Acorus calamus var. angustatus Besser. Front Pharmacol 13:1004529\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeć KB, Grabska J, Huck CW (2022) Miniaturized NIR Spectroscopy in Food Analysis and Quality Control: Promises, Challenges, and Perspectives. Foods, \u003cem\u003e11\u003c/em\u003e(10)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBetlejewska-Kielak K, Bednarek E, Budzianowski A, Michalska K, Maurin JK (2021) Comprehensive Characterisation of the Ketoprofen-β-Cyclodextrin Inclusion Complex Using X-ray Techniques and NMR Spectroscopy. Molecules, \u003cem\u003e26\u003c/em\u003e(13)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCanova LDS, Vallese FD, Pistonesi MF, de Ara\u0026uacute;jo Gomes A (2023) An improved successive projections algorithm version to variable selection in multiple linear regression. Anal Chim Acta 1274:341560\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao C, Deng C, Xuan F, Zhou Y (2023) Structural characterization and molecular dynamics simulation of inclusion complex of large ring cyclodextrin and alpha-tocopherol. China Oils Fats 48(5):49\u0026ndash;55 Article 1003\u0026ndash;7969(2023)48:5\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen H, Peng J, Zhou Y, Li L, Pan Z (2014) Extreme learning machine for ranking: generalization analysis and applications. Neural Netw 53:119\u0026ndash;126\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDastin-van Rijn EM, Widge AS (2025) Failure modes and mitigations for Bayesian optimization of neuromodulation parameters. J Neural Eng, \u003cem\u003e22\u003c/em\u003e(3)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGraves A, Schmidhuber J (2005) Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Netw 18(5\u0026ndash;6):602\u0026ndash;610\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo P, Su Y, Cheng Q, Pan Q, Li H (2011) Crystal structure determination of the β-cyclodextrin-p-aminobenzoic acid inclusion complex from powder X-ray diffraction data. Carbohydr Res 346(7):986\u0026ndash;990\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHsieh TJ, Yeh WC (2011) Knowledge discovery employing grid scheme least squares support vector machines based on orthogonal design bee colony algorithm. IEEE Trans Syst Man Cybern B Cybern 41(5):1198\u0026ndash;1212\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang G, Huang GB, Song S, You K (2015) Trends in extreme learning machines: a review. Neural Netw 61:32\u0026ndash;48\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKato M, Kim J, Oh J, Shimizu D, Fukui N, Shinokubo H (2023) Near-Infrared-Responsive Hydrocarbons Designed by π-Extension of Indeno[1,2,3,4-pgra] perylene at the 1,2,12-Positions. Chemistry, 29(23), e202300249\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Liang Y, Xu Q, Cao D (2009) Key wavelengths screening using competitive adaptive reweighted sampling method for multivariate calibration. Anal Chim Acta 648(1):77\u0026ndash;84\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu CH, He Z, Ruchlin C, Che Y, Somers K, Perepichka DF (2023) Thiele's Fluorocarbons: Stable Diradicaloids with Efficient Visible-to-Near-Infrared Fluorescence from a Zwitterionic Excited State. J Am Chem Soc 145(29):15702\u0026ndash;15707\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Li F, Cai HY, Chen X, Sun W, Shen WY (2016) Structural characterization of inclusion complex of arbutin and hydroxypropyl-β-cyclodextrin. Trop J Pharm Res 15(10):2227\u0026ndash;2233\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Z, Liu F, Yang W, Peng S, Zhou J (2022) A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects. IEEE Trans Neural Netw Learn Syst 33(12):6999\u0026ndash;7019\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLomarat P, Phechkrajang C, Sunghad P, Anantachoke N (2024) Raman spectroscopy coupled with the PLSR model: A rapid method for analyzing gamma-oryzanol content in rice bran oil. Food Chem X 24:101923\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLorenz R, Hampshire A, Leech R (2017) Neuroadaptive Bayesian Optimization and Hypothesis Testing. Trends Cogn Sci 21(3):155\u0026ndash;167\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMalashin I, Tynchenko V, Gantimurov A, Nelyub V, Borodulin A (2024) Applications of Long Short-Term Memory (LSTM) Networks in Polymeric Sciences: A Review. Polym (Basel), \u003cem\u003e16\u003c/em\u003e(18)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMansi, Khanna P, Yadav S, Singh A, Khanna L (2025) Inclusion complexes of novel formyl chromone Schiff bases with β-Cyclodextrin: Synthesis, characterization, DNA binding studies and in-vitro release study. Carbohydr Polym, \u003cem\u003e347\u003c/em\u003e, 17, Article 122667.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNerome H, Machmudah S, Wahyudiono, Fukuzato R, Higashiura T, Youn YS, Lee YW, Goto M (2013) Nanoparticle formation of lycopene/β-cyclodextrin inclusion complex using supercritical antisolvent precipitation. J Supercrit Fluids 83:97\u0026ndash;103\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSankom A, Mahakarnchanakul W, Rittiron R, Sajjaanantakul T, Thongket T (2021) Detection of Profenofos in Chinese Kale, Cabbage, and Chili Spur Pepper Using Fourier Transform Near-Infrared and Fourier Transform Mid-Infrared Spectroscopies. ACS Omega 6(40):26404\u0026ndash;26415\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong X, Huang Y, Yan H, Xiong Y, Min S (2016) A novel algorithm for spectral interval combination optimization. Anal Chim Acta 948:19\u0026ndash;29\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun Y, Xue B, Zhang M, Yen GG, Lv J (2020) Automatically Designing CNN Architectures Using the Genetic Algorithm for Image Classification. IEEE Trans Cybern 50(9):3840\u0026ndash;3854\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu K, Ma F, Wei C, Gan F, Du C (2023) Rapid Determination of Nitrate Nitrogen Isotope in Water Using Fourier Transform Infrared Attenuated Total Reflectance Spectroscopy (FTIR-ATR) Coupled with Deconvolution Algorithm. Molecules, \u003cem\u003e28\u003c/em\u003e(2)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiaobo Z, Jiewen Z, Povey MJ, Holmes M, Hanpin M (2010) Variables selection methods in near-infrared spectroscopy. Anal Chim Acta 667(1\u0026ndash;2):14\u0026ndash;32\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXi X, Huang J, Zhang S, Lu Q, Fang Z, Li C, Zhang Q, Liu Y, Chen H, Liu A, Liu S, Wang C, Li S, Hu B (2023) Preparation and characterization of inclusion complex of Myristica fragrans Houtt. (nutmeg) essential oil with 2-hydroxypropyl-β-cyclodextrin. Food Chem 423:136316\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu F, Yang Q, Wu L, Qi R, Wu Y, Li Y, Tang L, Guo DA, Liu B (2017) Investigation of Inclusion Complex of Patchouli Alcohol with β-Cyclodextrin. PLoS ONE, 12(1), e0169578\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu Y, Bian S, Shang L, Wang X, Bai X, Zhang W (2024) Phytochemistry, pharmacological effects and mechanism of action of volatile oil from Panax ginseng C.A.Mey: a review. Front Pharmacol 15:1436624\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu Y, Liu J, Sun Y, Chen S, Miao X (2023) Fast detection of volatile fatty acids in biogas slurry using NIR spectroscopy combined with feature wavelength selection. Sci Total Environ 857(Pt 1):159282\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang J, Ma X, Guan H, Yang C, Zhang Y, Li G, Li Z, Lu Y (2024) A quality detection method of corn based on spectral technology and deep learning model. Spectrochim Acta Mol Biomol Spectrosc 305:123472\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu DX, Guo S, Zhang X, Yan H, Zhang ZY, Chen X, Chen JY, Jin SJ, Yang J, Duan JA (2022) Rapid detection of adulteration in powder of ginger (Zingiber officinale Roscoe) by FT-NIR spectroscopy combined with chemometrics. Food Chem X 15:100450\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu X, Sun S, Guo Y, Liu Y, Yang D, Li G, L\u0026uuml; S (2018) Citri Reticulatae Pericarpium (Chenpi): Botany, ethnopharmacology, phytochemistry, and pharmacology of a frequently used traditional Chinese medicine. J Ethnopharmacol 220:265\u0026ndash;282\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYun YH, Wang WT, Tan ML, Liang YZ, Li HD, Cao DS, Lu HM, Xu QS (2014) A strategy that iteratively retains informative variables for selecting optimal variable subset in multivariate calibration. Anal Chim Acta 807:36\u0026ndash;43\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZang B, Ding L, Feng Z, Zhu M, Lei T, Xing M, Zhou X (2021) CNN-LRP: Understanding Convolutional Neural Networks Performance for Target Recognition in SAR Images. \u003cem\u003eSensors (Basel)\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e(13)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu W, Wu J, Guo X, Sun X, Li Q, Wang J, Chen L (2020) Development and physicochemical characterization of chitosan hydrochloride/sulfobutyl ether-β-cyclodextrin nanoparticles for cinnamaldehyde entrapment. J Food Biochem, 44(6), e13197\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"molecular-diversity","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"modi","sideBox":"Learn more about [Molecular Diversity](http://link.springer.com/journal/11030)","snPcode":"11030","submissionUrl":"https://submission.nature.com/new-submission/11030/3","title":"Molecular Diversity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"FT-NIR, Artificial Intelligence Algorithms, Cyclodextrin, CASSA, supramolecular complexes","lastPublishedDoi":"10.21203/rs.3.rs-9167929/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9167929/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePreparing volatile component/cyclodextrin supramolecular complexes is a common method for enhancing the stability of volatile components. However, the quality assessment of supramolecular complexes is highly complex. This study first prepared paeonol/cyclodextrin supramolecular complexes and evaluated their overall quality. Then, Fourier Transform Near-Infrared Spectroscopy (FT-NIR) was combined with artificial intelligence (AI) to build a Support Vector Machine Classification (SVM) model and a Partial Least Squares Regression (PLSR) model. The SVM classification model reached 100% accuracy, whereas the PLSR quantitative model demonstrated R\u0026sup2; \u0026gt; 0.90 for both calibration and prediction sets. Results confirm that integrating FT-NIR with AI improves the accuracy and reliability of qualitative/quantitative models. Via the Comprehensive Analysis of Single Spectral Acquisition (CASSA) method, dual detection of the complexes\u0026rsquo; formation state and concentration was achieved, enabling rapid, comprehensive quality evaluation. This study demonstrates the excellent prospects of combining FT-NIR technology with the preparation of supramolecular complexes.\u003c/p\u003e","manuscriptTitle":"Rapid evaluation of the comprehensive quality of Paeonol/Cyclodextrin supramolecular complexes using CASSA based on near-infrared spectroscopy combined with artificial intelligence algorithm","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-01 20:40:06","doi":"10.21203/rs.3.rs-9167929/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-09T08:56:51+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-03T06:12:21+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-02T02:55:05+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"121850607068768812595918325210691235968","date":"2026-03-31T17:58:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"40022973200013934996489448682635350819","date":"2026-03-31T15:49:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"8246519357602794177232393278204251014","date":"2026-03-31T01:49:40+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-31T00:45:21+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-21T07:47:39+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-21T07:35:35+00:00","index":"","fulltext":""},{"type":"submitted","content":"Molecular Diversity","date":"2026-03-19T09:47:16+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"molecular-diversity","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"modi","sideBox":"Learn more about [Molecular Diversity](http://link.springer.com/journal/11030)","snPcode":"11030","submissionUrl":"https://submission.nature.com/new-submission/11030/3","title":"Molecular Diversity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"0a9dacc3-4be1-4ae6-a3ff-c83f5cef3937","owner":[],"postedDate":"April 1st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-28T15:39:50+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-01 20:40:06","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9167929","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9167929","identity":"rs-9167929","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00