On the Role of Quantum Entanglement in Capturing Long-Term Dependencies

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Given the current gap in the literature regarding whether quantum computing can effectively address long-term dependency modeling in sequential learning tasks, this study seeks to explicitly investigate the potential benefits of quantum entanglement in this context. We propose a hybrid Quantum Dilated Con- volutional Neural Network (QDCNN) that synergistically integrates quantum entanglement with classical processing, leveraging the complementary strengths of both computational paradigms. Our architecture employs quantum dilated convolutions with exponentially increasing dilation rates, enabling the efficient capture of long-range temporal dependencies without a proportional increase in quantum circuit depth. To address the limitations of near-term quantum hard- ware, we introduce a novel clipping mechanism that ensures physically realizable entanglement when dilation exceeds the quantum register size, while preserving global quantum correlations. The source code for the proposed quantum model is available at : https://github.com/HoceiniRihab/Quantum-dilated-CNN
Full text 87,331 characters · extracted from preprint-html · click to expand
On the Role of Quantum Entanglement in Capturing Long-Term Dependencies | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article On the Role of Quantum Entanglement in Capturing Long-Term Dependencies Rihab Hoceini This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7095714/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Given the current gap in the literature regarding whether quantum computing can effectively address long-term dependency modeling in sequential learning tasks, this study seeks to explicitly investigate the potential benefits of quantum entanglement in this context. We propose a hybrid Quantum Dilated Con- volutional Neural Network (QDCNN) that synergistically integrates quantum entanglement with classical processing, leveraging the complementary strengths of both computational paradigms. Our architecture employs quantum dilated convolutions with exponentially increasing dilation rates, enabling the efficient capture of long-range temporal dependencies without a proportional increase in quantum circuit depth. To address the limitations of near-term quantum hard- ware, we introduce a novel clipping mechanism that ensures physically realizable entanglement when dilation exceeds the quantum register size, while preserving global quantum correlations. The source code for the proposed quantum model is available at : https://github.com/HoceiniRihab/Quantum-dilated-CNN Quantum entanglement Temporal dependency quantum neural network sequential data Figures Figure 1 1 Introduction Temporal data processing presents formidable challenges in modern computational systems [1, 2]. Time series data and sequences of observations indexed in time order underpin numerous critical applications, including natural language processing [3], financial forecasting [4], biomedical signal analysis [5], and autonomous system control [6]. A fundamental challenge in temporal data processing is the capture and retention of long-term dependencies [2, 7]. Classical computational models often face the difficulty of maintaining contextual information in extended sequences, leading to performance degradation when analyzing complex temporal patterns that span numerous time steps [8]. This ”long-term dependency problem” manifests as an inability to correlate events separated by significant temporal gaps, resulting in suboptimal predictive performance in tasks requiring extensive historical context. To address these limitations, the research community has developed progressively complex architectures, including recurrent neural networks (RNNs) [9], long-short- term memory (LSTM) networks [10], and Temporal Convolutional Networks (TCNs) [11] specifically engineered to capture long-range patterns. However, these models are typically associated with substantial computational complexity [12, 13] and continue to exhibit significant constraints in effectively modeling extensive sequences, particularly when confronted with sparse signals or irregular sampling intervals. Recent investigations have also explored Quantum Machine Learning (QML) as an alternative paradigm for sequence modeling [14] . From a theoretical perspec- tive, QML offers distinct advantages attributable to intrinsically quantum phenomena. However, empirical evaluations of contemporary QML implementations, particularly hybrid classical-quantum architectures, have predominantly yielded performance met- rics comparable to classical counterparts, thus raising substantive questions regarding the practical utility of entanglement in these models. In numerous experimental stud- ies, while entanglement has been incorporated into model architectures, it has not been explicitly leveraged to enhance the model’s capacity to capture long-range depen- dencies. Consequently, the quantum characteristics of such systems frequently remain superficial, contributing minimal advantages beyond classical baseline methodologies. Although a variety of quantum machine learning (QML) architectures harness quantum phenomena to outperform classical models in terms of computational speed and accuracy [15–17], the underlying mechanisms driving these advantages are still not well understood. In particular, the role of quantum entanglement, a key nonclas- sical correlation between quantum subsystems—remains underexplored as a source of performance gains. Recent studies have begun to examine the potential link between entanglement metrics and model performance [18] In contrast, a study has shown that quantum models that incorporate entangle- ment can perform comparably to their classical counterparts [19], suggesting that the observed ”quantumness” may not yield substantial benefits, at least when trained on small datasets. However, the existing literature lacks a focused investigation of whether entanglement can enhance long-term dependency capture in sequential data. To address this gap, we aim to explicitly explore the link between entanglement and temporal information retention. For this purpose, we employ a quantum dilated con- volutional neural network (QDCNN), a model in which entanglement plays a central architectural role. To improve the performance of the model, we integrate classical and quantum layers, using hybrid architectures to enhance both learning flexibility and computational efficiency [20]. Building on this objective, our study offers two main contributions: We investigate how quantum entanglement influences a model’s ability to capture long-range temporal dependencies We propose an enhanced hybrid QDCNN architecture that integrates classical and quantum components to maximize both learning capacity and computational efficiency. We introduce a novel clipping mechanism within the QDCNN design to ensure physical feasibility at high dilation rates. 2 literature review Recent efforts to improve long-term time-series forecasting have highlighted the limi- tations of traditional recurrent models like LSTMs, particularly in handling long-range dependencies. In response to these challenges, the researchers proposed P-sLSTM [21], an enhanced structured LSTM architecture that introduces two critical innovations: patching the input sequence into smaller segments and enforcing channel independence during processing. Through extensive experimentation, P-sLSTM demonstrated con- sistent improvements over baseline LSTM and structured LSTM models across various datasets, validating the importance of localized temporal modeling and independent feature processing for scaling LSTM-based architectures to long-horizon prediction tasks. In an other study where the improvement has done on IoT security , they intro- duced a hybrid LSTM-CNN architecture designed for real-time intrusion detection. By combining LSTM layers to capture temporal dependencies with CNN layers for spatial feature extraction, the model achieved high accuracy (99.87%) and robustness against adversarial attacks. Also in the geospatial deformation analysis, a work introduced a modified Long Short-Term Memory (mLSTM) [22] model tailored for predicting InSAR-derived defor- mation time series in mining regions. Applied to the Khetri Copper Belt in India, the mLSTM outperformed traditional RNN and standard LSTM models, achieving a prediction accuracy of 98.57% and a reduced RMS error of 4.22 mm/year. A recent advancement in neuroengineering has focused on leveraging deep learning techniques for seizure prediction using EEG data[23]. Where the scientists explored the application of Long Short-Term Memory (LSTM) recurrent neural networks to classify unprocessed EEG signals for seizure prediction. They demonstrated that LSTM mod- els could effectively capture temporal dependencies in EEG data, leading to improved accuracy in predicting epileptic seizures. In the improvements of memory for long-term capture in Transformers, several studies have proposed novel architectures to enhance the model’s ability to pro- cess extended sequences. In hydrological forecasting [24] , scientists introduced a transformer-based model that integrates historical streamflow data with climatic vari- ables to enhance long-term streamflow prediction accuracy. Evaluated across five diverse U.S. basins, this transformer architecture consistently outperformed traditional LSTM models. In addressing the challenge of processing long sequences with Transformers [25] introduced TransformerFAM, a novel architecture that incorporates a feedback atten- tion mechanism, enabling the model to attend to its own latent representations, effectively functioning as a form of working memory. Notably, TransformerFAM achieves this without additional parameters, allowing seamless integration with exist- ing pre-trained models. Empirical evaluations demonstrated that TransformerFAM significantly improves performance on long-context tasks across various model sizes. In multi-object tracking (MOT), Gao and Wang [26] introduced MeMOTR, a Transformer-based model enhanced with long-term memory capabilities. By incorpo- rating a customized memory-attention layer, MeMOTR stabilizes and distinguishes object track embeddings over extended sequences, improving target association. Evaluations on DanceTrack demonstrated significant performance gains, surpassing state-of-the-art methods by 7.9% in HOTA and 13.0% in AssA metrics. The model also outperformed other Transformer-based approaches on MOT17 and generalized effectively to datasets like BDD100K and SportsMOT. In the pursuit of enhancing long-term memory capabilities in Recurrent Neural Networks (RNNs), recent studies have introduced innovative architectures to address the limitations of traditional RNNs. One study [27] proposed the Delayed Memory Unit (DMU), which incorporates a delay line structure and delay gates into vanilla RNNs. This design enables the direct distribution of input information to optimal future time instants, enhancing tempo- ral interactions and facilitating temporal credit assignment. The DMU demonstrated superior temporal modeling capabilities across various sequential tasks, such as speech recognition and ECG waveform segmentation, while utilizing fewer parameters than other state-of-the-art gated RNN models. Another approach [28] introduced the Parallel Gated Network (PGN) as a successor to traditional RNNs for long-range time series forecasting. PGN employs a Historical Information Extraction (HIE) layer to directly capture information from previous time steps and utilizes gated mechanisms to fuse it with current inputs. This approach reduces the information propagation path, effectively addressing the limitations of RNNs. The Temporal PGN (TPGN) framework further enhances performance by capturing both long-term periodic patterns and short-term information, achieving state-of-the-art results on multiple benchmark datasets. A futher contribution [29] developed the Light Recurrent Unit (LRU), focusing on interpretability and efficiency in modeling long-range dependencies. The LRU simpli- fies the recurrent structure, reducing complexity while maintaining performance. It introduces a light gating mechanism that balances the retention of long-term infor- mation with computational efficiency, making it suitable for applications requiring interpretable models with lower resource consumption. On the quantum front, recent research has begun addressing the limitations of classical models in long-term sequence modeling by leveraging quantum mechanisms. The study of the QuLTSF model [30] introduces a quantum circuit-based architecture specifically designed to enhance long-term time series forecasting performance. By exploiting quantum entanglement and circuit expressivity, the model shows improved capability in capturing temporal dependencies over extended horizons. Similarly, another work integrates quantum kernel functions into the LSTM framework [31], allowing the model to encode complex temporal features in a high-dimensional quan- tum space—resulting in more efficient learning with fewer parameters. Complementing these efforts, researchers present QSegRNN [32], a hybrid architecture replacing clas- sical RNN units with quantum-enhanced recurrent cells. It demonstrates that even under resource constraints, quantum models can effectively handle sequential fore- casting tasks, suggesting a promising pathway toward scalable quantum memory modeling. 3 Methodology 3.1 Background The Quantum Dilated Convolutional Neural Network (QDCNN) [33] is a quantum architecture that extends the principles of dilated convolutions into the quantum regime using entanglement to enhance the receptive field across temporally distant input points. In QDCNN, input sequences are first encoded into quantum states and then processed through parametrized quantum circuits, where each quantum layer serves as an analog to a convolutional kernel. Dilation is introduced by controlling the entangling connectivity pattern: rather than restricting interactions to adjacent qubits (i.e., local time points), controlled quantum gates are applied between qubits with increasing separation, effectively expanding the temporal field of interaction with- out increasing circuit depth proportionally. The variational parameters of the unitary operations are optimized using classical gradient-based methods, enabling task-specific learning. 3.2 Implementation 3.2.1 Temporal Encoding Strategy To bridge classical time series inputs with quantum computation, we employ a hybrid temporal encoding strategy that combines amplitude-independent rotation- based encoding with position-aware phase modulation . Each input sample is first preprocessed using a linear transformation followed by ReLU activation to reduce the feature dimension to the number of available qubits n qubits . The transformed features are then mapped onto quantum states through two encoding stages: first, each qubit is initialized in superposition via Hadamard gates, and second, a PhaseShift operations of the form PhaseShift( θ t ) and RY( θ t ) rotations are applied, where the angle θ t is a function of the input value at timestep t . Specifically, temporal phase encoding incor- porates the qubit index as a modulation term, enabling time-sensitive transformation of each feature into the quantum Hilbert space. Moreover, we integrate a positional sinusoidal modulation into the rotational and phase layers using sin · π , thereby introducing a global positional context across the sequence. This strategy enables the quantum circuit to retain fine-grained temporal variation while simultaneously encoding broader sequence structure, estab- lishing a strong representational foundation for downstream entanglement-based operations. 3.2.2 Quantum Dilated Convolution Architecture The quantum convolutional layers in our model follow a dilated entanglement pat- tern analogous to classical dilated convolutions, but implemented in the quantum circuit via exponentially increasing qubit interactions. The quantum circuit com- prises n layers = 5, with each layer applying entangling gates between non-adjacent qubits at exponentially spaced intervals governed by the dilation rate d l = 2 2 l , where l ∈ {0 , 1 , ..., n layers − 1}. This dilation pattern ensures that the quantum model cap- tures long-term dependencies without significantly increasing circuit depth, as each layer reaches progressively farther in temporal space. Within each layer, parameterized rotations (RY) and phase shifts (PhaseShift) are applied to every qubit, where the parameters are modulated by their relative positions. These operations are driven by a learnable weight vector w ∈ R 2 ·n qubits · n layers , optimized using gradient descent during training. The final entanglement layer, a linear chain of CNOTs, reinforces global quantum correlations, acting as a final pass to bind all encoded temporal features before measurement. The quantum circuit outputs three expectation values—PauliZ, PauliX, and PauliY—capturing orthogonal projections of the quantum state. These are fed into a deep post-processing network composed of fully connected layers with batch nor- malization and dropout, enabling the model to refine and interpret quantum-encoded temporal patterns for regression or classification tasks. In a Quantum Dilated Convolutional Neural Network (QDCNN), the dilation rate d l dictates the spacing between entangled qubits, effectively expanding the temporal receptive field. When the dilation rate exceeds the number of qubits n qubits , certain entangling operations would nominally target qubits beyond the available index range. To maintain physical feasibility, we constrain the entanglement to valid qubit pairs by clipping the target index, such that the control qubit q i is entangled with q j = min( i + d l , n qubits − 1). This sparse entanglement strategy preserves the intent of long-range temporal interaction by coupling early qubits with the farthest accessible ones. Despite the reduced number of entangling operations in high-dilation layers, this mechanism effec- tively simulates long-term memory, allowing the model to capture extended temporal dependencies with minimal increase in circuit depth or qubit count. As a result, the QDCNN achieves broader temporal coverage in a resource-efficient manner. To clarify how dilation is implemented in our quantum dilated convolution archi- tecture, the table below provides a schematic illustration of entangling qubits using exponentially increasing dilation rates. When the dilation exceeds the number of qubits, the entanglement is clipped to the maximum valid index, preserving the integrity of the quantum circuit. In this architecture, dilation determines the distance between entangled qubit pairs in each layer. For the layerr l , the dilation is set as d l = 2 2 l , which grows exponentially. Table 1 Entanglement patterns across layers Layer Dilation d l Entanglement Pattern (control → target) 0 1 q 0 → q 1 , q 1 → q 2 , q 2 → q 3 , q 3 → q 4 1 4 q 0 → q 4 , q 1 → q 4 (clipped), q 2 → q 4 (clipped) 2 16 q 0 → q 4 (clipped), q 1 → q 4 (clipped) If a qubit q i is meant to entangle with a qubit q j = i + d l , but j ≥ n qubits , we clip the target to q n qubits − 1, the last qubit. In classical dilated convolutions, dilation is achieved by applying convolutional filters over inputs with fixed gaps between elements, allowing the network to expand its receptive field without increasing filter size or stacking deeper layers. In the quantum analogue, this expansion is realized through entanglement between qubits that are increasingly spaced apart in each layer. Instead of sliding filters, dilation manifests as controlled quantum gates—typically entangling gates such as CNOT or CZ—applied between non-adjacent qubits. For example, a dilation rate of d l = 4 implies that qubit q i is entangled with qubit q i +4 , where permitted by the number of available qubits. Unlike classical convolutions, which require explicit memory management across time steps, quantum entanglement enables a more compact representation of dependencies by encoding long-range correlations directly into the quantum state. This intrinsic non-locality makes the quantum dilation mechanism potentially more efficient, as it captures extended temporal features without increasing network depth or the number of parameters, thus reducing overall circuit complexity while maintaining expressive 4 Results and Discussion 4.1 Experiments To validate our hypothesis that quantum entanglement enhances a QML model’s ability to capture long-term temporal dependencies, we empirically evaluated the pro- posed hybrid quantum dilated convolutional neural network (QDCNN) against several state-of-the-art classical architectures, including LSTM, RNN, GRU, TCN[ 11 ], and R-Transformer[ 34 ]. These models were chosen for their proven performance in sequen- tial data modeling. Notably, the Temporal Convolutional Network (TCN) incorporates dilated convolutions an architectural feature it shares with our QDCNN making the comparison particularly relevant, as it evaluates the added value of quantum entangle- ment on top of an already optimized classical baseline. To ensure fairness, all models are trained and tested under identical conditions using two benchmark datasets: one with short-term dependencies and another designed to challenge long-range memory capabilities. Consistent hyperparameters and data preprocessing pipelines are applied across all experiments. 4.1.1 Dataset 1: The Adding Problem The Adding Problem [ 10 ] is a standard synthetic benchmark designed to evaluate a model’s ability to capture long-term dependencies in sequential data. Each input sample is a sequence of length n with dimensionality 2. The first dimension contains real-valued numbers randomly sampled from a uniform distribution over [0, 1], while the second dimension is a binary mask initialized to zeros, with exactly two positions randomly set to one. The task is to predict the sum of the two values in the first dimension whose corresponding positions in the mask (second dimension) are equal to one. 4.1.2 Dataset 2: Nottingham Polyphonic Music The Nottingham dataset [ 35 ] is a widely used benchmark in sequence modeling, con- sisting of approximately 1,200 British and American folk tunes encoded in a polyphonic music format. Each piece is represented as a binary piano-roll matrix, where each row corresponds to a time step and each column indicates whether a particular note is played. This multi-label temporal structure makes the dataset particularly suitable for evaluating a model’s ability to learn complex temporal correlations across multiple concurrent signals. Table 2 Performance comparison of various models on the Adding Problem and Music Data Model Adding Problem Data Music Data LSTM 1 0.164 3.29 GRU 1 5.3e-5 3.46 RNN 1 0.177 4.05 TCN 1 5.8e-5 3.07 R-Transformer 2 - 2.37 Hybird QDCNN 0.174 1.04 1 Example for a first table footnote. This is an example of table footnote. 2 Example for a second table footnote. This is an example of table footnote. 4.2 Discussion The experimental results presented in Table 2 support our hypothesis regarding the role of quantum entanglement in enhancing temporal modeling capabilities. On the Adding Problem dataset—a synthetic task that does not demand significant long- term memory—the performance of our Hybrid Quantum Dilated Convolutional Neural Network (QDCNN) aligns closely with classical models such as LSTM and GRU. This outcome is consistent with existing literature, where quantum models often demonstrate performance parity with classical counterparts on small-scale, low-complexity tasks. However, the distinction becomes pronounced in the more complex Music dataset, where long-range temporal dependencies are essential. In this setting, QDCNN signif- icantly outperforms all classical benchmarks, including the R-Transformer and TCN. This result underscores the potential of quantum entanglement—integrated in the QDCNN architecture—not merely as a novelty, but as a meaningful contributor to enhanced sequence modeling, even under current limitations in quantum resources and circuit depth. Beyond accuracy, the QDCNN model also benefits from reduced parameteriza- tion and increased computational efficiency. With a total number of parameters not exceeding 3000, the QDCNN offers a significantly leaner architecture compared to the best compared classical models, which often require tens or hundreds of thousands of parameters for similar tasks. The design not only accelerates training and inference times but also reduces memory overhead, making it highly suitable for deployment on resource-constrained quantum hardware. Such efficiency opens new pathways for addressing computational challenges in temporal data analysis, especially in domains like weather prediction, where the most efficient models currently achieve reliable forecasts up to only 15 days [ 36 , 37 ]. 5 Conclusion In this work, we introduced a hybrid Quantum Dilated Convolutional Neural Net- work (QDCNN) that systematically investigates the role of quantum entanglement in capturing long-range temporal dependencies in sequential data. By synergistically integrating classical and quantum components, our model capitalizes on the comple- mentary strengths of hybrid architectures, aligning with the practical constraints and opportunities of near-term quantum devices. Our results demonstrate that quantum entanglement significantly enhances the model’s ability to process long-term temporal information, particularly when combined with the computational efficiency of classical processing. Through rigorous evaluation on benchmark tasks, including the Adding Problem and Music datasets, we demonstrated that the proposed QDCNN consistently outperforms its purely classical counterpart. Our architecture not only advances the design of quantum-enhanced temporal models but also underscores the importance of architectural refinements such as the clipping mechanism in optimizing learning performance within hybrid quantum-classical systems. This study paves the way for future research into efficient hybrid quantum-classical models for complex sequential tasks and contributes to the broader effort of integrating quantum computing into practical machine learning workflows. Declarations Author Contribution All aspects of the work were completed solely by Rihab Hoceini Acknowledgement I would like to sincerely thank Dr. Ahmed Bouida for his guidance, support, and willingness to answer my questions throughout this work. His insights and encouragement were greatly appreciated. References Mojtahedi, F.F., Yousefpour, N., Chow, S., Cassidy, M.: Deep learning for time series forecasting: Review and applications in geotechnics and geosciences. Archives of Computational Methods in Engineering, 1–31 (2025) Casolaro, A., Capone, V., Iannuzzo, G., Camastra, F.: Deep learning for time series forecasting: Advances and open problems. Information 14 (11), 598 (2023) Luan, Y., Zhang, H., Zhang, C., Mu, Y., Wang, W.: Stock price prediction with sentiment analysis for chinese market. In: Proceedings of the Joint Work- shop of the 7th Financial Technology and Natural Language Processing, the 5th Knowledge Discovery from Unstructured Data in Financial Services, and the 4th Workshop on Economics and Natural Language Processing@ LREC-COLING 2024, pp. 167–177 (2024) Zhang, Y., Yang, W., Wang, J., Ma, Q., Xiong, J.: Camef: Causal- augmented multi-modality event-driven financial forecasting by integrating time series patterns and salient macroeconomic announcements. arXiv preprint arXiv:2502.04592 (2025) Li, X., Sakevych, M., Atkinson, G., Metsis, V.: Biodiffusion: A versatile diffusion model for biomedical signal synthesis. Bioengineering 11 (4), 299 (2024) Wang, C.-C.: T 2-lstm-based ai system for early detection of motor failure in chemical plants. Mathematics (2227-7390) 12 (17) (2024) Su, L., Zuo, X., Li, R., Wang, X., Zhao, H., Huang, B.: A systematic review for transformer-based long-term series forecasting. Artificial Intelligence Review 58 (3), 80 (2025) Johnston, L., Patel, V., Cui, Y., Balaprakash, P.: Revisiting the problem of learn- ing long-term dependencies in recurrent neural networks. Neural Networks 183 , 106887 (2025) Elman, J.L.: Finding structure in time. Cognitive science 14 (2), 179–211 (1990) Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural computation 9 (8), 1735–1780 (1997) Bai, S., Kolter, J.Z., Koltun, V.: An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271 (2018) Waqas, M., Humphries, U.W.: A critical review of rnn and lstm variants in hydrological time series predictions. MethodsX, 102946 (2024) Haider, A., Lee, G., Jafri, T.H., Yoon, P., Piao, J., Jhang, K.: Enhancing accuracy of groundwater level forecasting with minimal computational complexity using temporal convolutional network. Water 15 (23), 4041 (2023) Peral-Garc´ıa, D., Cruz-Benito, J., Garc´ıa-Pen˜alvo, F.J.: Systematic literature review: Quantum machine learning and its applications. Computer Science Review 51 , 100619 (2024) Senokosov, A., Sedykh, A., Sagingalieva, A., Kyriacou, B., Melnikov, A.: Quan- tum machine learning for image classification. Machine Learning: Science and Technology 5 (1), 015040 (2024) Corli, S., Moro, L., Dragoni, D., Dispenza, M., Prati, E.: Quantum machine learn- ing algorithms for anomaly detection: A review. Future Generation Computer Systems, 107632 (2024) Macaluso, A.: Quantum supervised learning. KI-Ku¨nstliche Intelligenz, 1–15 (2024) Nguyen, N., Chen, K.-C.: Quantum embedding search for quantum machine learning. IEEE Access 10 , 41444–41456 (2022) Bowles, J., Ahmed, S., Schuld, M.: Better than classical? the subtle art of bench- marking quantum machine learning models. arXiv preprint arXiv:2403.07059 (2024) Broughton, M., Verdon, G., McCourt, T., Martinez, A.J., Yoo, J.H., Isakov, S.V., Massey, P., Halavati, R., Niu, M.Y., Zlokapa, A., et al.: Tensorflow quan- tum: A software framework for quantum machine learning. arXiv preprint arXiv:2003.02989 (2020) Kong, Y., Wang, Z., Nie, Y., Zhou, T., Zohren, S., Liang, Y., Sun, P., Wen, Q.: Unlocking the power of lstm for long term time series forecasting. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, pp. 11968–11976 (2025) Soni, R., Alam, M.S., Vishwakarma, G.K.: Prediction of insar deformation time- series using improved lstm deep learning model. Scientific Reports 15 (1), 5333 (2025) Chambers, J.D., Cook, M.J., Burkitt, A.N., Grayden, D.B.: Using long short-term memory (lstm) recurrent neural networks to classify unprocessed eeg for seizure prediction. Frontiers in Neuroscience 18 , 1472747 (2024) Singh, R., Lu, D., Tayal, K.: Temporal sequence transformer to advance long-term streamflow prediction. In: Proceedings of the NeurIPS 2024 Workshop on Tackling Climate Change with Machine Learning, Vancouver, British Columbia, Canada (2024). Part of the Conference on Neural Information Processing Systems, Dec. 10–15, 2024 Hwang, D., Wang, W., Huo, Z., Sim, K.C., Mengibar, P.M.: Transformerfam: Feedback attention is working memory. arXiv preprint arXiv:2404.09173 (2024) Gao, R., Wang, L.: Memotr: Long-term memory-augmented transformer for multi-object tracking. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9901–9910 (2023) Sun, P., Wu, J., Zhang, M., Devos, P., Botteldooren, D.: Delayed memory unit: modeling temporal dependency through delay gate. IEEE Transactions on Neural Networks and Learning Systems (2024) Jia, Y., Lin, Y., Yu, J., Wang, S., Liu, T., Wan, H.: Pgn: The rnn’s new successor is effective for long-range time series forecasting. Advances in Neural Information Processing Systems 37 , 84139–84168 (2024) Ye, H., Zhang, Y., Liu, H., Li, X., Chang, J., Zheng, H.: Light recurrent unit: Towards an interpretable recurrent neural network for modeling long-range dependency. Electronics 13 (16), 3204 (2024) Chittoor, H.H.S., Griffin, P.R., Neufeld, A., Thompson, J., Gu, M.: Qultsf: Long- term time series forecasting with quantum machine learning. arXiv preprint arXiv:2412.13769 (2024) Hsu, Y.-C., Chen, N.-Y., Li, T.-Y., Chen, K.-C., et al.: Quantum kernel- based long short-term memory for climate time-series forecasting. arXiv preprint arXiv:2412.08851 (2024) Moon, K.-H., Jeong, S.-G., Hwang, W.-J.: Qsegrnn: quantum segment recurrent neural network for time series forecasting. EPJ Quantum Technology 12 (1), 32 (2025) Chen, Y.: Quantum dilated convolutional neural networks. IEEE Access 10 , 20240–20246 (2022) Wang, Z., Ma, Y., Liu, Z., Tang, J.: R-transformer: Recurrent neural network enhanced transformer. arXiv preprint arXiv:1907.05572 (2019) Boulanger-Lewandowski, N., Bengio, Y., Vincent, P.: Modeling temporal depen- dencies in high-dimensional sequences: Application to polyphonic music genera- tion and transcription. arXiv preprint arXiv:1206.6392 (2012) Chen, L., Zhong, X., Zhang, F., Cheng, Y., Xu, Y., Qi, Y., Li, H.: Fuxi: A cas- cade machine learning forecasting system for 15-day global weather forecast. npj climate and atmospheric science 6 (1), 190 (2023) Chen, K., Han, T., Gong, J., Bai, L., Ling, F., Luo, J.-J., Chen, X., Ma, L., Zhang, T., Su, R., et al.: Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948 (2023) Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7095714","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":514954793,"identity":"ff1d3ff5-1304-4695-bb12-3cfc6ad384ff","order_by":0,"name":"Rihab Hoceini","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9UlEQVRIiWNgGAWjYBACAxCRwCDBwAZEDB8MbIBcxsYDRGthnFGRBtLSQFgLGAC1MPOcOQxm49Vizt788MHDHRb5fNLNzx7wtp23W9t+GGhLjU00Li2WPceMDRLPSFi2yRwzN5Bsu5287UwiUMuxtNwGXA67kcMmkdgmYcAmkWAmYQjUYnYAqIWx4TA+Lew/IFrSvwH1nks2O/+QoBY2BoiWHDOJA2cO2JndIGTLmWPGUIfllBs2VCQnmN0A2pKAzy/Hmx9+/NlWZyA/I33b4z8GdvZm59MfPvhQY4NTCwZIBKtMIFY5CNiTongUjIJRMApGBgAAD1xh0pE27e0AAAAASUVORK5CYII=","orcid":"","institution":"Independent researcher","correspondingAuthor":true,"prefix":"","firstName":"Rihab","middleName":"","lastName":"Hoceini","suffix":""}],"badges":[],"createdAt":"2025-07-10 19:08:12","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7095714/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7095714/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":91422366,"identity":"2a49963a-e7dc-4e09-8bc7-2cd151368f8d","added_by":"auto","created_at":"2025-09-16 10:24:21","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":125559,"visible":true,"origin":"","legend":"\u003cp\u003ehybird QDCNN Model Architecture\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7095714/v1/d0797a8707bbb25a4713c9c3.png"},{"id":99316294,"identity":"fef2ea4c-439a-48f3-881d-8f08e0ecaab9","added_by":"auto","created_at":"2025-12-31 16:28:07","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":685430,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7095714/v1/99770e80-eab0-4583-9ef6-4865842f8ded.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"On the Role of Quantum Entanglement in Capturing Long-Term Dependencies","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eTemporal data processing presents formidable challenges in modern computational systems [1, 2]. Time series data and sequences of observations indexed in time order underpin numerous critical applications, including natural language processing [3], financial forecasting [4], biomedical signal analysis [5], and autonomous system control [6].\u003c/p\u003e\n\u003cp\u003eA fundamental challenge in temporal data processing is the capture and retention of long-term dependencies [2, 7]. Classical computational models often face the difficulty of maintaining contextual information in extended sequences, leading to performance degradation when analyzing complex temporal patterns that span numerous time steps [8]. This \u0026rdquo;long-term dependency problem\u0026rdquo; manifests as an inability to correlate events separated by significant temporal gaps, resulting in suboptimal predictive performance in tasks requiring extensive historical context.\u003c/p\u003e\n\u003cp\u003eTo address these limitations, the research community has developed progressively complex architectures, including recurrent neural networks (RNNs) [9], long-short- term memory (LSTM) networks [10], and Temporal Convolutional Networks (TCNs)\u003c/p\u003e\n\u003cp\u003e[11] specifically engineered to capture long-range patterns. However, these models are typically associated with substantial computational complexity [12, 13] and continue to exhibit significant constraints in effectively modeling extensive sequences, particularly when confronted with sparse signals or irregular sampling intervals.\u003c/p\u003e\n\u003cp\u003eRecent investigations have also explored Quantum Machine Learning (QML) as an alternative paradigm for sequence modeling [14] . From a theoretical perspec- tive, QML offers distinct advantages attributable to intrinsically quantum phenomena. However, empirical evaluations of contemporary QML implementations, particularly hybrid classical-quantum architectures, have predominantly yielded performance met- rics comparable to classical counterparts, thus raising substantive questions regarding the \u003cem\u003epractical utility \u003c/em\u003eof entanglement in these models. In numerous experimental stud- ies, while entanglement has been incorporated into model architectures, it has not been explicitly leveraged to enhance the model\u0026rsquo;s capacity to capture long-range depen- dencies. Consequently, the quantum characteristics of such systems frequently remain superficial, contributing minimal advantages beyond classical baseline methodologies. Although a variety of quantum machine learning (QML) architectures harness quantum phenomena to outperform classical models in terms of computational speed and accuracy [15\u0026ndash;17], the underlying mechanisms driving these advantages are still not well understood. In particular, the role of quantum entanglement, a key nonclas- sical correlation between quantum subsystems\u0026mdash;remains underexplored as a source of performance gains. Recent studies have begun to examine the potential link between entanglement metrics and model performance [18]\u003c/p\u003e\n\u003cp\u003eIn contrast, a study has shown that quantum models that incorporate entangle- ment can perform comparably to their classical counterparts [19], suggesting that the observed \u0026rdquo;quantumness\u0026rdquo; may not yield substantial benefits, at least when trained on small datasets. However, the existing literature lacks a focused investigation of whether entanglement can enhance long-term dependency capture in sequential data. To address this gap, we aim to explicitly explore the link between entanglement and temporal information retention. For this purpose, we employ a quantum dilated con- volutional neural network (QDCNN), a model in which entanglement plays a central architectural role. To improve the performance of the model, we integrate classical and quantum layers, using hybrid architectures to enhance both learning flexibility and computational efficiency [20].\u003c/p\u003e\n\u003cp\u003eBuilding on this objective, our study offers two main contributions:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eWe investigate how quantum entanglement influences a model\u0026rsquo;s ability to capture long-range temporal dependencies\u003c/li\u003e\n\u003cli\u003eWe propose an enhanced hybrid QDCNN architecture that integrates classical and quantum components to maximize both learning capacity and computational efficiency.\u003c/li\u003e\n\u003cli\u003eWe introduce a novel clipping mechanism within the QDCNN design to ensure physical feasibility at high dilation rates.\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"2 literature review","content":"\u003cp\u003eRecent efforts to improve long-term time-series forecasting have highlighted the limi- tations of traditional recurrent models like LSTMs, particularly in handling long-range dependencies. In response to these challenges, the researchers proposed P-sLSTM [21], an enhanced structured LSTM architecture that introduces two critical innovations: patching the input sequence into smaller segments and enforcing channel independence during processing. Through extensive experimentation, P-sLSTM demonstrated con- sistent improvements over baseline LSTM and structured LSTM models across various datasets, validating the importance of localized temporal modeling and independent feature processing for scaling LSTM-based architectures to long-horizon prediction tasks.\u003c/p\u003e\n\u003cp\u003eIn an other study where the improvement has done on IoT security , they intro- duced a hybrid LSTM-CNN architecture designed for real-time intrusion detection. By combining LSTM layers to capture temporal dependencies with CNN layers for spatial feature extraction, the model achieved high accuracy (99.87%) and robustness against adversarial attacks.\u003c/p\u003e\n\u003cp\u003eAlso in the geospatial deformation analysis, a work introduced a modified Long Short-Term Memory (mLSTM) [22] model tailored for predicting InSAR-derived defor- mation time series in mining regions. Applied to the Khetri Copper Belt in India, the mLSTM outperformed traditional RNN and standard LSTM models, achieving a prediction accuracy of 98.57% and a reduced RMS error of 4.22 mm/year.\u003c/p\u003e\n\u003cp\u003eA recent advancement in neuroengineering has focused on leveraging deep learning techniques for seizure prediction using EEG data[23]. Where the scientists explored the application of Long Short-Term Memory (LSTM) recurrent neural networks to classify unprocessed EEG signals for seizure prediction. They demonstrated that LSTM mod- els could effectively capture temporal dependencies in EEG data, leading to improved accuracy in predicting epileptic seizures.\u003c/p\u003e\n\u003cp\u003eIn the improvements of memory for long-term capture in Transformers, several studies have proposed novel architectures to enhance the model’s ability to pro- cess extended sequences. In hydrological forecasting [24] , scientists introduced a transformer-based model that integrates historical streamflow data with climatic vari- ables to enhance long-term streamflow prediction accuracy. Evaluated across five diverse U.S. basins, this transformer architecture consistently outperformed traditional LSTM models.\u003c/p\u003e\n\u003cp\u003eIn addressing the challenge of processing long sequences with Transformers [25] introduced TransformerFAM, a novel architecture that incorporates a feedback atten- tion mechanism, enabling the model to attend to its own latent representations, effectively functioning as a form of working memory. Notably, TransformerFAM achieves this without additional parameters, allowing seamless integration with exist- ing pre-trained models. Empirical evaluations demonstrated that TransformerFAM significantly improves performance on long-context tasks across various model sizes.\u003c/p\u003e\n\u003cp\u003eIn multi-object tracking (MOT), Gao and Wang [26] introduced MeMOTR, a Transformer-based model enhanced with long-term memory capabilities. By incorpo- rating a customized memory-attention layer, MeMOTR stabilizes and distinguishes object track embeddings over extended sequences, improving target association. Evaluations on DanceTrack demonstrated significant performance gains, surpassing state-of-the-art methods by 7.9% in HOTA and 13.0% in AssA metrics. The model also outperformed other Transformer-based approaches on MOT17 and generalized effectively to datasets like BDD100K and SportsMOT.\u003c/p\u003e\n\u003cp\u003eIn the pursuit of enhancing long-term memory capabilities in Recurrent Neural Networks (RNNs), recent studies have introduced innovative architectures to address the limitations of traditional RNNs.\u003c/p\u003e\n\u003cp\u003eOne study [27] proposed the Delayed Memory Unit (DMU), which incorporates a delay line structure and delay gates into vanilla RNNs. This design enables the direct distribution of input information to optimal future time instants, enhancing tempo- ral interactions and facilitating temporal credit assignment. The DMU demonstrated superior temporal modeling capabilities across various sequential tasks, such as speech recognition and ECG waveform segmentation, while utilizing fewer parameters than other state-of-the-art gated RNN models.\u003c/p\u003e\n\u003cp\u003eAnother approach [28] introduced the Parallel Gated Network (PGN) as a successor to traditional RNNs for long-range time series forecasting. PGN employs a Historical Information Extraction (HIE) layer to directly capture information from previous time steps and utilizes gated mechanisms to fuse it with current inputs. This approach reduces the information propagation path, effectively addressing the limitations of RNNs. The Temporal PGN (TPGN) framework further enhances performance by capturing both long-term periodic patterns and short-term information, achieving state-of-the-art results on multiple benchmark datasets.\u003c/p\u003e\n\u003cp\u003eA futher contribution [29] developed the Light Recurrent Unit (LRU), focusing on interpretability and efficiency in modeling long-range dependencies. The LRU simpli- fies the recurrent structure, reducing complexity while maintaining performance. It introduces a light gating mechanism that balances the retention of long-term infor- mation with computational efficiency, making it suitable for applications requiring interpretable models with lower resource consumption.\u003c/p\u003e\n\u003cp\u003eOn the quantum front, recent research has begun addressing the limitations of classical models in long-term sequence modeling by leveraging quantum mechanisms. The study of the QuLTSF model [30] introduces a quantum circuit-based architecture specifically designed to enhance long-term time series forecasting performance. By exploiting quantum entanglement and circuit expressivity, the model shows improved capability in capturing temporal dependencies over extended horizons. Similarly,\u003c/p\u003e\n\u003cp\u003eanother work integrates quantum kernel functions into the LSTM framework [31], allowing the model to encode complex temporal features in a high-dimensional quan- tum space—resulting in more efficient learning with fewer parameters. Complementing these efforts, researchers present QSegRNN [32], a hybrid architecture replacing clas- sical RNN units with quantum-enhanced recurrent cells. It demonstrates that even under resource constraints, quantum models can effectively handle sequential fore- casting tasks, suggesting a promising pathway toward scalable quantum memory modeling.\u003c/p\u003e"},{"header":"3\tMethodology","content":"\u003ch2\u003e3.1 Background\u003c/h2\u003e\n\u003cp\u003eThe Quantum Dilated Convolutional Neural Network (QDCNN) [33] is a quantum architecture that extends the principles of dilated convolutions into the quantum regime using entanglement to enhance the receptive field across temporally distant input points. In QDCNN, input sequences are first encoded into quantum states and then processed through parametrized quantum circuits, where each quantum layer serves as an analog to a convolutional kernel. Dilation is introduced by controlling the entangling connectivity pattern: rather than restricting interactions to adjacent qubits (i.e., local time points), controlled quantum gates are applied between qubits with increasing separation, effectively expanding the temporal field of interaction with- out increasing circuit depth proportionally. The variational parameters of the unitary operations are optimized using classical gradient-based methods, enabling task-specific learning.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e3.2 Implementation\u003c/h2\u003e\n\u003ch3\u003e3.2.1 Temporal Encoding Strategy\u003c/h3\u003e\n\u003cp\u003eTo bridge classical time series inputs with quantum computation, we employ a hybrid temporal encoding strategy that combines \u003cem\u003eamplitude-independent rotation- based encoding\u0026nbsp;\u003c/em\u003ewith \u003cem\u003eposition-aware phase modulation\u003c/em\u003e. Each input sample is first preprocessed using a linear transformation followed by ReLU activation to reduce the feature dimension to the number of available qubits \u003cem\u003en\u003c/em\u003e\u003csub\u003equbits\u003c/sub\u003e. The transformed features are then mapped onto quantum states through two encoding stages: first, each qubit is initialized in superposition via Hadamard gates, and second, a PhaseShift operations of the form PhaseShift(\u003cem\u003e\u0026theta;\u003csub\u003et\u003c/sub\u003e\u003c/em\u003e) and RY(\u003cem\u003e\u0026theta;\u003csub\u003et\u003c/sub\u003e\u003c/em\u003e) rotations are applied, where the angle \u003cem\u003e\u0026theta;\u003csub\u003et\u003c/sub\u003e\u0026nbsp;\u003c/em\u003eis a function of the input value at timestep \u003cem\u003et\u003c/em\u003e. Specifically, temporal phase encoding incor- porates the qubit index as a modulation term, enabling time-sensitive transformation of each feature into the quantum Hilbert space.\u003c/p\u003e\n\u003cp\u003eMoreover, we integrate a \u003cem\u003epositional sinusoidal modulation\u0026nbsp;\u003c/em\u003einto the rotational and phase layers using sin \u003cu\u003e\u0026nbsp;\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAEIAAAAfCAYAAABTRBvBAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAAFiUAABYlAUlSJPAAAAJRSURBVGhD7Zg9ruowEIWP3zqo4qyBAgUKGmcJgR0kC2AR2QASFRVmC7Q4omANOAViH/OK51jJABduHn9XN59kKZmxLTgaTzwjiIjQgT/c8FvphHB0Qjg6IRydEI5OCMfLhMiyDEIIlGXJXV8SxzGEEMiyjLseC70RYwxd+wlaawJAxhiSUlKapnzKQ3lZRAghsF6vufkqSZKAiBBFEXc9hZcIURQFAKDf73PXx/ASIXa7HZRSCIIARVFACAEhBIbDIeCipRrv4iVCLBYLjMdjAEAURSAiEBGMMcC/JOHHu3i6EGVZwlqLwWDAXZ8Fz56Ppsr+l7jnq8GHlJJPfQiC3hmPH8TTj8ZPoRPC0Qnh6IRwdEI4OiEc/yVEGIbIsgx5nvsrchiGfNqPoLUQ1Y1xPp8D7ppsjIG19ltV5qfQ+kL1zgKpDbf+ZuuISNMUYAWT1vrM1nZIKaG1PrPzobWGlPLMzsctWgtRliWUUg3bdrs9s/H8Ecex70/w1l0Yho1jtVwuL5boWZb53DSZTGCtbbTz6mvyPPfrvoQXH/cCgLTWDRtvqfGiylrr22/k9rDWer+U0u/J90rTlJRS/rnyaa0bhZjW2s/7Dq0i4lLHqUqeo9HI206nUyNCgiCAlNK/36K+13Q6xeFwaPgvkSQJxuPxWRTdopUQVXMlCAJvC4IARIQkSbyt1+ths9n490qsOvv93j9z3/F49M+r1eruT/NsNvP5I45j7r4MD5FHk6Zpo5cgpfRHg/cclFKNo6GUavjre9aPTeWv7PU1/Phe4y8Za6Wr4lztdgAAAABJRU5ErkJggg==\" width=\"66\" height=\"31\"\u003e\u003c/u\u003e\u0026middot; \u003cem\u003e\u0026pi;\u0026nbsp;\u003c/em\u003e, thereby introducing a global positional context across the sequence. This strategy enables the quantum circuit to retain fine-grained\u003c/p\u003e\n\u003cp\u003etemporal variation while simultaneously encoding broader sequence structure, estab- lishing a strong representational foundation for downstream entanglement-based operations.\u003c/p\u003e\n\u003ch3\u003e3.2.2 Quantum Dilated Convolution Architecture\u003c/h3\u003e\n\u003cp\u003eThe quantum convolutional layers in our model follow a \u003cem\u003edilated entanglement pat- tern\u0026nbsp;\u003c/em\u003eanalogous to classical dilated convolutions, but implemented in the quantum circuit via exponentially increasing qubit interactions. The quantum circuit com- prises \u003cem\u003en\u003c/em\u003e\u003csub\u003elayers\u003c/sub\u003e = 5, with each layer applying entangling gates between non-adjacent qubits at exponentially spaced intervals governed by the dilation rate \u003cem\u003ed\u003csub\u003el\u003c/sub\u003e\u0026nbsp;\u003c/em\u003e= 2\u003csup\u003e2\u003cem\u003el\u003c/em\u003e\u003c/sup\u003e, where \u003cem\u003el\u0026nbsp;\u003c/em\u003e\u0026isin; {0\u003cem\u003e,\u0026nbsp;\u003c/em\u003e1\u003cem\u003e, ..., n\u003c/em\u003e\u003csub\u003elayers\u003c/sub\u003e \u0026minus; 1}. This dilation pattern ensures that the quantum model cap- tures \u003cem\u003elong-term dependencies\u0026nbsp;\u003c/em\u003ewithout significantly increasing circuit depth, as each layer reaches progressively farther in temporal space.\u003c/p\u003e\n\u003cp\u003eWithin each layer, \u003cem\u003eparameterized rotations\u0026nbsp;\u003c/em\u003e(RY) and \u003cem\u003ephase shifts\u0026nbsp;\u003c/em\u003e(PhaseShift) are applied to every qubit, where the parameters are modulated by their relative positions. These operations are driven by a learnable weight vector \u003cem\u003ew\u0026nbsp;\u003c/em\u003e\u0026isin; R\u003csup\u003e2\u003c/sup\u003e\u003cem\u003e\u003csup\u003e\u0026middot;n\u003c/sup\u003e\u003c/em\u003equbits\u003cem\u003e\u0026middot;\u003c/em\u003e\u003cem\u003en\u003c/em\u003elayers , optimized using gradient descent during training. The final entanglement layer, a linear chain of CNOTs, reinforces global quantum correlations, acting as a final pass to bind all encoded temporal features before measurement.\u003c/p\u003e\n\u003cp\u003eThe quantum circuit outputs three expectation values\u0026mdash;PauliZ, PauliX, and PauliY\u0026mdash;capturing orthogonal projections of the quantum state. These are fed into a \u003cem\u003edeep post-processing network\u0026nbsp;\u003c/em\u003ecomposed of fully connected layers with batch nor- malization and dropout, enabling the model to refine and interpret quantum-encoded temporal patterns for regression or classification tasks.\u003c/p\u003e\n\u003cp\u003eIn a Quantum Dilated Convolutional Neural Network (QDCNN), the dilation rate \u003cem\u003ed\u003csub\u003el\u003c/sub\u003e\u0026nbsp;\u003c/em\u003edictates the spacing between entangled qubits, effectively expanding the temporal receptive field. When the dilation rate exceeds the number of qubits \u003cem\u003en\u003c/em\u003e\u003csub\u003equbits\u003c/sub\u003e, certain entangling operations would nominally target qubits beyond the available index range. To maintain physical feasibility, we constrain the entanglement to valid qubit pairs by clipping the target index, such that the control qubit \u003cem\u003eq\u003csub\u003ei\u003c/sub\u003e\u0026nbsp;\u003c/em\u003eis entangled with \u003cem\u003eq\u003csub\u003ej\u003c/sub\u003e\u0026nbsp;\u003c/em\u003e= min(\u003cem\u003ei\u0026nbsp;\u003c/em\u003e+ \u003cem\u003ed\u003csub\u003el\u003c/sub\u003e, n\u003c/em\u003e\u003csub\u003equbits\u003c/sub\u003e \u0026minus; 1).\u003c/p\u003e\n\u003cp\u003eThis sparse entanglement strategy preserves the intent of long-range temporal interaction by coupling early qubits with the farthest accessible ones. Despite the reduced number of entangling operations in high-dilation layers, this mechanism effec- tively simulates long-term memory, allowing the model to capture extended temporal dependencies with minimal increase in circuit depth or qubit count. As a result, the QDCNN achieves broader temporal coverage in a resource-efficient manner.\u003c/p\u003e\n\u003cp\u003eTo clarify how dilation is implemented in our quantum dilated convolution archi- tecture, the table below provides a schematic illustration of entangling qubits using exponentially increasing dilation rates. When the dilation exceeds the number of qubits, the entanglement is clipped to the maximum valid index, preserving the integrity of the quantum circuit.\u003c/p\u003e\n\u003cp\u003eIn this architecture, dilation determines the distance between entangled qubit pairs in each layer. For the layerr \u003cem\u003el\u003c/em\u003e, the dilation is set as \u003cem\u003ed\u003csub\u003el\u003c/sub\u003e\u003c/em\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e= 2\u003csup\u003e2\u003cem\u003el\u003c/em\u003e\u003c/sup\u003e, which grows exponentially.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1\u0026nbsp;\u003c/strong\u003eEntanglement patterns across layers\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLayer\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDilation\u0026nbsp;\u003c/strong\u003e\u003cem\u003ed\u003csub\u003el\u003c/sub\u003e\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 252px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEntanglement Pattern (control\u0026nbsp;\u003c/strong\u003e\u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cstrong\u003etarget)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 252px;\"\u003e\n \u003cp\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e0\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e1\u003c/sub\u003e, \u003cem\u003eq\u003c/em\u003e\u003csub\u003e1\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e2\u003c/sub\u003e, \u003cem\u003eq\u003c/em\u003e\u003csub\u003e2\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e3\u003c/sub\u003e, \u003cem\u003eq\u003c/em\u003e\u003csub\u003e3\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e4\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 252px;\"\u003e\n \u003cp\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e0\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e4\u003c/sub\u003e, \u003cem\u003eq\u003c/em\u003e\u003csub\u003e1\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e4\u003c/sub\u003e (clipped), \u003cem\u003eq\u003c/em\u003e\u003csub\u003e2\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e4\u003c/sub\u003e (clipped)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e16\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 252px;\"\u003e\n \u003cp\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e0\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e4\u003c/sub\u003e (clipped), \u003cem\u003eq\u003c/em\u003e\u003csub\u003e1\u003c/sub\u003e \u003cem\u003e\u0026rarr;\u0026nbsp;\u003c/em\u003e\u003cem\u003eq\u003c/em\u003e\u003csub\u003e4\u003c/sub\u003e (clipped)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eIf a qubit \u003cem\u003eq\u003csub\u003ei\u003c/sub\u003e\u0026nbsp;\u003c/em\u003eis meant to entangle with a qubit \u003cem\u003eq\u003csub\u003ej\u003c/sub\u003e\u0026nbsp;\u003c/em\u003e= \u003cem\u003ei\u0026nbsp;\u003c/em\u003e+ \u003cem\u003ed\u003csub\u003el\u003c/sub\u003e\u003c/em\u003e, but \u003cem\u003ej\u0026nbsp;\u003c/em\u003e\u0026ge; \u003cem\u003en\u003c/em\u003e\u003csub\u003equbits\u003c/sub\u003e, we clip the target to \u003cem\u003eq\u003csub\u003en\u003c/sub\u003e\u003c/em\u003equbits\u003cem\u003e\u0026minus;\u003c/em\u003e1, the last qubit.\u003c/p\u003e\n\u003cp\u003eIn classical dilated convolutions, dilation is achieved by applying convolutional\u003c/p\u003e\n\u003cp\u003efilters over inputs with fixed gaps between elements, allowing the network to expand its receptive field without increasing filter size or stacking deeper layers. In the quantum analogue, this expansion is realized through \u003cem\u003eentanglement\u0026nbsp;\u003c/em\u003ebetween qubits that are increasingly spaced apart in each layer. Instead of sliding filters, dilation manifests as controlled quantum gates\u0026mdash;typically entangling gates such as CNOT or CZ\u0026mdash;applied between non-adjacent qubits. For example, a dilation rate of \u003cem\u003ed\u003csub\u003el\u003c/sub\u003e\u0026nbsp;\u003c/em\u003e= 4 implies that qubit \u003cem\u003eq\u003csub\u003ei\u003c/sub\u003e\u0026nbsp;\u003c/em\u003eis entangled with qubit \u003cem\u003eq\u003csub\u003ei\u003c/sub\u003e\u003c/em\u003e\u003csub\u003e+4\u003c/sub\u003e, where permitted by the number of available qubits. Unlike classical convolutions, which require explicit memory management across time steps, quantum entanglement enables a more \u003cem\u003ecompact representation\u0026nbsp;\u003c/em\u003eof dependencies by encoding long-range correlations directly into the quantum state. This intrinsic non-locality makes the quantum dilation mechanism potentially more efficient, as it captures extended temporal features without increasing network depth or the number of parameters, thus reducing overall circuit complexity while maintaining expressive\u003c/p\u003e"},{"header":"4 Results and Discussion","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e4.1 Experiments\u003c/h2\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eTo validate our hypothesis that quantum entanglement enhances a QML model\u0026rsquo;s ability to capture long-term temporal dependencies, we empirically evaluated the pro- posed hybrid quantum dilated convolutional neural network (QDCNN) against several state-of-the-art classical architectures, including LSTM, RNN, GRU, TCN[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], and R-Transformer[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. These models were chosen for their proven performance in sequen- tial data modeling. Notably, the Temporal Convolutional Network (TCN) incorporates dilated convolutions an architectural feature it shares with our QDCNN making the comparison particularly relevant, as it evaluates the added value of quantum entangle- ment on top of an already optimized classical baseline. To ensure fairness, all models are trained and tested under identical conditions using two benchmark datasets: one with short-term dependencies and another designed to challenge long-range memory capabilities. Consistent hyperparameters and data preprocessing pipelines are applied across all experiments.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003cdiv id=\"Sec7\" class=\"Section3\"\u003e\u003ch2\u003e4.1.1 Dataset 1: The Adding Problem\u003c/h2\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eThe Adding Problem [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] is a standard synthetic benchmark designed to evaluate a model\u0026rsquo;s ability to capture long-term dependencies in sequential data. Each input sample is a sequence of length \u003cem\u003en\u003c/em\u003e with dimensionality 2. The first dimension contains real-valued numbers randomly sampled from a uniform distribution over [0, 1], while the second dimension is a binary mask initialized to zeros, with exactly two positions randomly set to one. The task is to predict the sum of the two values in the first dimension whose corresponding positions in the mask (second dimension) are equal to one.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section3\"\u003e\u003ch2\u003e4.1.2 Dataset 2: Nottingham Polyphonic Music\u003c/h2\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eThe Nottingham dataset [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] is a widely used benchmark in sequence modeling, con- sisting of approximately 1,200 British and American folk tunes encoded in a polyphonic music format. Each piece is represented as a binary piano-roll matrix, where each row corresponds to a time step and each column indicates whether a particular note is played. This multi-label temporal structure makes the dataset particularly suitable for evaluating a model\u0026rsquo;s ability to learn complex temporal correlations across multiple concurrent signals.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003ePerformance comparison of various models on the Adding Problem and Music Data\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAdding Problem Data\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMusic Data\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLSTM\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.164\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.29\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGRU\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e5.3e-5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.46\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRNN\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.177\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4.05\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTCN\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003e5.8e-5\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.07\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eR-Transformer\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.37\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHybird QDCNN\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.174\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e1.04\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e\u003csup\u003e1\u003c/sup\u003eExample for a first table footnote. This is an example of table footnote.\u003c/p\u003e\u003cp\u003e\u003csup\u003e2\u003c/sup\u003eExample for a second table footnote. This is an example of table footnote.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e4.2 Discussion\u003c/h2\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eThe experimental results presented in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e support our hypothesis regarding the role of quantum entanglement in enhancing temporal modeling capabilities. On the Adding Problem dataset\u0026mdash;a synthetic task that does not demand significant long- term memory\u0026mdash;the performance of our Hybrid Quantum Dilated Convolutional Neural Network (QDCNN) aligns closely with classical models such as LSTM and GRU.\u003c/p\u003e\u003cp\u003eThis outcome is consistent with existing literature, where quantum models often demonstrate performance parity with classical counterparts on small-scale, low-complexity tasks.\u003c/p\u003e\u003cp\u003eHowever, the distinction becomes pronounced in the more complex Music dataset, where long-range temporal dependencies are essential. In this setting, QDCNN signif- icantly outperforms all classical benchmarks, including the R-Transformer and TCN. This result underscores the potential of quantum entanglement\u0026mdash;integrated in the QDCNN architecture\u0026mdash;not merely as a novelty, but as a meaningful contributor to enhanced sequence modeling, even under current limitations in quantum resources and circuit depth.\u003c/p\u003e\u003cp\u003eBeyond accuracy, the QDCNN model also benefits from reduced parameteriza- tion and increased computational efficiency. With a total number of parameters not exceeding 3000, the QDCNN offers a significantly leaner architecture compared to the best compared classical models, which often require tens or hundreds of thousands of parameters for similar tasks. The design not only accelerates training and inference times but also reduces memory overhead, making it highly suitable for deployment on resource-constrained quantum hardware. Such efficiency opens new pathways for addressing computational challenges in temporal data analysis, especially in domains like weather prediction, where the most efficient models currently achieve reliable forecasts up to only 15 days [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"5 Conclusion","content":"\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eIn this work, we introduced a hybrid Quantum Dilated Convolutional Neural Net- work (QDCNN) that systematically investigates the role of quantum entanglement in capturing long-range temporal dependencies in sequential data. By synergistically integrating classical and quantum components, our model capitalizes on the comple- mentary strengths of hybrid architectures, aligning with the practical constraints and opportunities of near-term quantum devices. Our results demonstrate that quantum entanglement significantly enhances the model\u0026rsquo;s ability to process long-term temporal information, particularly when combined with the computational efficiency of classical processing. Through rigorous evaluation on benchmark tasks, including the Adding Problem and Music datasets, we demonstrated that the proposed QDCNN consistently outperforms its purely classical counterpart. Our architecture not only advances the design of quantum-enhanced temporal models but also underscores the importance of architectural refinements such as the clipping mechanism in optimizing learning performance within hybrid quantum-classical systems. This study paves the way for future research into efficient hybrid quantum-classical models for complex sequential tasks and contributes to the broader effort of integrating quantum computing into practical machine learning workflows.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll aspects of the work were completed solely by Rihab Hoceini\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eI would like to sincerely thank Dr. Ahmed Bouida for his guidance, support, and willingness to answer my questions throughout this work. His insights and encouragement were greatly appreciated.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMojtahedi, F.F., Yousefpour, N., Chow, S., Cassidy, M.: Deep learning for time series forecasting: Review and applications in geotechnics and geosciences. Archives of Computational Methods in Engineering, 1\u0026ndash;31 (2025)\u003c/li\u003e\n\u003cli\u003eCasolaro, A., Capone, V., Iannuzzo, G., Camastra, F.: Deep learning for time series forecasting: Advances and open problems. Information \u003cstrong\u003e14\u003c/strong\u003e(11), 598 (2023)\u003c/li\u003e\n\u003cli\u003eLuan, Y., Zhang, H., Zhang, C., Mu, Y., Wang, W.: Stock price prediction with sentiment analysis for chinese market. In: Proceedings of the Joint Work- shop of the 7th Financial Technology and Natural Language Processing, the 5th Knowledge Discovery from Unstructured Data in Financial Services, and the 4th Workshop on Economics and Natural Language Processing@ LREC-COLING 2024, pp. 167\u0026ndash;177 (2024)\u003c/li\u003e\n\u003cli\u003eZhang, Y., Yang, W., Wang, J., Ma, Q., Xiong, J.: Camef: Causal- augmented multi-modality event-driven financial forecasting by integrating time series patterns and salient macroeconomic announcements. arXiv preprint arXiv:2502.04592 (2025)\u003c/li\u003e\n\u003cli\u003eLi, X., Sakevych, M., Atkinson, G., Metsis, V.: Biodiffusion: A versatile diffusion model for biomedical signal synthesis. Bioengineering \u003cstrong\u003e11\u003c/strong\u003e(4), 299 (2024)\u003c/li\u003e\n\u003cli\u003eWang, C.-C.: T 2-lstm-based ai system for early detection of motor failure in chemical plants. Mathematics (2227-7390) \u003cstrong\u003e12\u003c/strong\u003e(17) (2024)\u003c/li\u003e\n\u003cli\u003eSu, L., Zuo, X., Li, R., Wang, X., Zhao, H., Huang, B.: A systematic review for transformer-based long-term series forecasting. Artificial Intelligence Review \u003cstrong\u003e58\u003c/strong\u003e(3), 80 (2025)\u003c/li\u003e\n\u003cli\u003eJohnston, L., Patel, V., Cui, Y., Balaprakash, P.: Revisiting the problem of learn- ing long-term dependencies in recurrent neural networks. Neural Networks \u003cstrong\u003e183\u003c/strong\u003e, 106887 (2025)\u003c/li\u003e\n\u003cli\u003eElman, J.L.: Finding structure in time. Cognitive science \u003cstrong\u003e14\u003c/strong\u003e(2), 179\u0026ndash;211 (1990)\u003c/li\u003e\n\u003cli\u003eHochreiter, S., Schmidhuber, J.: Long short-term memory. Neural computation \u003cstrong\u003e9\u003c/strong\u003e(8), 1735\u0026ndash;1780 (1997)\u003c/li\u003e\n\u003cli\u003eBai, S., Kolter, J.Z., Koltun, V.: An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271 (2018)\u003c/li\u003e\n\u003cli\u003eWaqas, M., Humphries, U.W.: A critical review of rnn and lstm variants in hydrological time series predictions. MethodsX, 102946 (2024)\u003c/li\u003e\n\u003cli\u003eHaider, A., Lee, G., Jafri, T.H., Yoon, P., Piao, J., Jhang, K.: Enhancing accuracy of groundwater level forecasting with minimal computational complexity using temporal convolutional network. Water \u003cstrong\u003e15\u003c/strong\u003e(23), 4041 (2023)\u003c/li\u003e\n\u003cli\u003ePeral-Garc\u0026acute;ıa, D., Cruz-Benito, J., Garc\u0026acute;ıa-Pen\u0026tilde;alvo, F.J.: Systematic literature review: Quantum machine learning and its applications. Computer Science Review \u003cstrong\u003e51\u003c/strong\u003e, 100619 (2024)\u003c/li\u003e\n\u003cli\u003eSenokosov, A., Sedykh, A., Sagingalieva, A., Kyriacou, B., Melnikov, A.: Quan- tum machine learning for image classification. Machine Learning: Science and Technology \u003cstrong\u003e5\u003c/strong\u003e(1), 015040 (2024)\u003c/li\u003e\n\u003cli\u003eCorli, S., Moro, L., Dragoni, D., Dispenza, M., Prati, E.: Quantum machine learn- ing algorithms for anomaly detection: A review. Future Generation Computer Systems, 107632 (2024)\u003c/li\u003e\n\u003cli\u003eMacaluso, A.: Quantum supervised learning. KI-Ku\u0026uml;nstliche Intelligenz, 1\u0026ndash;15 (2024)\u003c/li\u003e\n\u003cli\u003eNguyen, N., Chen, K.-C.: Quantum embedding search for quantum machine learning. IEEE Access \u003cstrong\u003e10\u003c/strong\u003e, 41444\u0026ndash;41456 (2022)\u003c/li\u003e\n\u003cli\u003eBowles, J., Ahmed, S., Schuld, M.: Better than classical? the subtle art of bench- marking quantum machine learning models. arXiv preprint arXiv:2403.07059 (2024)\u003c/li\u003e\n\u003cli\u003eBroughton, M., Verdon, G., McCourt, T., Martinez, A.J., Yoo, J.H., Isakov, S.V., Massey, P., Halavati, R., Niu, M.Y., Zlokapa, A., et al.: Tensorflow quan- tum: A software framework for quantum machine learning. arXiv preprint arXiv:2003.02989 (2020)\u003c/li\u003e\n\u003cli\u003eKong, Y., Wang, Z., Nie, Y., Zhou, T., Zohren, S., Liang, Y., Sun, P., Wen, Q.: Unlocking the power of lstm for long term time series forecasting. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, pp. 11968\u0026ndash;11976 (2025)\u003c/li\u003e\n\u003cli\u003eSoni, R., Alam, M.S., Vishwakarma, G.K.: Prediction of insar deformation time- series using improved lstm deep learning model. Scientific Reports \u003cstrong\u003e15\u003c/strong\u003e(1), 5333 (2025)\u003c/li\u003e\n\u003cli\u003eChambers, J.D., Cook, M.J., Burkitt, A.N., Grayden, D.B.: Using long short-term memory (lstm) recurrent neural networks to classify unprocessed eeg for seizure prediction. Frontiers in Neuroscience \u003cstrong\u003e18\u003c/strong\u003e, 1472747 (2024)\u003c/li\u003e\n\u003cli\u003eSingh, R., Lu, D., Tayal, K.: Temporal sequence transformer to advance long-term streamflow prediction. In: Proceedings of the NeurIPS 2024 Workshop on Tackling Climate Change with Machine Learning, Vancouver, British Columbia, Canada (2024). Part of the Conference on Neural Information Processing Systems, Dec. 10\u0026ndash;15, 2024\u003c/li\u003e\n\u003cli\u003eHwang, D., Wang, W., Huo, Z., Sim, K.C., Mengibar, P.M.: Transformerfam: Feedback attention is working memory. arXiv preprint arXiv:2404.09173 (2024)\u003c/li\u003e\n\u003cli\u003eGao, R., Wang, L.: Memotr: Long-term memory-augmented transformer for multi-object tracking. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9901\u0026ndash;9910 (2023)\u003c/li\u003e\n\u003cli\u003eSun, P., Wu, J., Zhang, M., Devos, P., Botteldooren, D.: Delayed memory unit: modeling temporal dependency through delay gate. IEEE Transactions on Neural Networks and Learning Systems (2024)\u003c/li\u003e\n\u003cli\u003eJia, Y., Lin, Y., Yu, J., Wang, S., Liu, T., Wan, H.: Pgn: The rnn\u0026rsquo;s new successor is effective for long-range time series forecasting. Advances in Neural Information Processing Systems \u003cstrong\u003e37\u003c/strong\u003e, 84139\u0026ndash;84168 (2024)\u003c/li\u003e\n\u003cli\u003eYe, H., Zhang, Y., Liu, H., Li, X., Chang, J., Zheng, H.: Light recurrent unit: Towards an interpretable recurrent neural network for modeling long-range dependency. Electronics \u003cstrong\u003e13\u003c/strong\u003e(16), 3204 (2024)\u003c/li\u003e\n\u003cli\u003eChittoor, H.H.S., Griffin, P.R., Neufeld, A., Thompson, J., Gu, M.: Qultsf: Long- term time series forecasting with quantum machine learning. arXiv preprint arXiv:2412.13769 (2024)\u003c/li\u003e\n\u003cli\u003eHsu, Y.-C., Chen, N.-Y., Li, T.-Y., Chen, K.-C., et al.: Quantum kernel- based long short-term memory for climate time-series forecasting. arXiv preprint arXiv:2412.08851 (2024)\u003c/li\u003e\n\u003cli\u003eMoon, K.-H., Jeong, S.-G., Hwang, W.-J.: Qsegrnn: quantum segment recurrent neural network for time series forecasting. EPJ Quantum Technology \u003cstrong\u003e12\u003c/strong\u003e(1), 32 (2025)\u003c/li\u003e\n\u003cli\u003eChen, Y.: Quantum dilated convolutional neural networks. IEEE Access \u003cstrong\u003e10\u003c/strong\u003e, 20240\u0026ndash;20246 (2022)\u003c/li\u003e\n\u003cli\u003eWang, Z., Ma, Y., Liu, Z., Tang, J.: R-transformer: Recurrent neural network enhanced transformer. arXiv preprint arXiv:1907.05572 (2019)\u003c/li\u003e\n\u003cli\u003eBoulanger-Lewandowski, N., Bengio, Y., Vincent, P.: Modeling temporal depen- dencies in high-dimensional sequences: Application to polyphonic music genera- tion and transcription. arXiv preprint arXiv:1206.6392 (2012)\u003c/li\u003e\n\u003cli\u003eChen, L., Zhong, X., Zhang, F., Cheng, Y., Xu, Y., Qi, Y., Li, H.: Fuxi: A cas- cade machine learning forecasting system for 15-day global weather forecast. npj climate and atmospheric science \u003cstrong\u003e6\u003c/strong\u003e(1), 190 (2023)\u003c/li\u003e\n\u003cli\u003eChen, K., Han, T., Gong, J., Bai, L., Ling, F., Luo, J.-J., Chen, X., Ma, L., Zhang, T., Su, R., et al.: Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948 (2023)\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Quantum entanglement, Temporal dependency, quantum neural network, sequential data","lastPublishedDoi":"10.21203/rs.3.rs-7095714/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7095714/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eGiven the current gap in the literature regarding whether quantum computing can effectively address long-term dependency modeling in sequential learning tasks, this study seeks to explicitly investigate the potential benefits of quantum entanglement in this context. We propose a hybrid Quantum Dilated Con- volutional Neural Network (QDCNN) that synergistically integrates quantum entanglement with classical processing, leveraging the complementary strengths of both computational paradigms. Our architecture employs quantum dilated convolutions with exponentially increasing dilation rates, enabling the efficient capture of long-range temporal dependencies without a proportional increase in quantum circuit depth. To address the limitations of near-term quantum hard- ware, we introduce a novel clipping mechanism that ensures physically realizable entanglement when dilation exceeds the quantum register size, while preserving global quantum correlations. The source code for the proposed quantum model is available at : \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/HoceiniRihab/Quantum-dilated-CNN\u003c/span\u003e\u003cspan address=\"https://github.com/HoceiniRihab/Quantum-dilated-CNN\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e","manuscriptTitle":"On the Role of Quantum Entanglement in Capturing Long-Term Dependencies","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-16 10:24:17","doi":"10.21203/rs.3.rs-7095714/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"eb4664be-3a31-49a1-bccc-8cc663d187cc","owner":[],"postedDate":"September 16th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-12-29T19:53:31+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-16 10:24:17","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7095714","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7095714","identity":"rs-7095714","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00