Enhancing CNN-Attention-LSTM with Novel LSTM Gate Mechanism: Evaluation on Benchmark and Brain Tumor Imaging Datasets

preprint OA: closed
View at publisher

Abstract

Abstract This work introduces the gating mechanisms of the Long Short-Term Memory (LSTM) architecture, proposing an innovative architecture. Traditionally, LSTM models control information flow and address gradient vanishing and explosion problems by incorporating sigmoid and tanh activation functions into their input, forget, and output gates. In earlier studies, the sigmoid function was modified with keeping and discarding of relevant and irrelevant information, respectively, and the tanh function was adjusted to solve the gradient problem. These methods have, however, usually only focused on one of the functions, leaving an opportunity for additional optimization. This work provides a novel method that combines changes to the sigmoid and tanh functions inside the LSTM gates. This dual adjustment is meant to maximise the benefits of both improved gradient stability and increased information gating. Prior to the LSTM model, it included a CNN layer to extract and remove irrelevant data. The attention layer focusses on the most essential features. By combining the advantages of LSTMs for temporal dependency control, attention layer with the query and key values for focusing on the relationship between the different neighbourhood pixels and CNNs for handling spatial data, this combination creates a strong hybrid model (CNN-attention-LSTM). It performs extensive evaluations on benchmark datasets, such as Fashion MNIST, MNIST, and CIFAR-10, to check the impact of the proposed approach. When compared to previously proposed models with different activation functions, our results show improved accuracy and lower loss. To verify the applicability and generalisability of our model in an essential domain, it further extends the evaluation to a medical dataset, namely the brain tumor dataset in NIfTI format. The experimental results confirm that the hybrid CNN-attention-LSTM model and novel activation functions provide enhanced stability in gradient propagation in addition to higher accuracy.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00