Prediction Model For Digital Image Tampering Using Customized Deep Neural Network Techniques

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This study evaluated customized and standard CNN architectures for image tampering detection, with the proposed model achieving a 96% F1 score using separable convolutional layers.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

This paper studies automated digital image tampering detection using convolutional neural networks, comparing ResNet50V2, InceptionNetV3, MobileNetV2, and a custom CNN designed for classification of four image types: copy-move, inpaint, splicing, and normal images. The authors train and evaluate the models using standardized metrics including accuracy, precision, recall, and F1-score, and they specifically explore the use of separable convolutional layers within the custom architecture to improve scalability and effectiveness. They report that the proposed customized model achieves the highest validation and training accuracy among those compared and attains an F1 score of 96%, indicating strong discrimination of tampered regions with fewer false positives. The paper does not state a clear dataset size, external validation approach, or peer-reviewed limitations in the provided text, beyond noting it is a preprint/published work. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Image tampering detection is a critical area of research, given the widespread use of manipulated images for deceptive purposes. Convolutional Neural Networks (CNNs) have shown significant potential in automating the identification of tampered images. This paper presents customized deep learning model to detect tampering class with comparative analysis of CNN architectures - ResNet50V2, InceptionNetV3, MobileNetV2, and the proposed CNN, for image tampering detection. The proposed approach encompasses a dataset comprising four distinct classes: copy-move, inpaint, splicing, and normal images. This study sheds light on the comparative strengths and weaknesses of these CNN architectures. The dataset encompasses the key tampered classes, offering a holistic assessment of each model's ability to identify various tampering techniques. The custom CNN architecture is specifically tailored for this task, aiming to evaluate its efficiency compared to the established CNNs. Metrics for training and evaluation are standardized to generate equitable comparisons, encompassing performance indicators such as accuracy, precision, recall, and F1-score.This research contributes the knowledge in the field of image tampering detection, offering a comprehensive evaluation of multiple CNN architectures. Additionally, the effectiveness of separable convolutional layers is explored in deep neural networks, showcasing their potential to enhance scalability and effectiveness across various tasks in machine learning and computer vision. The proposed model, designed with separable convolution layers, exhibits superior validation accuracy and training accuracy compared to the other models under evaluation. Notably, The proposed customized model achieved an impressive F1 score of 96%, highlighting its proficiency in accurately detecting tampered regions within images while minimizing false positives.
Full text 97,762 characters · extracted from preprint-html · click to expand
Prediction Model For Digital Image Tampering Using Customized Deep Neural Network Techniques | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Prediction Model For Digital Image Tampering Using Customized Deep Neural Network Techniques Sachin Saxena, Archana Singh, Shailesh Tiwari This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4273139/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 19 Jul, 2024 Read the published version in International Journal of System Assurance Engineering and Management → Version 1 posted 5 You are reading this latest preprint version Abstract Image tampering detection is a critical area of research, given the widespread use of manipulated images for deceptive purposes. Convolutional Neural Networks (CNNs) have shown significant potential in automating the identification of tampered images. This paper presents customized deep learning model to detect tampering class with comparative analysis of CNN architectures - ResNet50V2, InceptionNetV3, MobileNetV2, and the proposed CNN, for image tampering detection. The proposed approach encompasses a dataset comprising four distinct classes: copy-move, inpaint, splicing, and normal images. This study sheds light on the comparative strengths and weaknesses of these CNN architectures. The dataset encompasses the key tampered classes, offering a holistic assessment of each model's ability to identify various tampering techniques. The custom CNN architecture is specifically tailored for this task, aiming to evaluate its efficiency compared to the established CNNs. Metrics for training and evaluation are standardized to generate equitable comparisons, encompassing performance indicators such as accuracy, precision, recall, and F1-score.This research contributes the knowledge in the field of image tampering detection, offering a comprehensive evaluation of multiple CNN architectures. Additionally, the effectiveness of separable convolutional layers is explored in deep neural networks, showcasing their potential to enhance scalability and effectiveness across various tasks in machine learning and computer vision. The proposed model, designed with separable convolution layers, exhibits superior validation accuracy and training accuracy compared to the other models under evaluation. Notably, The proposed customized model achieved an impressive F1 score of 96%, highlighting its proficiency in accurately detecting tampered regions within images while minimizing false positives. Convolutional Neural Network Image Tampering Fake Images Image Forensic Deep Learning Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1. Introduction In the digital age, the manipulation of images has become a prevalent and concerning issue. With the proliferation of image-editing software and the ease of sharing visuals on various platforms, image tampering has evolved into a significant challenge for preserving the integrity and authenticity of digital media. Detecting image tampering is not only essential for maintaining trust and credibility in various domains but is also crucial for criminal investigations, journalism, and ensuring the accuracy of visual content in social media. Traditional methods for image tampering detection have relied on manual examination, expert analysis, and the application of heuristics. While these approaches can be effective in some cases, they often fall short when dealing with sophisticated tampering techniques. Convolutional Neural Networks (CNNs), a potent category of deep learning models that have shown impressive abilities in tackling intricate computer vision challenges. This paper delves into the domain of image tampering detection and presents a complete performance analysis of four distinct CNN architectures: ResNet50V2, InceptionNetV3, MobileNetV2, and a custom-designed CNN architecture specifically crafted for this task. Image tampering is a powerful tool for creating deceptive visuals that can mislead, manipulate, or deceive the viewer. In the realm of journalism, manipulated images can be used to distort facts, sensationalize events, or even fabricate entirely fictitious stories. Misleading visuals can also be employed for political propaganda, harming the credibility of institutions and public trust [ 1 ]. In criminal investigations, the integrity of digital evidence is paramount. Manipulated images can be used to alter crime scenes, obscure identities, or erase incriminating evidence. Detecting such tampering is crucial for law enforcement agencies and legal proceedings, as a single manipulated image can alter the course of justice [ 2 ].The rise of social media platforms and online communication has given individuals the power to share images globally. However, this convenience has also made it easier for malicious actors to disseminate fake or manipulated visuals, leading to the spread of misinformation and disinformation. Detecting and addressing image tampering is essential for maintaining the credibility of content shared on these platforms [ 3 ]. In addressing the pressing concern of fake detection, the integration of Generative Adversarial Networks (GANs) in deep learning has demonstrated significant prowess, particularly in the manipulation of faces, commonly referred to as Deepfake. Deepfake, driven by GANs, enables the realistic substitution of faces, posing a threat of spreading fabricated events through online channels. To counter this, [ 10 ] proposes an intelligent forensic method for Deepfake detection. Their approach focuses on subtle texture disparities in image saliency, particularly in facial textures, utilizing a guided filter with a saliency map to enhance texture artifacts. The ResNet18 classification network efficiently learns these features, achieving state-of-the-art detection accuracy in distinguishing between real and fake face images [ 10 ]. Generative adversarial networks (GANs) have been investigated for their ability to discern real images from fake ones based on texture differences in tampered images. However, challenges persist, as the proposed methodology still necessitates large-scale data to yield satisfactory results [ 16 ]. Recent efforts, exemplified by [ 11 ], have explored novel strategies leveraging resampling features, Long-Short Term Memory (LSTM) cells, and encoder-decoder networks. Their approach adeptly captures artifacts like JPEG quality loss and spatial transformations, enabling discriminative analysis between manipulated and non-manipulated regions. The method achieves pixel-level localization, demonstrating high precision in identifying image manipulations. The introduction of a new large image splicing dataset for training further enhances the robustness of their approach. Through rigorous experimentation on diverse datasets, [ 11 ] showcase the efficacy of their method in addressing the intricate challenges posed by subtle image manipulations. Recent studies have identified various methods for tampering detection. However, a common challenge persists among these methods, as they often concentrate on specific manipulation techniques by extracting explicit features from the images [ 12 ]. Pre-trained models have been utilized for identifying manipulated images. ResNet50 attained higher accuracy; nevertheless, it suggested enhancing accuracy through the utilization of image patches [ 13 ]. Guided filters have been utilized for extracting texture features aimed at identifying tampered images [ 14 ]. InceptionNet and MobileNet have been applied to the classification of medical images. However, there is a need to further enhance the accuracy to ensure broader acceptance of detection methodologies [ 15 ]. The various approaches discussed above, perform automatic feature extraction and recognition from image data, however, challenges arise in tasks such as determining valid data, handling de-duplication, and verifying data authenticity before the application of machine learning. The proposed approach provided better accuracy by feeding verified image data into different layers of CNN. This uniqueness distinguishes the approach, as existing methods often rely on machine learning techniques and authentication mechanisms in isolation. The objective of this paper is to formulate a robust predictive model employing deep learning techniques for the discernment of counterfeit images. This model is designed to accurately identify distinct classes of tampering applied to images, ensuring a comprehensive and precise detection mechanism. Furthermore, this paper thoroughly explores sophisticated feature extraction methods, aiming to enhance the precision and nuanced classification of images according to distinct tampering classes like copy-move, region removal, and splicing unlikely done in the extant research. The proposed approach of work enhances the accuracy of the model in less computation time. The paper is organized as follows; Following the introduction which presents the pressing issue of digital tampering of images, recent work done in this area, challenges, and objectives of the paper. Section 2 mentioned the techniques and meth-ods used to propose the custom CNN architecture used in this paper. Section 3 pre-sent the research methodology describing the dataset used and the proposed 4 architecture. In section 4 , the experiment and results are mentioned followed by the conclusion in the last section 5 . 2. Methods and Techniques – Deep Neural Networks Convolutional Neural Networks (CNNs) have emerged as a groundbreaking technology in the field of computer vision [ 4 ], particularly in addressing complex image analysis tasks. CNNs are a class of deep neural networks designed to process and analyze visual data, making them highly effective in tasks such as image classification, object detection, and image tampering detection [ 4 ].CNNs consist of convolutional layers, which apply filters to input images, extracting features like edges and textures. These layers slide small filters over the image, computing dot products with local regions to generate feature maps. Pooling layers, like max-pooling and average-pooling, are integral components of CNNs and serve to decrease spatial dimensions, aiding computation and emphasizing important features. After convolution and pooling, fully connected layers capture advanced abstractions and relationships among features, facilitating intricate decision-making processes. Activation functions like ReLU [ 8 ] and Sigmoid introduce non-linearity, allowing CNNs to model intricate data relationships. CNNs undergo training through backpropagation, a process that involves adjusting internal parameters such as weights and biases to minimize a loss function, typically a measure of predicted vs. actual labels. In this paper analysis of four CNN architectures—ResNet50V2, InceptionNetV3, MobileNetV2, and a custom-designed CNN specifically tailored for image tampering detection is shown by experimental results. ResNet50V2 [ 5 ], short for Residual Network with 50 layers and Version 2 is particularly recognized for its outstanding performance in image classification. At its core, ResNet50V2 utilizes residual blocks with skip connections to mitigate the vanishing gradient problem in deep networks. The architecture incorporates bottleneck structures within these blocks to reduce computational load while maintaining representational capacity [ 5 ]. InceptionNetV3 [ 6 ], or Inception-v3, is recognized for exceptional performance in image classification and object recognition, Inception-v3 employs specialized inception modules, allowing multi-scale feature capture and intricate pattern detection. It incorporates global average pooling to reduce dimensionality, aiding in mitigating overfitting, and benefits from batch normalization for stabilized training [ 4 ]. MobileNetV2 [ 7 ], designed for efficient on-device computer vision on mobile and embedded devices, enhances its predecessor with innovative features. Linear bottlenecks and shortcut connections enhance information flow and stability [ 7 ]. The scalable design includes width and resolution multipliers for flexibility in balancing model size and accuracy. Efficient convolutions reduce computational overhead, making MobileNetV2 ideal for real-time applications like image classification, object detection, and semantic segmentation [ 7 ]. Addressing the challenges posed by image tampering requires advanced tools capable of automatically identifying manipulated images with precision and accuracy. Convolutional Neural Networks (CNNs) [ 4 ], a type of deep learning models, have shown immense promise in this regard. CNNs are designed to learn and extract intricate patterns and features directly from the data they are trained on [ 4 ]. This ability makes them highly suitable for image tampering detection, as tampering techniques often involve subtle changes in pixel values, textures, and patterns. Traditional methods for image tampering detection often rely on manually crafted features and heuristics. CNNs, on the other hand, have the capacity to automatically discover and leverage complex features that may be challenging to define explicitly. This feature extraction capability allows CNNs to excel in identifying tampered regions [ 4 ], even in cases where the manipulation is sophisticated or subtle. There are various methods exist for determining the authenticity of an image, and these approaches are briefly examined. In the realm of tampered text detection in document images, the critical role of information security has drawn increasing attention. While progress has been made, detecting visually consistent tampered text in photographed document images remains challenging. Document Tampering Detector (DTD) framework, leveraging a Frequency Perception Head (FP H) and a Multi-view Iterative Decoder (MID) focuses on visual feature challenges. The novel Curriculum Learning for Tampering Detection (CLTD) training paradigm is designed to enhance robustness, mitigating confusion during training [ 9 ]. In the proposed work, the CNN model and its variations ResNet50V2, Inception-NetV3 and MobileNetV2 were used to define the proposed architecture of the customized CNN model explained in the following section 3 . 3. Research Methodology In the proposed work, a more extensive approach is opted to fine-tune the pre-trained models. This involved unfreezing all layers of the models, not just the fully connected layers, allowing for modifications to both the convolutional and fully connected layers. The rationale behind this approach was to ensure that the models could adapt comprehensively to the unique demands of the image tampering detection task. The combination of the datasets, especially the augmented Defacto dataset, allows for a comprehensive assessment of CNNs in image tampering detection. The Defacto dataset provides tampered images with various scenarios and severity levels, further enhanced by rotation augmentation, while the COCO dataset offers diverse authentic images for assessing the models' ability to distinguish between tampered and normal images in real-world scenarios [ 17 ]. This dataset diversity and structure make them suitable for robust training, evaluation, and benchmarking of machine learning models dedicated to image tampering detection, offering an extensive and diverse dataset resource for research and experimentation in this field. 3.1. Dataset In the context of image tampering detection, the choice of an appropriate dataset is crucial for training and evaluating machine learning models, including Convolutional Neural Networks (CNNs). The research paper discusses two key datasets: the Defacto dataset and the COCO dataset [ 17 ]. The Defacto dataset[[ 18 ], containing 72,776 images per class after rotation augmentation, is specialized for image tampering detection. Multiple python scripts are developed for converting the raw images into tampered images of 4 categories such as copy-move, inpaint, and splicing,and normal with varying levels of tampering severity. Sample image form the each class is shown in the Fig. 1.Ground truth annotations for tampered regions within each image are provided, enabling supervised machine learning model training and performance evaluation. On the other hand, the COCO dataset serves as a valuable source for authentic, unaltered images, representing the "normal" class. It consists of diverse high-quality images capturing everyday scenes, objects, and contexts, with detailed annotations for object categories, locations, and segmentation masks. The combination of high-end GPUs and CPUs created a harmonious synergy that facilitated efficient model training, leading to optimal utilization of computational resources. In determining the training parameters, a batch size of 128 has been chosen and executed training for a span of 20 epochs. The selection of these parameters was a result of a carefully crafted balance between computational resources and model convergence. A batch size of 128 strikes a balance between using large enough batches for efficient gradient updates while not exceeding the memory capacity of the GPUs, which could lead to bottlenecks in training. The choice of 20 epochs for training duration was equally well-considered. It provided the models with ample opportunities to learn and adapt to the intricate features within the image tampering detection dataset while avoiding overfitting. 3.2. Proposed Architecture: In addition to the utilization of pre-trained models, custom CNN architecture is proposed, specifically designed to address the unique challenges posed by image tampering detection as shown in Fig. 2 . While pre-trained models serve as powerful baselines, the development of a custom architecture was a crucial component of proposed research, as it allowed us to tailor the network to the specific requirements of this task. At the foundation lies the input layer, accepting raw image data. The initial CNN block forms the base, composed of two CNN layers and a Batch Normalization layer, establishing the groundwork for feature extraction. A subsequent max-pooling layer condenses information, marking a pivotal transition. The following are separable convolution blocks, represented as sturdy columns, emphasizing the model's proficiency in discerning intricate features. Max-pooling layers intermittently guide the flow, each contributing to spatial reduction. The model evolves through these blocks, culminating in a compact global average pooling layer. Finally, a fully connected layer crowns the structure, epitomizing the culmination of feature extraction for precise classification of tampered class. A distinguishing feature of proposed customized CNN architecture is the incorporation of separable convolution layers. Traditional CNNs employ standard convolution layers, which involve a high degree of interdependence between convolutional filters. In contrast, separable convolution layers separate the spatial and depth-wise convolutions. This separation significantly reduces computational complexity and enhances the model's capacity to learn features more efficiently and rapidly. It's important to emphasize that this innovation wasn't just about computational efficiency; it also brought tangible benefits to the model's representational power. The custom CNN, along with separable convolution layers, became a crucial asset in the comparative analysis. This architecture not only demonstrated its potential for more efficient feature learning but also highlighted the importance of architectural innovations when tackling image tampering detection. 4. Experiments and Results A comprehensive evaluation of the model’s performance unveils distinctive trends for loss and accuracy as mentioned in Table 1 . The custom model, crafted with separable convolution layers, displays the highest validation accuracy and training accuracy among the models being assessed. It notably surpasses the performance of both InceptionNetV3 and ResNet50V2, attaining a validation accuracy that stands out significantly as shown in Fig. 6 . The results are also mirrored in its training accuracy, where the proposed customized model displays robust learning capabilities. Table 1 Accuracy and Loss of Deep Learning Models Model Validation Accuracy Training Accuracy Training Loss Validation Loss Custom model 0.9000 0.9688 0.0137 0.1043 InceptionNetV3 0.7148 0.9240 0.5184 0.4958 ResNet50V2 0.6716 0.7196 0.5068 0.5854 MobileNetV2 0.7727 0.9026 0.5718 0.6068 Furthermore, the custom model excels in minimizing training loss and validation loss, demonstrating its effectiveness in image tampering detection. In contrast, InceptionNetV3, though renowned for its feature extraction prowess [ 6 ], achieves 71.48% validation accuracy as shown in Fig. 3 . This is further emphasized by its comparatively higher training and validation loss values which are 51.84% and 49.58% respectively, indicating a struggle to achieve good precision. ResNet50V2 model known for its depth and skip connections [ 5 ], demonstrates 67.16% validation accuracy in comparison, emphasizing the advantages of customization as shown in Fig. 4 . Its training accuracy falls short of expectations, and the training and validation losses remain relatively high, indicating a challenge in capturing the intricate features essential for tampering detection. MobileNetV2 [ 7 ], as shown in Fig. 5 while showing a competitive validation accuracy and training accuracy which is 77.27% and 90.26% respectively, slight lags in the validation loss section, which could be attributed to its inherent trade-offs in terms of computational efficiency. Nonetheless, the findings underline the superior performance too, which successfully combines efficient feature learning with faster convergence, showcasing its potential as a promising solution for image tampering detection. In the evaluation of image tampering detection models, precision, recall, and F1 score serve as crucial performance metrics. In this context, the results highlight the exceptional performance of the custom CNN, showcasing its capability in achieving a harmonious balance between precision and recall. The custom model achieved an impressive F1 score of 0.96, indicating its proficiency in both identifying tampered regions within images and minimizing false positives as mentioned in Table 2. Table 2. Precision, Recall and F1 Score of Deep Learning Models Model Precision Recall F1 Score InceptionNetV3 0.8586 0.8541 0.8538 ResNet50V2 0.7409 0.7263 0.7278 MobileNetV2 0.8153 0.8158 0.8154 Custom model 0.9562 0.9683 0.9622 The proposed customized model demonstrates a strong F1 score of 0.9622, indicating a symmetric precision and recall. In comparison, InceptionNetV3 achieves an F1 score of 0.8538, ResNet50V2 records 0.7278, and MobileNetV2 attains 0.8154. These scores signify the nuanced trade-off between precision and recall for each model, offering insights into their respective strengths in identifying tampered regions within images, positioning it as an asset in applications. In results and analysis, it is evident that training times for each model significantly impact their practical utility. As shown in Fig. 7 , The proposed CNN outshines the competition with a training time of 573.4 seconds for 20 epochs, which is a substantial improvement over the other models. In comparison, MobileNetV2 required 2250.5 seconds, InceptionNetV3 took 2413.5 seconds, and ResNet50V2 demanded 2220.5 seconds for the same number of epochs. What's particularly striking is that proposed model, despite having a similar number of parameters as MobileNetV2, reduced the training time by nearly 75%. This reduction in training time showcases the efficiency of the proposed customized CNN, thanks in part to the utilization of separable convolution layers, which expedite the learning process and enhance overall training efficiency. A comprehensive evaluation of the model parameters underscores the efficiency and compactness of the proposed customized CNN architecture in image tampering detection. With a remarkably lean parameter count of 5,450,988, the custom model showcases an adept balance between complexity and performance as shown in Fig. 8 . In contrast, InceptionNetV3, ResNet50V2, and MobileNetV2 exhibit substantially larger parameter counts, registering at 25,878,802, 27,908,998, and 5,047,298, respectively. This notable discrepancy in parameter sizes emphasizes the streamlined architecture of custom model, demonstrating that good performance need not necessitate an excessive number of parameters. The judicious management of parameters in customized CNN not only contributes to computational efficiency but also underscores its potential for applications where resource constraints are a critical consideration. 4.1 Hyperparameter Optimization Hyper parameter optimization is a critical component of the proposed research methodology. To guarantee the optimal fine-tuning of custom CNN and the pretrained models (ResNet50V2, InceptionNetV3, and MobileNetV2) for image tampering detection, an extensive and systematic hyperparameter search has been used. The hyperparameter search involved running hundreds of training sessions, each characterized by a unique combination of hyperparameters. Key parameters that were subjected to this optimization process included learning rates, weight decay, dropout rates, and architecture-specific settings. Each of these hyperparameters plays a crucial role in the performance and behavior of a neural network. Therefore, determining the most suitable values for these hyperparameters was paramount. By systematically exploring the hyperparameter space, the model gained insights into the intricate interplay between these parameters and the models' ability to generalize and adapt effectively to the proposed work. This comprehensive approach to hyperparameter optimization reflects to ensuring that model, reached its full potential. It not only enhanced the performance of the models but also contributed to the reliability and robustness of comparative analysis, providing a solid foundation for the findings and conclusions. 4.2 Optimization Techniques: In proposed work, the choice of optimization techniques played a crucial role in fine-tuning the pretrained CNN architectures for image tampering detection. Adam optimizer has been used which is widely acclaimed and renowned for its efficacy in training deep neural networks. Adam combines the benefits of both the momentum-based updates. This adaptive optimization algorithm dynamically modifies the learning rates for individual parameters, leading to accelerated convergence and more effective training. To further enhance the training performance and prevent the models from converging prematurely or getting stuck in local minima, learning rate scheduling mechanism has been implemented. Learning rate scheduling is a dynamic adjustment of the learning rate during training. This mechanism monitored the model's performance, and when it detected a plateau in accuracy, it automatically decreased the learning rate. 4.3 Addressing Overfitting: Overfitting, a common hurdle in deep learning, occurs when a model excessively tailors itself to the training data, often incorporating noise and irrelevant patterns that hinder its ability to generalize effectively to unseen data. To tackle this issue, various strategies were executed with the aim of improving the models' ability to perform well on data they had not encountered before. First and foremost, L1 regularization has been applied which was a vital tool in arsenal against overfitting. L1 regularization works on to the network's weights, encouraging sparsity within the model. This regularization technique introduced a penalty on the magnitude of the weights, promoting a more parsimonious set of features. By doing so, L1 regularization encouraged the model to focus on the most relevant and informative features while discouraging the overemphasis on noise and irrelevant details in the training data. This process played a significant role in enhancing the model's performance with unseen data, a critical factor in image tampering detection where adaptability and accuracy are paramount. The dropout layers are incorporated at strategic points in the network architecture. These dropout layers played a pivotal role in preventing overfitting by randomly deactivating a fraction of neurons during training. This randomness introduced a degree of uncertainty into the model's learning process, effectively discouraging it from becoming overly reliant on specific features or neurons. Furthermore, batch normalization [ 14 ] was a key component of the proposed strategy to combat overfitting. This technique was applied to standardize the input to each layer during training. By doing so, batch normalization mitigated the effects of internal covariate shift [ 19 ], a phenomenon where the distribution of layer inputs changes during training. Batch normalization aided in improving the models' generalization by ensuring that the learning process was not hindered by abrupt shifts in data distributions. It contributed to the overall stability and robustness of the models, which is paramount in the context of image tampering detection where detecting subtle variations and manipulations are essential. 5. Conclusion & Future Scope In conclusion, the investigation into Convolutional Neural Network (CNN) architectures for image tampering detection has unfolded notable advancements, notably exemplified by the proposed customized model. This exclusive CNN exhibited discernible proficiency, particularly in validation accuracy, performing better than the other considered architectures. Its precision in identifying tampered regions underscores its potential significance in digital forensics and content verification. Moreover, the custom model's noteworthy efficiency, characterized by reduced training time, stands as a testament to the impact of leveraging separable convolution layers. This computational efficacy not only conserves resources but also extends the model's adaptability to more extensive datasets and complex tasks. While recognizing the merits of established architectures, the suggested exploration emphasizes the inventive potential and adaptability encapsulated in neural network customization. Beyond the pursuit of elevated accuracy and training efficiency, this research contributes to the ongoing evolution of image tampering detection methodologies, paving the way for nuanced applications and reinforcing the digital landscape's integrity and trustworthiness. In future different image pre-processing technique could be used to enhance the performance. Additionally, a potential future scope of work could encompass the detection of fake videos. This expansion would involve adapting and extending the proposed model to address the unique challenges posed by video content. References Molina MD, Sundar SS, Le T, Lee D (2021) Fake News Is Not Simply False Information: A Concept Explication and Taxonomy of Online Content. Am Behav Sci 65(2):180–212. https://doi.org/10.1177/0002764219878224 Moussa AF (2021) Electronic evidence and its authenticity in forensic evidence. Egypt J Forensic Sci 11(1). https://doi.org/10.1186/s41935-021-00234-6,hp Aïmeur E, Amri S, Brassard G (2023) Fake news, disinformation and misinformation in social media: a review. Social Netw Anal Min 13(1). https://doi.org/10.1007/s13278-023-01028-5 O’Shea K, Nash RR, An (2015) introduction to convolutional neural networks. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1511.08458 He K, Zhang X, Ren S, Sun J (2016) Identity mappings in deep residual networks. https://doi.org/10.48550/arxiv.1603.05027 . arXiv (Cornell University) Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2015) Rethinking the inception architecture for computer vision. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1512.00567 Sandler M, Howard A, Zhu M, Zhmoginov A, Chen L (2018) MobileNetV2: Inverted residuals and linear bottlenecks. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1801.04381 Agarap AF (2019) Deep Learning using Rectified Linear Units (ReLU). arXiv [Cs.NE]. Retrieved from http://arxiv.org/abs/1803.08375 Qu C, Liu C, Liu Y, Chen X, Peng D, Guo F, Jin L (2023) Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5937–5946 Yang J, Xiao S, Li A, Lan G, Wang (2021) Detecting fake images by identifying potential texture differences. Future Generation Comput Syst 125:127–135. https://doi.org/10.1016/j.future.2021.06.043 Bappy JH, Simons C, Nataraj L, Manjunath BS, Roy-Chowdhury AK (2019) Hybrid LSTM and Encoder–Decoder Architecture for Detection of Image Forgeries. IEEE Trans Image Process 28(7):3286–3300. 10.1109/tip.2019.2895466 Manjunatha S, Malini M, Patil (2021) Deep learning-based Technique for Image Tamper Detection, Proceedings of the Third International Conference on Intelligent Communication Technologies and Virtual Mobile NetworksICICV EEE Nagaveni K, Hebbar, Kunte AS (2021) Transfer Learning Approach For Splicing And Copy-Move Image Tampering Detection. Ictact Journal On Image And Video Processing, May Jiachen Ya, Xiao S (2021) Aiyun Li a, GuipengLan, HuihuiWangb, Detecting fake images by identifying potential texture difference. Future Generation Computer Systems, Elsevier Image classification and prediction using transfer learning in colab notebook, J Praveen Gujjara,∗, H R Prasanna Kumar b, Niranjan N, Chiplunkar (2021) Global Transitions Proceedings Jiachen Ya, Li SXA, Wang GLH (2021) Detecting fake images by identifying potential texture difference. Future Generation Computer Systems, ELSEVIER – Lin T, Maire M, Belongie S, Bourdev L, Girshick R, Hays J, Perona P, Ramanan D, Zitnick CL, Dollár P, Microsoft COCO (2015) Common Objects in context. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1405.0312 MAHFOUDI G, TAJINI B, RETRAINT F, MORAIN-NICOLIER F, DUGELAY JL and M. PIC, DEFACTO: Image and Face Manipulation Dataset, 2019 27th European Signal Processing Conference (EUSIPCO), Coruna A (2019) Spain, pp. 1–5, 10.23919/EUSIPCO.2019.8903181.( 2014) Ioffe S, Szegedy C (2015) Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. arXiv [Cs.LG]. Retrieved from http://arxiv.org/abs/1502.03167 Cite Share Download PDF Status: Published Journal Publication published 19 Jul, 2024 Read the published version in International Journal of System Assurance Engineering and Management → Version 1 posted Editorial decision: Minor revisions 24 Jun, 2024 Reviewers agreed at journal 07 Jun, 2024 Reviewers invited by journal 30 May, 2024 Editor invited by journal 07 May, 2024 First submitted to journal 06 May, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4273139","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":308789340,"identity":"04208575-0255-489f-b3ab-44ce38110bb1","order_by":0,"name":"Sachin Saxena","email":"","orcid":"","institution":"Amity University ASET: Amity University Noida School of Engineering \u0026 Technology","correspondingAuthor":false,"prefix":"","firstName":"Sachin","middleName":"","lastName":"Saxena","suffix":""},{"id":308789341,"identity":"00aa80e9-8163-4ddb-8cce-e1ddde1648c3","order_by":1,"name":"Archana Singh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA3UlEQVRIie3NvQrCMBDA8YRCp1LXFLF9hZRCJx8mWdxcxa2dzrGzb1FwcUzp0MGia0Zd3AQ7CAqC9gMdm46C+ZPhCPfjENLpfjIc128aJfUo2g8xjMzwOq6XxTDSqhyn4rOtIqMkh9NtaxhBseN5Bci1JesnRPJVMClNMyznqcgABY6KIIlh7IBlhaIjPFUR75A1hJAgubQkUhIqODgVUEpJd4VRFfElhzEGxoisr5R74q/LYz9xD8XZecCLjZL55rpcTD27UFxpMqzPhE2iXm8X79/xOUzodDrdf/UG16JSz5FO18UAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-5365-6351","institution":"Amity University","correspondingAuthor":true,"prefix":"","firstName":"Archana","middleName":"","lastName":"Singh","suffix":""},{"id":308789342,"identity":"4f252081-79e4-41b8-8498-ae8c459e81c7","order_by":2,"name":"Shailesh Tiwari","email":"","orcid":"","institution":"KIET Group of Institutions: Krishna Institute of Engineering \u0026 Technology","correspondingAuthor":false,"prefix":"","firstName":"Shailesh","middleName":"","lastName":"Tiwari","suffix":""}],"badges":[],"createdAt":"2024-04-16 04:34:51","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4273139/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4273139/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s13198-024-02420-w","type":"published","date":"2024-07-19T16:13:27+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":58308849,"identity":"bd14fa00-282e-40a0-8bdb-d2b3f91efb1d","added_by":"auto","created_at":"2024-06-13 18:49:58","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":131604,"visible":true,"origin":"","legend":"\u003cp\u003eImage Tampering Classes\u003c/p\u003e","description":"","filename":"F1.png","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/d7f157c0121f0cd8da8bf899.png"},{"id":58308848,"identity":"f3b406fd-f73a-4d99-9d96-21cb8aa21b10","added_by":"auto","created_at":"2024-06-13 18:49:58","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":82970,"visible":true,"origin":"","legend":"\u003cp\u003eProposed Customized CNN Architecture\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/af82bed05d5548438780f8ee.jpg"},{"id":58308847,"identity":"31fcb00e-3237-4c15-add1-36ae15a216fc","added_by":"auto","created_at":"2024-06-13 18:49:58","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":13323,"visible":true,"origin":"","legend":"\u003cp\u003eInceptionNetV3 Training and Validation Accuracy\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/fc118f0968a46ac9efccf198.jpg"},{"id":58308845,"identity":"e8c5c4fd-b33f-4309-af57-6c54129e5c87","added_by":"auto","created_at":"2024-06-13 18:49:58","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":14053,"visible":true,"origin":"","legend":"\u003cp\u003eResNet50V2 Training and Validation Accuracy\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/1c4c85b654448c8a968bfa3c.jpg"},{"id":58308851,"identity":"4d7f6c71-aa59-4858-aa37-013f521d81fa","added_by":"auto","created_at":"2024-06-13 18:49:59","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":15520,"visible":true,"origin":"","legend":"\u003cp\u003eMobileNetV2 Training and Validation Accuracy\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/9440b7effe74d2156ba4535c.jpg"},{"id":58309734,"identity":"8cd6c856-9121-4aa5-afb4-b0b4387e168a","added_by":"auto","created_at":"2024-06-13 18:57:58","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":15357,"visible":true,"origin":"","legend":"\u003cp\u003eProposed Customized CNN Training and Validation Accuracy\u003c/p\u003e","description":"","filename":"Picture5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/0e4a5dc245d84b6470c61553.jpg"},{"id":58308852,"identity":"0bbaba6c-01a3-443f-833e-abd1198548b6","added_by":"auto","created_at":"2024-06-13 18:49:59","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":53634,"visible":true,"origin":"","legend":"\u003cp\u003eEpoch Execution Time of Deep Learning Models\u003c/p\u003e","description":"","filename":"Picture6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/ddf876c096329496a4de4ef1.jpg"},{"id":58308850,"identity":"c70eb9cd-540a-469e-b1a3-43d375575a15","added_by":"auto","created_at":"2024-06-13 18:49:59","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":47838,"visible":true,"origin":"","legend":"\u003cp\u003eParameter Comparison of Deep Learning Models\u003c/p\u003e","description":"","filename":"Picture7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/dea4054abef2711dd4a4ab0f.jpg"},{"id":61595338,"identity":"bd2436b0-5384-4db0-a47c-4f94066fdc9a","added_by":"auto","created_at":"2024-08-01 17:22:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":826697,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4273139/v1/5c0c7795-617c-41a7-a168-ae0b577b078b.pdf"}],"financialInterests":"","formattedTitle":"Prediction Model For Digital Image Tampering Using Customized Deep Neural Network Techniques","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eIn the digital age, the manipulation of images has become a prevalent and concerning issue. With the proliferation of image-editing software and the ease of sharing visuals on various platforms, image tampering has evolved into a significant challenge for preserving the integrity and authenticity of digital media. Detecting image tampering is not only essential for maintaining trust and credibility in various domains but is also crucial for criminal investigations, journalism, and ensuring the accuracy of visual content in social media. Traditional methods for image tampering detection have relied on manual examination, expert analysis, and the application of heuristics. While these approaches can be effective in some cases, they often fall short when dealing with sophisticated tampering techniques. Convolutional Neural Networks (CNNs), a potent category of deep learning models that have shown impressive abilities in tackling intricate computer vision challenges. This paper delves into the domain of image tampering detection and presents a complete performance analysis of four distinct CNN architectures: ResNet50V2, InceptionNetV3, MobileNetV2, and a custom-designed CNN architecture specifically crafted for this task. Image tampering is a powerful tool for creating deceptive visuals that can mislead, manipulate, or deceive the viewer. In the realm of journalism, manipulated images can be used to distort facts, sensationalize events, or even fabricate entirely fictitious stories. Misleading visuals can also be employed for political propaganda, harming the credibility of institutions and public trust [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. In criminal investigations, the integrity of digital evidence is paramount. Manipulated images can be used to alter crime scenes, obscure identities, or erase incriminating evidence. Detecting such tampering is crucial for law enforcement agencies and legal proceedings, as a single manipulated image can alter the course of justice [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].The rise of social media platforms and online communication has given individuals the power to share images globally. However, this convenience has also made it easier for malicious actors to disseminate fake or manipulated visuals, leading to the spread of misinformation and disinformation. Detecting and addressing image tampering is essential for maintaining the credibility of content shared on these platforms [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn addressing the pressing concern of fake detection, the integration of Generative Adversarial Networks (GANs) in deep learning has demonstrated significant prowess, particularly in the manipulation of faces, commonly referred to as Deepfake. Deepfake, driven by GANs, enables the realistic substitution of faces, posing a threat of spreading fabricated events through online channels. To counter this, [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] proposes an intelligent forensic method for Deepfake detection. Their approach focuses on subtle texture disparities in image saliency, particularly in facial textures, utilizing a guided filter with a saliency map to enhance texture artifacts. The ResNet18 classification network efficiently learns these features, achieving state-of-the-art detection accuracy in distinguishing between real and fake face images [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Generative adversarial networks (GANs) have been investigated for their ability to discern real images from fake ones based on texture differences in tampered images. However, challenges persist, as the proposed methodology still necessitates large-scale data to yield satisfactory results [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Recent efforts, exemplified by [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], have explored novel strategies leveraging resampling features, Long-Short Term Memory (LSTM) cells, and encoder-decoder networks. Their approach adeptly captures artifacts like JPEG quality loss and spatial transformations, enabling discriminative analysis between manipulated and non-manipulated regions. The method achieves pixel-level localization, demonstrating high precision in identifying image manipulations. The introduction of a new large image splicing dataset for training further enhances the robustness of their approach. Through rigorous experimentation on diverse datasets, [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] showcase the efficacy of their method in addressing the intricate challenges posed by subtle image manipulations.\u003c/p\u003e \u003cp\u003eRecent studies have identified various methods for tampering detection. However, a common challenge persists among these methods, as they often concentrate on specific manipulation techniques by extracting explicit features from the images [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Pre-trained models have been utilized for identifying manipulated images. ResNet50 attained higher accuracy; nevertheless, it suggested enhancing accuracy through the utilization of image patches [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Guided filters have been utilized for extracting texture features aimed at identifying tampered images [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. InceptionNet and MobileNet have been applied to the classification of medical images. However, there is a need to further enhance the accuracy to ensure broader acceptance of detection methodologies [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe various approaches discussed above, perform automatic feature extraction and recognition from image data, however, challenges arise in tasks such as determining valid data, handling de-duplication, and verifying data authenticity before the application of machine learning. The proposed approach provided better accuracy by feeding verified image data into different layers of CNN. This uniqueness distinguishes the approach, as existing methods often rely on machine learning techniques and authentication mechanisms in isolation.\u003c/p\u003e \u003cp\u003eThe objective of this paper is to formulate a robust predictive model employing deep learning techniques for the discernment of counterfeit images. This model is designed to accurately identify distinct classes of tampering applied to images, ensuring a comprehensive and precise detection mechanism. Furthermore, this paper thoroughly explores sophisticated feature extraction methods, aiming to enhance the precision and nuanced classification of images according to distinct tampering classes like copy-move, region removal, and splicing unlikely done in the extant research. The proposed approach of work enhances the accuracy of the model in less computation time.\u003c/p\u003e \u003cp\u003eThe paper is organized as follows; Following the introduction which presents the pressing issue of digital tampering of images, recent work done in this area, challenges, and objectives of the paper. Section \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e mentioned the techniques and meth-ods used to propose the custom CNN architecture used in this paper. Section \u003cspan refid=\"Sec3\" class=\"InternalRef\"\u003e3\u003c/span\u003e pre-sent the research methodology describing the dataset used and the proposed 4\u003c/p\u003e \u003cp\u003earchitecture. In section \u003cspan refid=\"Sec6\" class=\"InternalRef\"\u003e4\u003c/span\u003e, the experiment and results are mentioned followed by the conclusion in the last section \u003cspan refid=\"Sec10\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e"},{"header":"2. Methods and Techniques – Deep Neural Networks","content":"\u003cp\u003e \u003cb\u003eConvolutional Neural Networks (CNNs)\u003c/b\u003e have emerged as a groundbreaking technology in the field of computer vision [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], particularly in addressing complex image analysis tasks. CNNs are a class of deep neural networks designed to process and analyze visual data, making them highly effective in tasks such as image classification, object detection, and image tampering detection [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].CNNs consist of convolutional layers, which apply filters to input images, extracting features like edges and textures. These layers slide small filters over the image, computing dot products with local regions to generate feature maps. Pooling layers, like max-pooling and average-pooling, are integral components of CNNs and serve to decrease spatial dimensions, aiding computation and emphasizing important features. After convolution and pooling, fully connected layers capture advanced abstractions and relationships among features, facilitating intricate decision-making processes. Activation functions like ReLU [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] and Sigmoid introduce non-linearity, allowing CNNs to model intricate data relationships. CNNs undergo training through backpropagation, a process that involves adjusting internal parameters such as weights and biases to minimize a loss function, typically a measure of predicted vs. actual labels. In this paper analysis of four CNN architectures\u0026mdash;ResNet50V2, InceptionNetV3, MobileNetV2, and a custom-designed CNN specifically tailored for image tampering detection is shown by experimental results.\u003cb\u003eResNet50V2\u003c/b\u003e [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], short for Residual Network with 50 layers and Version 2 is particularly recognized for its outstanding performance in image classification. At its core, ResNet50V2 utilizes residual blocks with skip connections to mitigate the vanishing gradient problem in deep networks. The architecture incorporates bottleneck structures within these blocks to reduce computational load while maintaining representational capacity [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. \u003cb\u003eInceptionNetV3\u003c/b\u003e [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], or Inception-v3, is recognized for exceptional performance in image classification and object recognition, Inception-v3 employs specialized inception modules, allowing multi-scale feature capture and intricate pattern detection. It incorporates global average pooling to reduce dimensionality, aiding in mitigating overfitting, and benefits from batch normalization for stabilized training [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. \u003cb\u003eMobileNetV2\u003c/b\u003e [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], designed for efficient on-device computer vision on mobile and embedded devices, enhances its predecessor with innovative features. Linear bottlenecks and shortcut connections enhance information flow and stability [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. The scalable design includes width and resolution multipliers for flexibility in balancing model size and accuracy. Efficient convolutions reduce computational overhead, making MobileNetV2 ideal for real-time applications like image classification, object detection, and semantic segmentation [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAddressing the challenges posed by image tampering requires advanced tools capable of automatically identifying manipulated images with precision and accuracy. Convolutional Neural Networks (CNNs) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], a type of deep learning models, have shown immense promise in this regard. CNNs are designed to learn and extract intricate patterns and features directly from the data they are trained on [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. This ability makes them highly suitable for image tampering detection, as tampering techniques often involve subtle changes in pixel values, textures, and patterns. Traditional methods for image tampering detection often rely on manually crafted features and heuristics. CNNs, on the other hand, have the capacity to automatically discover and leverage complex features that may be challenging to define explicitly. This feature extraction capability allows CNNs to excel in identifying tampered regions [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], even in cases where the manipulation is sophisticated or subtle.\u003c/p\u003e \u003cp\u003eThere are various methods exist for determining the authenticity of an image, and these approaches are briefly examined. In the realm of tampered text detection in document images, the critical role of information security has drawn increasing attention. While progress has been made, detecting visually consistent tampered text in photographed document images remains challenging. Document Tampering Detector (DTD) framework, leveraging a Frequency Perception Head (FP H) and a Multi-view Iterative Decoder (MID) focuses on visual feature challenges. The novel Curriculum Learning for Tampering Detection (CLTD) training paradigm is designed to enhance robustness, mitigating confusion during training [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn the proposed work, the CNN model and its variations ResNet50V2, Inception-NetV3 and MobileNetV2 were used to define the proposed architecture of the customized CNN model explained in the following section \u003cspan refid=\"Sec3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e"},{"header":"3. Research Methodology","content":"\u003cp\u003eIn the proposed work, a more extensive approach is opted to fine-tune the pre-trained models. This involved unfreezing all layers of the models, not just the fully connected layers, allowing for modifications to both the convolutional and fully connected layers. The rationale behind this approach was to ensure that the models could adapt comprehensively to the unique demands of the image tampering detection task. The combination of the datasets, especially the augmented Defacto dataset, allows for a comprehensive assessment of CNNs in image tampering detection. The Defacto dataset provides tampered images with various scenarios and severity levels, further enhanced by rotation augmentation, while the COCO dataset offers diverse authentic images for assessing the models\u0026apos; ability to distinguish between tampered and normal images in real-world scenarios [\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e]. This dataset diversity and structure make them suitable for robust training, evaluation, and benchmarking of machine learning models dedicated to image tampering detection, offering an extensive and diverse dataset resource for research and experimentation in this field.\u003c/p\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n \u003ch2\u003e3.1. Dataset\u003c/h2\u003e\n \u003cp\u003eIn the context of image tampering detection, the choice of an appropriate dataset is crucial for training and evaluating machine learning models, including Convolutional Neural Networks (CNNs). The research paper discusses two key datasets: the Defacto dataset and the COCO dataset [\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e]. The Defacto dataset[[\u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e], containing 72,776 images per class after rotation augmentation, is specialized for image tampering detection. Multiple python scripts are developed for converting the raw images into tampered images of 4 categories such as copy-move, inpaint, and splicing,and normal with varying levels of tampering severity. Sample image form the each class is shown in the Fig.\u0026nbsp;1.Ground truth annotations for tampered regions within each image are provided, enabling supervised machine learning model training and performance evaluation.\u003c/p\u003e\n \u003cp\u003eOn the other hand, the COCO dataset serves as a valuable source for authentic, unaltered images, representing the \u0026quot;normal\u0026quot; class. It consists of diverse high-quality images capturing everyday scenes, objects, and contexts, with detailed annotations for object categories, locations, and segmentation masks.\u003c/p\u003e\n \u003cp\u003eThe combination of high-end GPUs and CPUs created a harmonious synergy that facilitated efficient model training, leading to optimal utilization of computational resources. In determining the training parameters, a batch size of 128 has been chosen and executed training for a span of 20 epochs. The selection of these parameters was a result of a carefully crafted balance between computational resources and model convergence. A batch size of 128 strikes a balance between using large enough batches for efficient gradient updates while not exceeding the memory capacity of the GPUs, which could lead to bottlenecks in training. The choice of 20 epochs for training duration was equally well-considered. It provided the models with ample opportunities to learn and adapt to the intricate features within the image tampering detection dataset while avoiding overfitting.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n \u003ch2\u003e3.2. Proposed Architecture:\u003c/h2\u003e\n \u003cp\u003eIn addition to the utilization of pre-trained models, custom CNN architecture is proposed, specifically designed to address the unique challenges posed by image tampering detection as shown in Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. While pre-trained models serve as powerful baselines, the development of a custom architecture was a crucial component of proposed research, as it allowed us to tailor the network to the specific requirements of this task.\u003c/p\u003e\n \u003cp\u003eAt the foundation lies the input layer, accepting raw image data. The initial CNN block forms the base, composed of two CNN layers and a Batch Normalization layer, establishing the groundwork for feature extraction. A subsequent max-pooling layer condenses information, marking a pivotal transition. The following are separable convolution blocks, represented as sturdy columns, emphasizing the model\u0026apos;s proficiency in discerning intricate features. Max-pooling layers intermittently guide the flow, each contributing to spatial reduction. The model evolves through these blocks, culminating in a compact global average pooling layer. Finally, a fully connected layer crowns the structure, epitomizing the culmination of feature extraction for precise classification of tampered class.\u003c/p\u003e\n \u003cp\u003eA distinguishing feature of proposed customized CNN architecture is the incorporation of separable convolution layers. Traditional CNNs employ standard convolution layers, which involve a high degree of interdependence between convolutional filters. In contrast, separable convolution layers separate the spatial and depth-wise convolutions. This separation significantly reduces computational complexity and enhances the model\u0026apos;s capacity to learn features more efficiently and rapidly. It\u0026apos;s important to emphasize that this innovation wasn\u0026apos;t just about computational efficiency; it also brought tangible benefits to the model\u0026apos;s representational power. The custom CNN, along with separable convolution layers, became a crucial asset in the comparative analysis. This architecture not only demonstrated its potential for more efficient feature learning but also highlighted the importance of architectural innovations when tackling image tampering detection.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"4. Experiments and Results","content":"\u003cp\u003eA comprehensive evaluation of the model\u0026rsquo;s performance unveils distinctive trends for loss and accuracy as mentioned in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The custom model, crafted with separable convolution layers, displays the highest validation accuracy and training accuracy among the models being assessed. It notably surpasses the performance of both InceptionNetV3 and ResNet50V2, attaining a validation accuracy that stands out significantly as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003e. The results are also mirrored in its training accuracy, where the proposed customized model displays robust learning capabilities.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy and Loss of Deep Learning Models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabb\" border=\"1\"\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eValidation Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTraining Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTraining Loss\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eValidation Loss\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCustom model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9000\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9688\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0137\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1043\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInceptionNetV3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.7148\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9240\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.5184\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.4958\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResNet50V2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.6716\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.7196\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.5068\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.5854\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMobileNetV2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.7727\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9026\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.5718\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.6068\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eFurthermore, the custom model excels in minimizing training loss and validation loss, demonstrating its effectiveness in image tampering detection. In contrast, InceptionNetV3, though renowned for its feature extraction prowess [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], achieves 71.48% validation accuracy as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003e. This is further emphasized by its comparatively higher training and validation loss values which are 51.84% and 49.58% respectively, indicating a struggle to achieve good precision. ResNet50V2 model known for its depth and skip connections [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], demonstrates 67.16% validation accuracy in comparison, emphasizing the advantages of customization as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e4\u003c/span\u003e. Its training accuracy falls short of expectations, and the training and validation losses remain relatively high, indicating a challenge in capturing the intricate features essential for tampering detection.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eMobileNetV2 [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003e while showing a competitive validation accuracy and training accuracy which is 77.27% and 90.26% respectively, slight lags in the validation loss section, which could be attributed to its inherent trade-offs in terms of computational efficiency. Nonetheless, the findings underline the superior performance too, which successfully combines efficient feature learning with faster convergence, showcasing its potential as a promising solution for image tampering detection.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the evaluation of image tampering detection models, precision, recall, and F1 score serve as crucial performance metrics. In this context, the results highlight the exceptional performance of the custom CNN, showcasing its capability in achieving a harmonious balance between precision and recall. The custom model achieved an impressive F1 score of 0.96, indicating its proficiency in both identifying tampered regions within images and minimizing false positives as mentioned in Table\u0026nbsp;2.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabc\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"char\" char=\".\" colname=\"c1\"\u003e \u003cp\u003eTable\u0026nbsp;2. Precision, Recall and F1 Score of Deep Learning Models\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabd\" border=\"1\"\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1 Score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInceptionNetV3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.8586\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.8541\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.8538\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResNet50V2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.7409\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.7263\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.7278\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMobileNetV2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.8153\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.8158\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.8154\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCustom model\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.9562\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.9683\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.9622\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe proposed customized model demonstrates a strong F1 score of 0.9622, indicating a symmetric precision and recall. In comparison, InceptionNetV3 achieves an F1 score of 0.8538, ResNet50V2 records 0.7278, and MobileNetV2 attains 0.8154. These scores signify the nuanced trade-off between precision and recall for each model, offering insights into their respective strengths in identifying tampered regions within images, positioning it as an asset in applications. In results and analysis, it is evident that training times for each model significantly impact their practical utility. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e7\u003c/span\u003e, The proposed CNN outshines the competition with a training time of 573.4 seconds for 20 epochs, which is a substantial improvement over the other models. In comparison, MobileNetV2 required 2250.5 seconds, InceptionNetV3 took 2413.5 seconds, and ResNet50V2 demanded 2220.5 seconds for the same number of epochs. What's particularly striking is that proposed model, despite having a similar number of parameters as MobileNetV2, reduced the training time by nearly 75%.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThis reduction in training time showcases the efficiency of the proposed customized CNN, thanks in part to the utilization of separable convolution layers, which expedite the learning process and enhance overall training efficiency. A comprehensive evaluation of the model parameters underscores the efficiency and compactness of the proposed customized CNN architecture in image tampering detection. With a remarkably lean parameter count of 5,450,988, the custom model showcases an adept balance between complexity and performance as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e8\u003c/span\u003e. In contrast, InceptionNetV3, ResNet50V2, and MobileNetV2 exhibit substantially larger parameter counts, registering at 25,878,802, 27,908,998, and 5,047,298, respectively. This notable discrepancy in parameter sizes emphasizes the streamlined architecture of custom model, demonstrating that good performance need not necessitate an excessive number of parameters. The judicious management of parameters in customized CNN not only contributes to computational efficiency but also underscores its potential for applications where resource constraints are a critical consideration.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Hyperparameter Optimization\u003c/h2\u003e \u003cp\u003eHyper parameter optimization is a critical component of the proposed research methodology. To guarantee the optimal fine-tuning of custom CNN and the pretrained models (ResNet50V2, InceptionNetV3, and MobileNetV2) for image tampering detection, an extensive and systematic hyperparameter search has been used. The hyperparameter search involved running hundreds of training sessions, each characterized by a unique combination of hyperparameters. Key parameters that were subjected to this optimization process included learning rates, weight decay, dropout rates, and architecture-specific settings. Each of these hyperparameters plays a crucial role in the performance and behavior of a neural network. Therefore, determining the most suitable values for these hyperparameters was paramount. By systematically exploring the hyperparameter space, the model gained insights into the intricate interplay between these parameters and the models' ability to generalize and adapt effectively to the proposed work. This comprehensive approach to hyperparameter optimization reflects to ensuring that model, reached its full potential. It not only enhanced the performance of the models but also contributed to the reliability and robustness of comparative analysis, providing a solid foundation for the findings and conclusions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Optimization Techniques:\u003c/h2\u003e \u003cp\u003eIn proposed work, the choice of optimization techniques played a crucial role in fine-tuning the pretrained CNN architectures for image tampering detection. Adam optimizer has been used which is widely acclaimed and renowned for its efficacy in training deep neural networks. Adam combines the benefits of both the momentum-based updates. This adaptive optimization algorithm dynamically modifies the learning rates for individual parameters, leading to accelerated convergence and more effective training. To further enhance the training performance and prevent the models from converging prematurely or getting stuck in local minima, learning rate scheduling mechanism has been implemented. Learning rate scheduling is a dynamic adjustment of the learning rate during training. This mechanism monitored the model's performance, and when it detected a plateau in accuracy, it automatically decreased the learning rate.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Addressing Overfitting:\u003c/h2\u003e \u003cp\u003eOverfitting, a common hurdle in deep learning, occurs when a model excessively tailors itself to the training data, often incorporating noise and irrelevant patterns that hinder its ability to generalize effectively to unseen data. To tackle this issue, various strategies were executed with the aim of improving the models' ability to perform well on data they had not encountered before. First and foremost, L1 regularization has been applied which was a vital tool in arsenal against overfitting. L1 regularization works on to the network's weights, encouraging sparsity within the model. This regularization technique introduced a penalty on the magnitude of the weights, promoting a more parsimonious set of features. By doing so, L1 regularization encouraged the model to focus on the most relevant and informative features while discouraging the overemphasis on noise and irrelevant details in the training data. This process played a significant role in enhancing the model's performance with unseen data, a critical factor in image tampering detection where adaptability and accuracy are paramount. The dropout layers are incorporated at strategic points in the network architecture. These dropout layers played a pivotal role in preventing overfitting by randomly deactivating a fraction of neurons during training. This randomness introduced a degree of uncertainty into the model's learning process, effectively discouraging it from becoming overly reliant on specific features or neurons. Furthermore, batch normalization [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] was a key component of the proposed strategy to combat overfitting. This technique was applied to standardize the input to each layer during training. By doing so, batch normalization mitigated the effects of internal covariate shift [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], a phenomenon where the distribution of layer inputs changes during training. Batch normalization aided in improving the models' generalization by ensuring that the learning process was not hindered by abrupt shifts in data distributions. It contributed to the overall stability and robustness of the models, which is paramount in the context of image tampering detection where detecting subtle variations and manipulations are essential.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Conclusion \u0026 Future Scope","content":"\u003cp\u003eIn conclusion, the investigation into Convolutional Neural Network (CNN) architectures for image tampering detection has unfolded notable advancements, notably exemplified by the proposed customized model. This exclusive CNN exhibited discernible proficiency, particularly in validation accuracy, performing better than the other considered architectures. Its precision in identifying tampered regions underscores its potential significance in digital forensics and content verification. Moreover, the custom model's noteworthy efficiency, characterized by reduced training time, stands as a testament to the impact of leveraging separable convolution layers. This computational efficacy not only conserves resources but also extends the model's adaptability to more extensive datasets and complex tasks. While recognizing the merits of established architectures, the suggested exploration emphasizes the inventive potential and adaptability encapsulated in neural network customization. Beyond the pursuit of elevated accuracy and training efficiency, this research contributes to the ongoing evolution of image tampering detection methodologies, paving the way for nuanced applications and reinforcing the digital landscape's integrity and trustworthiness. In future different image pre-processing technique could be used to enhance the performance. Additionally, a potential future scope of work could encompass the detection of fake videos. This expansion would involve adapting and extending the proposed model to address the unique challenges posed by video content.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eMolina MD, Sundar SS, Le T, Lee D (2021) Fake News Is Not Simply False Information: A Concept Explication and Taxonomy of Online Content. Am Behav Sci 65(2):180\u0026ndash;212. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/0002764219878224\u003c/span\u003e\u003cspan address=\"10.1177/0002764219878224\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoussa AF (2021) Electronic evidence and its authenticity in forensic evidence. Egypt J Forensic Sci 11(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s41935-021-00234-6,hp\u003c/span\u003e\u003cspan address=\"10.1186/s41935-021-00234-6,hp\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA\u0026iuml;meur E, Amri S, Brassard G (2023) Fake news, disinformation and misinformation in social media: a review. Social Netw Anal Min 13(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s13278-023-01028-5\u003c/span\u003e\u003cspan address=\"10.1007/s13278-023-01028-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eO\u0026rsquo;Shea K, Nash RR, An (2015) introduction to convolutional neural networks. arXiv (Cornell University). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arxiv.1511.08458\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.1511.08458\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe K, Zhang X, Ren S, Sun J (2016) Identity mappings in deep residual networks. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arxiv.1603.05027\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.1603.05027\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. arXiv (Cornell University)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSzegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2015) Rethinking the inception architecture for computer vision. arXiv (Cornell University). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arxiv.1512.00567\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.1512.00567\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSandler M, Howard A, Zhu M, Zhmoginov A, Chen L (2018) MobileNetV2: Inverted residuals and linear bottlenecks. arXiv (Cornell University). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arxiv.1801.04381\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.1801.04381\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgarap AF (2019) Deep Learning using Rectified Linear Units (ReLU). arXiv [Cs.NE]. Retrieved from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1803.08375\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1803.08375\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQu C, Liu C, Liu Y, Chen X, Peng D, Guo F, Jin L (2023) Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5937\u0026ndash;5946\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang J, Xiao S, Li A, Lan G, Wang (2021) Detecting fake images by identifying potential texture differences. Future Generation Comput Syst 125:127\u0026ndash;135. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.future.2021.06.043\u003c/span\u003e\u003cspan address=\"10.1016/j.future.2021.06.043\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBappy JH, Simons C, Nataraj L, Manjunath BS, Roy-Chowdhury AK (2019) Hybrid LSTM and Encoder\u0026ndash;Decoder Architecture for Detection of Image Forgeries. IEEE Trans Image Process 28(7):3286\u0026ndash;3300. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/tip.2019.2895466\u003c/span\u003e\u003cspan address=\"10.1109/tip.2019.2895466\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eManjunatha S, Malini M, Patil (2021) Deep learning-based Technique for Image Tamper Detection, Proceedings of the Third International Conference on Intelligent Communication Technologies and Virtual Mobile NetworksICICV EEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNagaveni K, Hebbar, Kunte AS (2021) Transfer Learning Approach For Splicing And Copy-Move Image Tampering Detection. Ictact Journal On Image And Video Processing, May\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiachen Ya, Xiao S (2021) Aiyun Li a, GuipengLan, HuihuiWangb, Detecting fake images by identifying potential texture difference. Future Generation Computer Systems, Elsevier\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eImage classification and prediction using transfer learning in colab notebook, J Praveen Gujjara,\u0026lowast;, H R Prasanna Kumar b, Niranjan N, Chiplunkar (2021) Global Transitions Proceedings\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiachen Ya, Li SXA, Wang GLH (2021) Detecting fake images by identifying potential texture difference. Future Generation Computer Systems, ELSEVIER \u0026ndash;\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin T, Maire M, Belongie S, Bourdev L, Girshick R, Hays J, Perona P, Ramanan D, Zitnick CL, Doll\u0026aacute;r P, Microsoft COCO (2015) Common Objects in context. arXiv (Cornell University). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arxiv.1405.0312\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.1405.0312\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMAHFOUDI G, TAJINI B, RETRAINT F, MORAIN-NICOLIER F, DUGELAY JL and M. PIC, DEFACTO: Image and Face Manipulation Dataset, 2019 27th European Signal Processing Conference (EUSIPCO), Coruna A (2019) Spain, pp. 1\u0026ndash;5, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.23919/EUSIPCO.2019.8903181.(\u003c/span\u003e\u003cspan address=\"10.23919/EUSIPCO.2019.8903181.(\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e2014)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIoffe S, Szegedy C (2015) Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. arXiv [Cs.LG]. Retrieved from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1502.03167\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1502.03167\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"international-journal-of-system-assurance-engineering-and-management","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ijsa","sideBox":"Learn more about [International Journal of System Assurance Engineering and Management](http://link.springer.com/journal/13198)","snPcode":"13198","submissionUrl":"https://www.editorialmanager.com/ijsa/default2.aspx","title":"International Journal of System Assurance Engineering and Management","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Convolutional Neural Network, Image Tampering, Fake Images, Image Forensic, Deep Learning","lastPublishedDoi":"10.21203/rs.3.rs-4273139/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4273139/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eImage tampering detection is a critical area of research, given the widespread use of manipulated images for deceptive purposes. Convolutional Neural Networks (CNNs) have shown significant potential in automating the identification of tampered images. This paper presents customized deep learning model to detect tampering class with comparative analysis of CNN architectures - ResNet50V2, InceptionNetV3, MobileNetV2, and the proposed CNN, for image tampering detection. The proposed approach encompasses a dataset comprising four distinct classes: copy-move, inpaint, splicing, and normal images. This study sheds light on the comparative strengths and weaknesses of these CNN architectures. The dataset encompasses the key tampered classes, offering a holistic assessment of each model's ability to identify various tampering techniques. The custom CNN architecture is specifically tailored for this task, aiming to evaluate its efficiency compared to the established CNNs. Metrics for training and evaluation are standardized to generate equitable comparisons, encompassing performance indicators such as accuracy, precision, recall, and F1-score.This research contributes the knowledge in the field of image tampering detection, offering a comprehensive evaluation of multiple CNN architectures. Additionally, the effectiveness of separable convolutional layers is explored in deep neural networks, showcasing their potential to enhance scalability and effectiveness across various tasks in machine learning and computer vision. The proposed model, designed with separable convolution layers, exhibits superior validation accuracy and training accuracy compared to the other models under evaluation. Notably, The proposed customized model achieved an impressive F1 score of 96%, highlighting its proficiency in accurately detecting tampered regions within images while minimizing false positives.\u003c/p\u003e","manuscriptTitle":"Prediction Model For Digital Image Tampering Using Customized Deep Neural Network Techniques","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-13 18:49:49","doi":"10.21203/rs.3.rs-4273139/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Minor revisions","date":"2024-06-24T11:51:00+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2024-06-07T05:11:44+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-05-30T18:13:53+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"International Journal of System Assurance Engineering and Management","date":"2024-05-07T08:04:51+00:00","index":"","fulltext":""},{"type":"submitted","content":"International Journal of System Assurance Engineering and Management","date":"2024-05-06T08:56:15+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"international-journal-of-system-assurance-engineering-and-management","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ijsa","sideBox":"Learn more about [International Journal of System Assurance Engineering and Management](http://link.springer.com/journal/13198)","snPcode":"13198","submissionUrl":"https://www.editorialmanager.com/ijsa/default2.aspx","title":"International Journal of System Assurance Engineering and Management","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"74bcab6f-3ce4-476c-bf6a-35d764c26e32","owner":[],"postedDate":"June 13th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-08-01T16:20:59+00:00","versionOfRecord":{"articleIdentity":"rs-4273139","link":"https://doi.org/10.1007/s13198-024-02420-w","journal":{"identity":"international-journal-of-system-assurance-engineering-and-management","isVorOnly":false,"title":"International Journal of System Assurance Engineering and Management"},"publishedOn":"2024-07-19 16:13:27","publishedOnDateReadable":"July 19th, 2024"},"versionCreatedAt":"2024-06-13 18:49:49","video":"","vorDoi":"10.1007/s13198-024-02420-w","vorDoiUrl":"https://doi.org/10.1007/s13198-024-02420-w","workflowStages":[]},"version":"v1","identity":"rs-4273139","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4273139","identity":"rs-4273139","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0