A Comparative Analysis of Convolutional Neural Network and Vision Transformer Embeddings on a Novel Domain-Specific Task
preprint
OA: closed
Abstract
Abstract The Vision Transformer (ViT) architecture utilized the self attention and transformer architecture originally designed for natural lan guage processing (NLP) enables ViTs to capture global relationships and long-range dependencies within images. The purpose of our study was to compare the performance of embeddings generated by Convolutional Neural Network (CNN) and Vision Transformers (ViT) on a novel domain-specific task which was not presented at any point to the models prior to the process of the embeddings generation. The pretrained CNN model was MobileNetV2, and the pretrained vision transformer was ViT-B16. The accuracy and F score obtained from the embeddings generated by MobileNetV2 were 0.64 and 0.69 respectively. The accuracy and F-score obtained from the embed dings generated by ViT-B16 were 0.81 and 0.79 respectively. Our study suggests that ViT might perform better in unseen domain specific problems which were not presented in the pretraining. ViT utilizing the self-attention mechanisms capture rich and generic visual representations that might gen eralize well to unseen domain specific problems.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00