Temporal Convulutional Network & Content-based Frame Sampling Fusion for Semantically Enriched Video Summarization

preprint OA: closed
View at publisher

Abstract

Due to the exponential growth of video data, which makes it difficult to handle massive volumes of video content, video summarizing techniques have drawn a lot of interest recently. Although there are numerous video summary techniques, summarizing lengthy videos is still difficult since it takes a long time to process hundreds of frames. The Distributed Temporal Convolutional Networks method, which takes into account the shot length distribution's properties to effectively summarize lengthy films, is proposed in this paper as a novel solution to tackle this difficulty. By completely using the temporal coherence of lengthy films, the Content-Based Frame Sampling [3] technique is additionally suggested to improve the system's throughput. The fusion approach has a significant implication for improving video summarization techniques and enabling efficient video processing, paving the way for new applications in areas such as video surveillance, entertainment, and education. All authors contributed to the writing and revision of the manuscript.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00